Categorizing Crowd Emotions based on Cross Division Expressions and Anomalies

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract The crowd emotion sensing is a critical element in surveillance and management of the crowd in different environments. With exploding populations, and developing nations, the crowd in urban cities mandate state of art surveillance methodologies involving continuous monitoring and reporting of criminal activities. The research article presents a novel technique to compute the spatial and temporal features obtained from the crowd environments and combine the novelty of neural networks for detecting the emotions of crowds with better accuracy and swiftness. The features are obtained from the continuous feed of surveillance videos typically categorized into the common features of human beings namely anger, sadness, disgust, surprise, fear, happiness and obviously neutrality. Such features are extracted after careful background separation which are typically difficult in crowded environments, using techniques namely SIFT, and FAST termed to be the visual descriptors. Once the features are extracted, spatial and temporal features are classified into individual and combined features as defined in the cross-division environment in order to portray the crowd dynamics and characteristics. Cross division environment computes the necessary features for identifying the anomalies in the crowded situations in a neural network, after a series of operations such as dimensionality reduction, and principal component analysis. From the semantic information, crowd behaviours are detected based on interactive features in a dynamic environment and the proposed technique has demonstrated effective results in terms of 98.9% accuracy in detecting especially violence in crowd datasets collected from UMN.
Full text 151,816 characters · extracted from preprint-html · click to expand
Categorizing Crowd Emotions based on Cross Division Expressions and Anomalies | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Categorizing Crowd Emotions based on Cross Division Expressions and Anomalies Manojkumar K, Suji Helen L This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5709790/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The crowd emotion sensing is a critical element in surveillance and management of the crowd in different environments. With exploding populations, and developing nations, the crowd in urban cities mandate state of art surveillance methodologies involving continuous monitoring and reporting of criminal activities. The research article presents a novel technique to compute the spatial and temporal features obtained from the crowd environments and combine the novelty of neural networks for detecting the emotions of crowds with better accuracy and swiftness. The features are obtained from the continuous feed of surveillance videos typically categorized into the common features of human beings namely anger, sadness, disgust, surprise, fear, happiness and obviously neutrality. Such features are extracted after careful background separation which are typically difficult in crowded environments, using techniques namely SIFT, and FAST termed to be the visual descriptors. Once the features are extracted, spatial and temporal features are classified into individual and combined features as defined in the cross-division environment in order to portray the crowd dynamics and characteristics. Cross division environment computes the necessary features for identifying the anomalies in the crowded situations in a neural network, after a series of operations such as dimensionality reduction, and principal component analysis. From the semantic information, crowd behaviours are detected based on interactive features in a dynamic environment and the proposed technique has demonstrated effective results in terms of 98.9% accuracy in detecting especially violence in crowd datasets collected from UMN. crowd emotion sensing anomaly detection spatial temporal behaviour cross division Figures Figure 1 Figure 2 Figure 3 1 INTRODUCTION In developing nations, the rapid urbanization has become the regions more crowded and prone to unsafe environments for vulnerable individuals. Despite have numerous measures for video surveillance and monitoring of the crowded environments, the crime rate has been alarmingly high in cities. Major events primarily the religious, sports, and festive seasons have been prone to increasing criminal occurrences in urban cities, according to reports in cities like Delhi, Chennai and Pune [ 1 ]. The chances of detecting the anomalies in a crowded environment become more challenging due to the dynamic features of a crowd and hence need for automated techniques to detect the anomalies becomes a primary concern for governments and other organizations. Since the crowd features are completely dynamic and unpredictable, determining the future actions of every individual becomes more challenging. The recent developments in machine learning algorithms involving deep learning, computer vision and automated techniques have opened up a new research domain named as crowd behaviour analysis [ 2 ]. Detecting the anomalies has been a booming area of research for monitoring the strange and abnormal activities of certain individuals in a crowded environment. As soon as the domain has recognized enough anomalies, public safety can be ensured by preventing the accidents, controlling the flow of individuals in the crowded environments and thereby protecting the innocent individuals in chaotic situations. The detection of challenging situations has to be in real-time and thereby alerting the officials in order to prevent further damages to lives and properties. In most of the models, the feature engineering processes play a critical and inevitable role in determining the accuracy of predicting the chaotic situations in a crowded environment. Feature engineering is a process that identifies significant elements that are found to be potential features for representing the crowd behaviour [ 3 ]. These features clearly discriminate the chaotic movement and behaviour from a normal vs abnormal crowd. Yet, the accuracy of the determination is completely based on the quality of images, surveillance videos, background features along with the specificity of the abnormal event that has to be detected. This research work aims to consider the shortcomings of the existing methods, and contribute to the betterment of quicker detection of abnormal events in the crowds. The amount of information surrounding every individual is extremely huge and it takes only a glance for an average individual to process the information and put it to use. The ability to observe, contemplate, and perceive the information obtained from the surroundings has been extended to powerful ensemble algorithms. From the recovered information, statistical data such as mean, median, variance, distribution, relationship between the data and much more are estimated for understanding the situation of the current scenarios. The previous studies on ensemble algorithms have been extensively covered to illustrate the power of visual key descriptors for image and video analytics. The commonly observed attributes of primary importance are the size, orientation, hue, with respect to facial emotions, motions of the individuals and as a crowd, background information along with economic values [ 4 ]. The emotions of the individuals and the crowd depends on the features observed from the backgrounds too. From the existing research works, the ensemble algorithms considered the frequently changing features from a rich feature set obtained from the crowd surveillance videos. The threshold value for each emotion is identified around the mean value for each continuum of feature dimensions. The features may be considered as the immediate change of one emotion to another within a stipulated time, ranging from one to another, small to huge, low intensity to high intensity, thereby guiding the ensemble algorithms to determine the average emotion from emotion dimensions. Unique member identification from a set of emotions can thus be generated by approximating the information collected from the set of features, by recognizing the standalone set of emotions. Despite the measure of change of emotions being estimated in quantity, the stimuli play an important factor in the frequency of changes from happy to sad or to anger with respect to quantity again. All such facial expressions, are further enhanced to improve the quality of detections. Since the proposed work contemplates the cross dimension of various feature sets, every emotion is defined with a distinct emotion observed from the approximation from the feature sets along with the captured stimuli. However, the relationships between the individuals of the crowd were not considered in the ensemble algorithms based on the cross dimensions and categories. Given this scenario, the generalized feature extraction algorithms could not derive the minor overlapping of the feature sets, and hence the need for ensemble algorithms arise. The cross dimensions of the features are hence combined with a feature mapping across various dimensions forming a cross category for better evaluations of emotions. From the previous research work, the models were able to determine the actions and emotions based on low level stimuli from the environments. The observations were made to classify the actions from circled emotions captured from simultaneous frames that are processed in a spatial-temporal sets of features [ 6 ]. The processes ensured that the individuals from the crowded environments are categorized into different classes based on the unique characteristics and thus forming appropriate subsets of feature dimensions. Statistics obtained from each subset depict the summary of various dimensions uniquely assigned to every emotion. Depending on the characteristics of individuals namely the tallness, shortness, fairness, presence of objects and other combinations, multiple cross category feature sets were computed for including monitoring numerous situations accordingly. From the updated cross category feature dimensions, it is observed that the quality of observations improved with considerable accuracy. The feature distribution was formed based on the combinations of different individuals from different dimensions including the category items. Numerically, the peak distributions were observed twice indicating the two-peaks of approximations, in place of a single distribution. The other works extended the previous research works to measure the number of hues collected from the continuous perceptions levels thereby establishing the relationships between various individuals and members of the feature dimension sets. A notable experiment was conducted to identify the hue value and categorize the dimension based on the unique feature set from every cross-category feature sets. Every individual hue property was compared to identify whether it is a new hue property or it exists in other dimension sets. From the investigative results of the previous works, the hue properties were the results of average extraction of individual and unique hue properties from actually derived hues of multiple hue properties. All these factors contributed to the need of ensemble algorithms for deriving multiple feature sets and cross category dimension sets [ 7 ]. The extent to which the category relationship of set members with high-level features, influences the combination of algorithm remains unclear wherever the facial features are used. While low-level features may differ fundamentally from processing of the high level features through the algorithms, it is challenging to directly apply findings from low-level feature studies to high-level feature scenarios. In everyday life, groups of interacting faces are often diverse in their categorical expressions, both spatially and temporally. For instance, individuals within a group may exhibit varying attitudes toward a specific event, or an individual's expression may shift significantly due to unforeseen circumstances. This leads to scenarios involving crowds with mixed-category emotions presented simultaneously or sequentially. Understanding whether perceivers can form averaged representations from such heterogeneous emotional expressions is essential, as the ability to derive statistical information from a group of emotional faces plays a critical role in daily life and overall well-being. This study aimed to investigate whether perceivers could extract average expressions from sets of cross-category facial expressions through two experiments, where facial expressions were presented either spatially or temporally [ 8 ][ 9 ]. The focus was on happy and fearful expressions rather than the more commonly studied pair of happy and angry expressions. Happiness and anger are both associated with approach motivation—happiness with well-being and anger with attack. In contrast, fear belongs to a distinct emotional category and is associated with avoidance motivation, representing the opposite end of the motivational spectrum from happiness. This distinction makes happy and fearful expressions more clearly categorized and widely used in studies on the categorical perception of facial expressions. Faces near the categorical boundary were selected as set members for two main reasons. First, in real-life interactions, individuals often display subtle and ambiguous expressions rather than distinct, prototypical emotions. Second, prior research indicates that increased variance within a set diminishes the ability to derive an average representation, which could obscure potential effects of ensemble coding in cross-category groups. Additionally, it is uncommon to encounter groups of individuals simultaneously expressing entirely different facial emotions or an individual rapidly transitioning between extreme emotional categories. While studies on categorical perception of emotional faces support basic emotion theory, ensemble coding of facial expressions aligns with dimensional emotion theory. This perspective suggests that observers perceive facial expressions as part of a continuous spectrum rather than discrete categories. 2 RELATED WORKS The objective of the anomaly detection algorithms is to detect the changes in momentum, velocity and obviously the emotions of individuals in a crowded environment. A significant benefit of the anomaly detection algorithm is to detect the changes immediately and report the abnormal events to the concerned officials. Various methods have been proposed and the primary methods contemplate the anomaly detection based on objects and holistic approaches [ 10 ]. Object based approaches concentrate on the segmentation of individuals into smaller groups, monitoring the trajectories, predicting the trajectories and focus on object-based attributes to extract the probable behaviours of individuals and as a crowd. However, the factors such as occlusion and the presence of multiple target objects are inhibiting the visibility of anomaly detection algorithms. These factors greatly affect the accuracy of predictions and hence the need for better approaches. Holistic approaches, on the other hand, operate on a wider network of individuals where all are interconnected to provide more meaning to the scenes. The low level and mid-level features are extracted for monitoring the crowd behaviours. Holistic approaches concentrated on optical flow fields, that highlighted the low-level and mid-level features for predicting the outcome with better accuracy. An automated method for deriving the optical flow histograms were proposed for monitoring the overall motion of individuals, thereby connecting them in a wide network of people. In chaotic situations such as the stampedes [ 5 ], the proposed histograms were fruitful in identifying the potentially critical situations especially in a crowded environment. Another technique suggested the utilization of low to mid-level features in the motion aspect specifically to visualize the magnitude of the crowd along with the direction of dispersions. The segmentation algorithms proposed in the model were enabled to predict the motion of the crowd during chaotic situations by modelling the regions and probable motion of individuals during normal vs abnormal events. A probabilistic model based on Riemannian detection [ 11 ] for detecting crowd anomalies was proposed that handled the optical flow with respect to walking, running and dispersion at different velocities in various directions. Comparatively, the performance of models that have taken the regions of interest from the respective frames has been higher than the models that processed the entire video or frames of the videos [ 12 ]. The technique was to consider the particle advection specifically for quicker processing and deriving the prediction outcomes. Such techniques also considered the trajectory information of every individual with respect to spatial-temporal features and changes observed in different time frames. Motion patterns identified from the different video inputs have shown significant similarities in chaotic situations, thereby establishing a remarkable similarity in crowd anomaly detections. A technique named Histogram of Oriented Tracklets (HOT) was suggested based on the motion described by the histograms that illustrated the motion features, magnitude and the direction of individuals. In case of highly populated areas, the patterns were able to track the anomalies better than the conventional methods. Crowd Anomaly Detection has gained popularity in the research domain in recent years owing to automated surveillance techniques, and numerous techniques have been introduced. The challenges in the current techniques are also accounted for and the advent of machine learning, deep learning and computer vision algorithms has simplified the entire process flow. Various articles contemplated the list of techniques used for monitoring the individual behaviours and specifically segmenting the abnormal events in crowded environments. The techniques contemplated in the literature review have been collectively describing the various crowded environments, anomalies and target objects that can be the reason for the anomalies [ 13 ]. Such techniques emphasized the need for machine learning, deep learning and computer vision algorithms for automated monitoring systems that can be a live-saving measure for crowded environments. Video analysis techniques played a significant role in anomaly detection models for adding more intelligence to the algorithms. Primarily, the intelligent systems for monitoring the anomalies and crowds used the spatial and temporal features along with perturbations, yet the models needed to address the uneven, non-repeating, unique and rare cases of abnormal events. A better approach implementing a Convolutional Neural Network with multiple optimization techniques was designed and delivered to improve the accuracy of crowd anomaly detection. The need for automated surveillance techniques and thus the accuracy of detection and prediction was justified in the article. Crowded environments were processed in a neural network as continuous streams of input video, right from the cameras and the algorithm detected the anomalies with various threshold features. Ability to process the crowded environments from surveillance videos was tested in a technique that enforced the optimized version of multiple Convolutional Neural Networks. Despite the huge number of individuals in the crowded environments, and a diverse range of anomalies, the optimized Convolutional Neural Network ensemble performed considerably better than the previous techniques said in the literature survey. Another significant benefit of this approach was the considerably low computational cost and the efficiency when compared to the traditional techniques, apart from the difficulties faced by the other techniques in handling real-time surveillance videos. The next technique approached the anomalies detection problem with a two-step process, where the objects and individuals were identified using a You Only Look Once (YOLOv5) version 5 model and a model named DeepSORT was implemented for tracking the trajectories of the individuals [ 14 ]. Optical flow, yet again, was helpful in determining the features that highlight the abnormal behaviours and events of the individuals in the crowded environment. The spatial features were derived from the bounding boxes, compared with the threshold features and the abnormalities were determined. Support Vector Machines were the classifiers that compared the features of the captured events, and compared against the threshold features. From the experimental results, the SVM exhibited an Area Under the Curve metric of 88.9% and the same model has outperformed the other conventional approaches of detecting abnormal events [ 15 ][ 16 ]. The same model has proven to perform remarkably well while processing the videos of Hajj yatras, indicating the significant performance during the presence of a dense crowd. Moreover, the model has been able to identify seven distinct abnormal events captured during the yatras, even in the huge crowd. Since the further categorization has been narrowed down to seven more distinctive categories, the accuracy of classifying the abnormal events has improved significantly. Combination of multiple individual techniques have showcased the increase in accuracy and area under the curve of classification of abnormal events. Ensemble of state of art machine learning techniques [ 17 ] and multiple modalities of classification algorithms, bounding box techniques and feature extraction techniques have been adapting to fit the current landscape of crowd anomaly detection in the recent years [ 18 ]. The following Table 1 has detailed all the models and techniques discussed in the literature survey that have produced the models for crowd anomaly detections. Table 1 Summary of Models and Techniques used for Crowd Anomaly Detection and their performance Considered Datasets Models / Techniques Used Performance Metric Outcome ShangaiTech Violent Flow Generative Adversarial Networks Area Under Curve 73.8% Recurrent Neural Networks, 2D Convolutional Neural Networks Accuracy 93.5% Optical Flow Accuracy 73.56% Optical Flow GAN Accuracy 79.6% Area Under Curve 98.1% UCF Crime CNN Residual Logn Short Term Memory Area Under Curve 70.4% CNN, Random Forest Area Under Curve 88.98% Hajj Yatra Optical Flow Area Under Curve 88.96% UMN SVM Area Under Curve 88.29% 3 PROPOSED METHODOLOGY The proposed model in the research article contemplates the functioning of crowd anomaly detection methodology with two unique techniques namely, the illustration of the crowded environments, their attributes, definition of various behaviours, followed by the dimensionality reduction and classification of behaviours according to the attributes [ 19 ]. The surveillance video is processed in a series of frames, from which a set of feature sets are derived to form a list of probably trajectories and Delaunay triangles. From the frames, the visual attributes are further categorized into spatial and temporal features primarily being the velocity, density, motion descriptors and other feature attributes indicating the state of the crowd members. These features are critically acclaimed to the important features that define the quality background upon which the classifiers act upon and determine the abnormal events observed in the given environments [ 20 ]. The second part of the methodology involves the processing of the feature sets into a set of histograms after aggregation for computational purposes. This section applies the process of dimensionality reduction through the Principal Component Analysis and autoencoders enabling the classification process. Classification is a typical element of neural networks which categorizes and lists the captured events into two primary classes known to be the normal and abnormal events. The following Fig. 1 illustrates the functionality of the proposed model and the following sections explain the processes in detail. 3.1 Crowd Behaviour Analysis The processes commence with the definition of visual descriptors that explain the characteristics of individuals and the entire crowd. Typically for an anomaly detection system, the input videos are processed directly from the surveillance videos captured in real-time. Given the installation of surveillance cameras in major cities, homes and crowded public places, the availability of surveillance videos from places has increased in recent years [ 21 ]. The frames are thoroughly analysed for regions of interest that may hold the potential of analysis followed by the definition of the feature points. Pixels of specific regions of interest with clear and concise events of abnormality are defined from such regions of interest and thus the other regions are removed from the processing of neural networks. The commonly applied techniques for identifying such events from the frames are FAST, SIFT and AKAZE [ 22 ], that are known for the quickness and accuracy of detecting such feature sets from subsequent images. Every frame is analysed through these renowned algorithms to narrow down the features of interest and following them in a series of images thereby ensuring that the abnormal event is prolonged for a specific duration. Such feature sets usually depict the characteristics of objects, individuals and the surroundings detected in a surveillance video. The identical feature sets in subsequent frames describe the valuable information that has to be processed for features tracking and predicting the trajectories of every individual in the crowd. In order to add more sense to the predictions, the objects are considered as well, to detect the abnormal events well in advance. The spatial features are better derived from the Delaunay triangulation strategy, where the spatial elements are highlighted and differentiated between the subsequent frames. Such features are further defined as the visual descriptors which are classified into individual and combined features accordingly [ 23 ]. 3.1.1 Feature Extraction and Tracking In the proposed approach, the Features from Accelerated Segment Test (FAST) technique was applied to identify the pixels of interest from the frames of significant importance and the features were compared with the continuous frames to match for similarities [ 24 ]. The variations of intensity around the circled region of interest explains the continuous changes in the trajectories and on the other hand, the similar activities in the regions of interest in the neighbouring frames indicate the capture of abnormal events. In order to reduce the computational complexity, the corner criteria technique is applied to identify the predetermined areas of interest on a specific frame. As soon as a specific feature is identified using FAST algorithm, Scale Invariant Feature Transform (SIFT) algorithm is applied to derive a 16x16 neighbourhood matrix around the detected feature thereby deriving a 128-bin value. The equation represents the Difference of Gaussians (DoG) [ 25 ], which approximates the Laplacian of Gaussian (LoG) for edge detection in image processing. This difference isolates features that vary in size between the scales, highlighting image details at a specific range of scales. The resulting difference is then convolved (*) with the input image I(x,y) to detect edges or features using the following expression 1, where x, y are the source data to which the Difference of Gaussians filter is applied to detect features. Convolution integrates the filter with the image to detect the desired features at each pixel. The scales 𝑘𝜎 and 𝜎 indicate the difference between the two functions. D(x,y,σ)=[G(x,y,kσ) − G(x,y,σ)]∗I(x,y) (1) In this approach, the feature extraction capabilities of SIFT and FAST along with AKAZE were individually analysed, highlighting their unique contributions to anomaly detection in sparse feature tracking. SIFT is well-known for its ability to remain robust under variations in scale, rotation, and lighting conditions. This method is particularly effective in identifying distinct key points within crowd images, allowing it to capture intricate patterns that signal deviations from typical crowd behaviour. As a result, it is highly suitable for detecting anomalies of different sizes. FAST, on the other hand, is optimized for quick corner detection, enabling the rapid identification of key features in crowd images [ 26 – 28 ]. While it lacks the scale and rotation invariance of SIFT and FAST, its high speed makes it an excellent choice for real-time anomaly detection applications, where quick responses are essential. The efficient detection of key points by FAST enhances its utility in sparse feature tracking for recognizing unusual crowd behaviours. Sparse feature tracking is a widely utilized technique in computer vision for tracking a subset of key features across video frames. Unlike dense tracking methods, which monitor every single pixel, this approach zeroes in on distinctive features within the frames. This focus on unique features such as corners or structural elements enables it to handle challenges like occlusion (where parts of the crowd are hidden) and dynamic changes in crowd behaviours more effectively. Sparse tracking excels by identifying and following these prominent key points [ 29 ], ensuring robust performance even when parts of the crowd are obscured or when movement patterns shift unpredictably. One notable advantage of sparse feature tracking is its ability to re-detect and associate features over time. This ensures the system adapts to evolving crowd dynamics, allowing for accurate long-term tracking. By prioritizing features that are simple to identify and track, our approach maintains its effectiveness across various scenarios [ 30 ]. The Accelerated-KAZE (AKAZE) algorithm extends the original KAZE algorithm by using a computationally efficient framework known as Fast Explicit Diffusion (FED) to create its non-linear scale spaces. Built upon non-linear diffusion filtering, AKAZE utilizes the determinant of the Hessian matrix for feature detection. To enhance rotation invariance, it employs Scharr filters. The maximum responses from these detectors pinpoint specific feature point locations, forming the basis for AKAZE’s strong and distinctive feature detection capabilities. The AKAZE descriptor [ 31 ] leverages the Modified Local Difference Binary (MLDB) algorithm, renowned for its power and efficiency. Due to the non-linear nature of AKAZE’s scale spaces, the algorithm exhibits invariance to scale, rotation, and limited affine transformations. Furthermore, the distinctiveness of AKAZE’s features is maintained and even enhanced as they are scaled up or down. Our methodology capitalizes on these local features and incorporates the Lucas–Kanade optical flow algorithm. This algorithm is particularly adept at handling the challenges posed by unrestricted optical flow, making it ideal for applications like crowd anomaly detection. The Lucas–Kanade method stands out due to its precision, robustness, and adaptability to complex environments [ 32 ]. By emphasizing sparse feature tracking, we substantially reduce computational overhead while maintaining high accuracy, perfect for real-time applications. The Lucas–Kanade algorithm enhances its effectiveness in detecting anomalous behaviours by capturing subtle motion variations within crowded environments. Ensuring temporal coherence in feature tracking stabilizes the system and minimizes false positives, thereby increasing the reliability of anomaly detection mechanisms. This approach operates on the assumption that neighbouring pixels within a small, localized region share consistent optical flow values. Instead of analysing every pixel in the frame, it calculates optical flow based on these groups. This method provides several benefits, such as faster computations and the efficient generation of training data. Mathematically as expressed in Eq. 2, the optical flow constraint for a set of pixels moving at the same velocity can be expressed by a specific equation, ensuring a cohesive and efficient analysis of movement within the crowd [ 33 ]. By leveraging these principles, our methodology strikes a balance between computational efficiency and accuracy, making it a powerful tool for real-time crowd monitoring and anomaly detection. I x (x 1 , y 1 ) ・ v x + I y (x 1 , y 1 ) ・ v y = − I t (x 1 , y 1 ) I x (x 2 , y 2 ) ・ v x + I y (x 2 , y 2 ) ・ v y = − It(x 2 , y 2 ) .. . I x (x n , y n ) ・ v x + I y (x n , y n ) ・ v y = − I t (x n , y n ) (2) This efficient and adaptive approach to sparse optical flow makes it a promising solution for crowd anomaly detection in diverse surveillance and monitoring scenarios [ 34 ]. 3.1.2 Trajectory Detection The Lucas Kanade method may not function in case of object detection especially in motion and the trajectories may not be detected effectively. This raises the requirement to include the gradient transformation for considering the neighbouring pixels of the respective frames to detect the object motions. The proposed approach introduces a novel approach for considering the pixels from regions of interest in form of a pyramidal format. This technique contemplates the techniques of down-sampling, passing through a low-pass filtering technique and applying a factor of 2. The purpose of optical flow inclusion is justified when the low-quality images are processed first [ 35 ], followed by the high-quality frames from the same videos, in order to increase the optical flow accuracy. Once the spatial features are computed, the significant features are forwarded to the next stage of processing as tracklets. The series of frames, where the objects are defined to be the primary features of consideration, are transformed into a graph with mapped tracklets. Delaunay Triangulation graph technique [ 36 ] processes the features in omnidirectional nodes in the neighbouring frames in order to trace the local, spatial and temporal features or tracklets. The number of nodes depends on the tracklets and the node connections are represented as \(\:{ϵ}^{n}\) and the relationships between the tracklets are represented by \(\:{g}^{n}\left({\vartheta\:}^{n},{ϵ}^{n},{F}^{n}\right).\:\) The number of triplet features are represented as \(\:{F}^{n}\) . Temporal features are effectively described with respect to the topographical features over the different states of time without affecting the shape of the triangle in a graph eliminating the noise and occlusion factors. A graph containing the tracklets derived from the previous stages, the local tracklets are identified as cliques and neighbouring nodes are identified as seed points, which forms the tracking points for objects and individuals in a crowded environment. The cliques are represented according to the following Eq. 3 . $$\:C\left({V}_{i}^{k}\right)=\:\left\{{V}_{i}^{k}\right\}\:\cup\:\:\left\{{V}_{j}^{k},\:\forall\:\:\left({V}_{i}^{k},\:{V}_{j}^{k}\right)\:\in\:\:{ϵ}^{n}\right\}$$ 3 The connections between the cliques or the tracklets are connected through the short-term or long-term connections representing the probable trajectories. In terms of spatial features, the cliques are connected on varying terms increasing the dynamic ability of the crowd. With a more diversified crowd, the number of cliques with temporal and spatial features are comparatively higher than insignificant cliques. 3.1.3 Individual vs Crowd Behaviour Visual Descriptors Identification Once the visual descriptors are defined and mapped in a graph, the characteristics of objects and individuals are further analysed for classification. Visual descriptors indicate the semantic information about the crowd participants, typically the spatial and temporal aspects. Information captured from the input videos are highly dynamic and critical for the further analysis [ 37 ]. The proposed system carefully analyses the individual and collective features, thereby contemplating the entire set of features from the surveillance videos. Such visual descriptors are further classified into individual and entire crowd-based features namely the collective features. Individual behaviours are processed to segment the participants of the crowd, understanding their activities by observing the dynamic properties such as direction of the flow and velocity of individuals. The direction of motions, describes the individuals with respect to normal and extreme conditions in varying tracklets motions [ 38 ]. A complete structure of the tracklets, neighbouring nodes, are represented by the F segment as described in the following Eq. 4. $$\:\left\{{S}_{i}^{n},\:{S}_{i}^{n-{\tau\:}_{2}},\dots\:.\:{S}_{i}^{n-\left(F-1\right){\tau\:}_{2}}\right\}\:\left(4\right)$$ The changes in the directions of the individuals are observed by monitoring the angular variations of every trajectory and tracklets. The following Eq. 5 provides the calculation for predicting the change of direction in a given trajectory. $$\:{D}^{var\:}\left({V}_{i}^{k}\right)=\:\frac{1}{F}\:.\:\sum\:_{0}^{F-2}{d}_{\theta\:}\left({S}_{i}^{n-f{\tau\:}_{2}},\:{S}_{i}^{n-(f+1){\tau\:}_{2}}\right)$$ 5 Where \(\:\theta\:\) represents the angular variations, F is represented by vectors of the graph. On the other critical element for determining the individual characteristics, velocity of motion vectors in a specific direction explains the seriousness of the situation occurred in the crowded environment [ 39 ]. The accuracy of the visual descriptors depends on the considered frame in the history of frames. The motion vectors suitably describe the information about the motion of individuals using the following Eq. (6). Euclidean distance is the measure between the tracklets in any direction, preferably between the current node and origin, instead of measuring the sum of all the tracklets between the nodes. $$\:{D}^{velocity}\left({V}_{i}^{k}\right)=\:\frac{1}{{\tau\:}_{1}}\:.\:‖\underset{{V}_{i}^{k-\tau\:1}{V}_{i}^{k}}{\to\:}‖\:\left(6\right)$$ In a crowded environment, it is extremely important to derive the collective characteristics of all the participants and the collective behaviours in this section explains the need for collective visual descriptors. The characteristics of the crowded collective features are derived from the stability, collectiveness, density and uniformity [ 40 ]. The concept of stability in crowd analysis reflects the degree of consistency in the crowd's topological structure over time. It evaluates how individuals within a crowd maintain their proximity to the same neighbours as time goes on. By examining the stability property, we can glean valuable insights into persistent patterns and relationships within the crowd, thereby enhancing our understanding of its dynamics and behaviours. The collectiveness property pertains to how pedestrians move as a cohesive group. This property is measured by calculating each individual's directional deviation from the overall movement of the group. Traditionally, coherent motion has been assessed using predefined collective transitions. However, in this approach, cliques are utilized for the local computation of this descriptor, offering a nuanced alternative. Conflict, an important property in crowd analysis, captures interactions among individuals, especially when they are in close proximity [ 41 ]. Similar to the computation of the collectiveness descriptor, the conflict property is also determined locally. The local density descriptor focuses specifically on the spatial distribution aspect of the model. Unlike previous interactive descriptors, it emphasizes the spatial arrangement of individuals within the scene. This descriptor captures a key characteristic of crowd behaviours: the distribution of individuals. An approximate measure of local density can be obtained by evaluating the proximity of nearby features. This is based on the observation that when nearby features converge, it signifies a higher probability of a larger crowd forming in that area. The uniformity descriptor assesses the coherence of the spatial distribution of regional features. It indicates whether a group has a tendency to cluster together in a uniform manner or to fragment into smaller subgroups, reflecting non-uniform behaviours. 3.1.4 Dimensionality Reduction and Classification The previous section explained the list of descriptors that worked upon the features with spatial and temporal features that were predominant in identifying the abnormal events in the crowded environments. Dimensionality reduction [ 42 ] is a renowned technique for processing the potential list of features by reducing the rich feature sets and processing only the required critical information. The proposed approach reduced the dimensionalities using two remarkable techniques namely principal component analysis and autoencoding. These two approaches are common in computer vision applications and have been proven to reduce the computational cost associated with processing images and videos. The first technique employed was principal component analysis (PCA). PCA aims to transform the original features into a new set of uncorrelated variables, known as principal components, while preserving as much variance as possible. By projecting the data onto these components, we effectively reduce the information into a lower-dimensional space. The detection of anomalies within crowd dynamics is a critical challenge in visual crowd analysis. This paper presents a comprehensive approach that leverages dimensionality reduction techniques and neural network architectures to effectively identify unusual crowd behaviours. The process begins with the application of principal component analysis to reduce the dimensionality of the input data. By capturing the most significant features, PCA helps to prepare the data for the subsequent classification stage. The second approach involves using autoencoders, a type of neural network designed to learn efficient representations of the input data. Autoencoders consist of an encoder network that compresses the input into a latent space representation and a decoder network that reconstructs the original input from this representation. By training the autoencoder to minimize the reconstruction error, it learns to capture the most important features of the data in the latent space. After applying dimensionality reduction using PCA and autoencoders, we proceed to the classification stage. In this step, we utilize the reduced-dimensional feature set to train neural network classifiers. Specifically, we employ neural networks like the multi-layer perceptron. Neural networks are well-suited to handle non-linear problems and extract intricate patterns from the input data. The adaptability of neural networks is especially useful in detecting anomalies in crowd behaviours. They excel at uncovering subtle relationships within the data and identifying unusual crowd dynamics that might be overlooked by traditional methods. Due to their hierarchical architectures, neural networks can capture both detailed and higher-level representations, providing a comprehensive understanding of crowd behaviours. Throughout our study, we evaluated neural networks with varying numbers of hidden layers to determine the optimal configuration for our analysis. Through systematic testing and performance comparisons, we identified the most effective approach for crowd anomaly detection (Favarelli & Giorgetti, 2020). This rigorous approach ensured that our methodology was robust and capable of addressing the complexities of detecting anomalies within crowd dynamics. 4 RESULTS AND DISCUSSIONS The proposed approach exemplifies the performance of crowd anomaly detection through a series of investigations with state of art datasets known to have captured crowd activities. The datasets possess various crowd activities and is a standard dataset to be used in this specific research domain. University of Minnesota (UMN) along with other datasets that have captured violent and unsettling events have been used for measuring the performance of the proposed approach. Various scenes in the 11 videos of the datasets depict the normal and abnormal (violent) behaviours in subsequent frames, thereby having scenes of people running in a single direction, different directions, moving from the same point of origin with high velocity, and other atypical scenes of abnormal nature. YouTube is also the repository that contains numerous videos of unsettling crowded environments with different tension among the participants, and other violent scenes. Surveillance videos are curated into a list of videos containing a balanced number of normal videos and violent videos. From the UMN datasets, the proposed approach has delivered the visual descriptors for sensing the different behaviours and the YouTube videos were used to measure the accuracy of predicting the abnormal events in challenging situations. The following Fig. 2 illustrates the spatial features captured by the Delaunay triangulation technique which explains the distribution of crowd using the following metrics. Each individual in the crowd is marked as a distinct point. Triangles are formed such that no point lies inside the circumcircle of any triangle. The edges of these triangles represent the proximity of individuals to one another, capturing the spatial relationships within the crowd. The spatial features between individuals and cluster of people will be tracked using the visual descriptors identified by the feature extraction techniques of the proposed model. As soon as the features are marked and identified to be potentially significant features, the Delaunay Triangulation technique will mark the presence of individuals as distinct points, a group of such distinct points marks a crowd with normal activities. Upon the presence of abnormal activities, the events are marked as red, with varying velocity, the Delaunay triangulation will mark the edges in green coloured lines representing different clusters of people/participants. The cohesion among the group participants will be lesser in case of an abnormal event, thereby dispersing the crowd in different directions. The clusters will be less crowded owing to the events and the Delaunay triangulation clearly differentiates the normal vs abnormal crowds as shown in Figs. 2 (a) and 2(b) respectively. When the crowd is found to be agitated due to an abnormal event, the distance between the nodes/distinctive points is found to be longer than the nodes with normal events. The length between the nodes will be consistent in a normal crowd and uniformity is a critical element. Length of the green lines will be longer and scattered when the participants are dispersed or scattered in different directions when a chaotic environment is found in the crowd. The following Table 2 has listed the performance measures in terms of accuracy and area under the curve (AUC) metrics based on the quality of feature extraction techniques applied in the proposed model. The configuration of the neural networks is tested for three different scenarios where a normal neural network is applied without any dimensionality reduction techniques, and an autoencoding technique is applied in the next configuration where the dimensions are reduced to 128 bits. The last variant is the performance of the proposed method where the contrast features are reduced based on Principal Component Analysis (PCA) where the variance value is considered to be 95. For testing purposes, a 5-fold cross-validation technique is applied for classifying the behaviours. Table 2 Performance of different techniques for detecting chaotic scenes in input videos Techniques Neural Network 128-Bits Dimensionality Reduction NN PCA – NN Accuracy AUC Accuracy AUC Accuracy AUC SIFT 0.803 0.805 0.799 0.802 0.848 0.846 FAST 0.848 0.853 0.844 0.847 0.844 0.851 A-KAZE 0.858 0.862 0.865 0.890 0.923 0.898 It is observed that SIFT feature extraction technique has performed better than the other two techniques in terms of crowded environments and in terms of dimensionality reduction, the PCA based neural network has shown the highest accuracy and area under the curve. The spatial information and crowd features were best described with a SIFT based neural network, configured with principal component analysis-based dimensionality reduction. From the classification results tabulated in Table 2 , the videos taken from the YouTube sources were processed better than the inputs from UMN datasets. Given that the challenges in the YouTube videos were present, the proposed model has shown promising results as highlighted in the given table. The following Figs. 3 illustrates the results captured across the three methods and showcases the classification accuracy of the three techniques. Descriptors obtained through the processing of SIFT, FAST and AKAZE upon the violent videos and UMN dataset, crowd anomaly detection was more balanced. From the inputs of UMN dataset, the videos were more structured and the scenarios were limited, posing a clear-cut approach for training the models. The events in the UMN dataset were more controlled and hence the possessed lesser challenging environments. On the other hand, the videos from the YouTube repository were highly dynamic and more challenging. The quality of the videos was comparatively lesser on the YouTube, and hence added more constraints and longer processing time for understanding the scenes better. Conversely, the reduced performance on the YouTube dataset could be due to issues such as poor video quality, occlusion, and varied surveillance scenarios. Addressing these specific challenges within each dataset is crucial to enhance the effectiveness of crowd anomaly detection methods. The proposed model achieved an AUC of 89.67% and 97.50% for scenes with normal activities and abnormal activities, respectively when the input from UMN dataset was processed. The input videos from the UMN dataset were processed, where the accuracy and AUC performance parameters are tabulated in the below Table 3 . From the comparative analysis, our approach has shown superior performance relative to other methods as listed below. Specifically, our approach achieved 99.5%, 96.5%, and 99% for normal and abnormal activities of the UMN dataset, respectively. Moreover, our approach shows a more pronounced advantage on the YouTube videos, where the proposed approach has delivered 89.67% AUC and 88.5% accuracy in detecting the abnormal activities. From the investigative results, consistent performance of our approach across both datasets indicates its high effectiveness in classification tasks, leveraging unique features and techniques that contribute to its superior performance. The following Table 3 lists down the comparative results of classification accuracy with other state of art techniques and models. Table 3 Performance of different Models for classification Models AUC Accuracy Optical Flow 84% 87.3% Sparse Reconstruction 90.1% 88.1% Visual Descriptors 88% 89.1% GAN 87.65% 87.2% Proposed Approach 89.67% 88.5% 5 CONCLUSION This study introduces an innovative model by leveraging the combination of visual descriptors, feature extraction and neural networks for detecting the abnormal events occurring in a crowded environment. The proposed method's effectiveness, demonstrated through the investigative results on the renowned UMN and crowd activities datasets on YouTube, highlights its potential in accurately detecting abnormal crowd behaviours. SIFT and AKAZE descriptors along with the neural networks, along with the impressive performance of the neural network configurations, underscores the robustness of our approach. Consequently, the proposed work makes a significant contribution to the evolving field, laying a solid foundation for future advancements in detecting the abnormal events in a crowded environment. The enhancements to the conventional neural networks with better feature descriptors and architectural updates to the deep learning systems improve the accuracy of detecting abnormal events. The patterns of abnormal behaviours are processed collectively to provide meaningful insights to the models for future detection. The proposed model has been tested against the real time videos curated from YouTube and UMN datasets for applicability in real time environments. The model has been tested recursively with different scenarios for addressing the various complex situations and other factors limiting the processing ability of the proposed model. In order to enhance the real time implications, the proposed model has to incorporate multimodal inputs from various sources to be more adaptable and promising. Declarations Author Contribution We hereby declare that this manuscript is the result of our collaborative effort. Manojkumar K contributed to the conceptualization, methodology, and writing of the manuscript. Suji Helen provided significant contributions to data analysis, interpretation, and critical revisions of the manuscript. Both authors have read and approved the final version of the manuscript and agree to its submission to the Iranian Journal of Science and Technology. References Manoj Kumar K, Sujihelen L (2022) Recognising Actions with Segmentation and Prediction Techniques in ROI based Deep Learning Framework. Math Stat Eng Appl 71(4):4072–4090. https://doi.org/10.17762/msea.v71i4.977 Sujihelen L (2021) Behavioural Analysis For Prospects In Crowd Emotion Sensing: A Survey, Third International Conference on Inventive Research in Computing Applications (ICIRCA) , Coimbatore, India, 2021, pp. 735–743. 10.1109/ICIRCA51532.2021.9544607 Halboob W, Altaheri H, Derhab A, Almuhtadi J (2024) Crowd Management Intelligence Framework: Umrah Use Case, in IEEE Access , vol. 12, pp. 6752–6767. 10.1109/ACCESS.2024.3350188 Liu D, Liu W, Yuan X, Jiang Y (2023) Conscious and Unconscious Processing of Ensemble Statistics Oppositely Modulate Perceptual Decision-Making. Am Psychol 78:346–357 List of Human Stampedes and Crushes (2022) [online] Available: https://en.wikipedia.org/wiki/List_of_human_stampedes_and_crushes Aljuaid H, Akhter I, Alsufyani N, Shorfuzzaman M, Alarfaj M, Alnowaiser K, Jalal A, Park J (2023) Postures anomaly tracking and prediction learning model over crowd data analytics. PeerJ Comput Sci. 9, e1355 Wang Y, Luo X, Zhou Z (Aug. 2024) Contrasting Estimation of Pattern Prototypes for Anomaly Detection in Urban Crowd Flow. IEEE Trans Intell Transp Syst 25(8):10231–10245. 10.1109/TITS.2024.3355143 Al-Shaery AM et al (2024) Open Dataset for Predicting Pilgrim Activities for Crowd Management During Hajj Using Wearable Sensors, in IEEE Access , vol. 12, pp. 72828–72846. 10.1109/ACCESS.2024.3402230 Luo L, Li Y, Yin H, Xie S, Hu R, Cai W (2023) Crowd-level abnormal behavior detection via multi-scale motion consistency learning, Proc. AAAI Conf. Artif. Intell. , pp. 8984–8992 Luo L, Xie S, Yin H, Peng C, Ong Y-S (2024) Detecting and Quantifying Crowd-Level Abnormal Behaviors in Crowd Events. IEEE Trans Inf Forensics Secur 19:6810–6823. 10.1109/TIFS.2024.3423388 Liao X-C, Chen W-N, Guo X-Q, Zhong J, Hu X-M (2023) Crowd management through optimal layout of fences: An ant colony approach based on crowd simulation, IEEE Trans. Intell. Transp. Syst. , vol. 24, no. 9, pp. 9137–9149, Sep Abdullah F, Abdelhaq M, Alsaqour R, Alatiyyah MH, Alnowaiser K, Alotaibi SS et al (2023) ,Context aware crowd tracking and anomaly detection via deep learning and social force model. IEEE Access 11:75884–75898 Ding X, He F, Lin Z, Wang Y, Guo H, Huang Y (2021) Crowd density estimation using fusion of multi-layer features, IEEE Trans. Intell. Transp. Syst. , vol. 22, no. 8, pp. 4776–4787, Aug Noor TH (2023) Behavior analysis-based IoT services for crowd management, Comput. J. , vol. 66, no. 9, pp. 2208–2219, Sep Alsubai S et al (2024) Design of Artificial Intelligence Driven Crowd Density Analysis for Sustainable Smart Cities, in IEEE Access , vol. 12, pp. 121983–121993. 10.1109/ACCESS.2024.3390049 Luo L, Li Y, Yin H, Xie S, Hu R, Cai W (2023) Crowd-level abnormal behavior detection via multi-scale motion consistency learning, Proc. AAAI Conf. Artif. Intell. , pp. 8984–8992 Zhao R et al (2022) Oct., Dynamic crowd accident-risk assessment based on internal energy and information entropy for large-scale crowd flow considering COVID-19 epidemic, IEEE Trans. Intell. Transp. Syst. , vol. 23, no. 10, pp. 17466–17478 Yang T, Wang C, Zhou T, Cai Z, Wu K, Hou B (2022) Identification of anomalous behavioral patterns in crowd scenes. Comput Mater Continua 71(1):925–939 Yadav S, Gulia P, Gill NS, Chatterjee JM (2022) A real-time crowd monitoring and management system for social distance classification and healthcare using deep learning, J. Healthcare Eng. , vol. pp. 1–11, Apr. 2022 Long J, Liang W, Li K-C, Wei Y, Marino MD (Feb. 2023) A regularized cross-layer ladder network for intrusion detection in industrial Internet of Things. IEEE Trans Ind Informat 19(2):1747–1755 Li Y, Xie Z, Li B, Mohiuddin M (2022) The impacts of in situ urbanization on housing mobility and employment of local residents in China, Sustainability , vol. 14, no. 15, p. 9058, Jul Yu G, Wang S, Cai Z, Liu X, Zhu E, Yin J (2023) Video anomaly detection via visual cloze tests. IEEE Trans Inf Forensics Secur 18:4955–4969 Zhang M, Li T, Yu Y, Li Y, Hui P, Zheng Y (2022) Urban anomaly analytics: Description detection and prediction, IEEE Trans. Big Data , vol. 8, no. 3, pp. 809–826, Jun Lalit R, Purwar RK (2022) Crowd abnormality detection using optical flow and GLCM-based texture features, J. Inf. Technol. Res. , vol. 15, no. 1, pp. 1–15, Jun Nguyen TN, Zeadally S (May 2022) Mobile crowd-sensing applications: Data redundancies challenges and solutions. ACM Trans Internet Technol 22(2):1–15 Zhang S, Yang Y, Liang W, Sandor VKA, Xie G, Choo KR (2023) MKSS: An effective multi-authority keyword search scheme for edge-cloud collaboration, J. Syst. Archit. , vol. 144 Martí P, Serrano-Estrada L, Nolasco-Cirugeda A, Baeza JL (2022) Revisiting the spatial definition of neighborhood boundaries: Functional clusters versus administrative neighborhoods, J. Urban Technol. , vol. 29, no. 3, pp. 73–94, Jul Chen C et al (2022) Jun., Comprehensive regularization in a bi-directional predictive network for video anomaly detection, Proc. AAAI Conf. Artif. Intell. , vol. 36, no. 1, pp. 230–238 Deng L, Lian D, Huang Z, Chen E (2022) Graph convolutional adversarial networks for spatiotemporal anomaly detection, IEEE Trans. Neural Netw. Learn. Syst. , vol. 33, no. 6, pp. 2416–2428, Jun Yang T, Wang C, Zhou T, Cai Z, Wu K, Hou B (2022) Identification of anomalous behavioral patterns in crowd scenes. Comput Mater Continua 71(1):925–939 Liu Y, Liu X, Li X, Li M, Li Y (2023) Participants recruitment for coverage maximization by mobility predicting in mobile crowd sensing, China Commun. , vol. 20, no. 8, pp. 163–176, Aug Chen Y, Deng A (2022) IEEE Access 10:93513–93524Using POI data and Baidu migration big data to modify nighttime light data to identify urban and rural area Cao C, Lu Y, Zhang Y (2024) Context recovery and knowledge retrieval: A novel two-stream framework for video anomaly detection. IEEE Trans Image Process 33:1810–1825 Liu CH et al (2021) Apr., Modeling citywide crowd flows using attentive convolutional LSTM, Proc. IEEE 37th Int. Conf. Data Eng. (ICDE) , pp. 217–228 Zhang J, Zhang X (2023) Multi-task allocation in mobile crowd sensing with mobility prediction, IEEE Trans. Mobile Comput. , vol. 22, no. 2, pp. 1081–1094, Feb Wang H, Zeng S, Li Y, Jin D (2021) Predictability and prediction of human mobility based on application-collected location data, IEEE Trans. Mobile Comput. , vol. 20, no. 7, pp. 2457–2472, Jul Liu T, Zhang C, Lam K-M, Kong J (2023) Decouple and resolve: Transformer-based models for online anomaly detection from weakly labeled videos. IEEE Trans Inf Forensics Secur 18:15–28 Woo G, Liu C, Sahoo D, Kumar A, Hoi S (2022) CoST: Contrastive learning of disentangled seasonal-trend representations for time series forecasting, Proc. Int. Conf. Learn. Represent. , pp. 1–18 Liu X, Liu J (2022) A truthful double auction mechanism for multi-resource allocation in crowd sensing systems, IEEE Trans. Serv. Comput. , vol. 15, no. 5, pp. 2579–2590, Sep./Oct Wu P et al (2024) VadCLIP: Adapting vision-language models for weakly supervised video anomaly detection, Proc. AAAI Conf. Artif. Intell. , pp. 6074–6082 Yang Y, Zhang C, Zhou T, Wen Q, Sun L (2023) DCdetector: Dual attention contrastive representation learning for time series anomaly detection, Proc. 29th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining (KDD) , pp. 3033–3045 Zhang S, He J, Liang W, Li K (2024) A secure and verifiable multimedia data search scheme for cloud-assisted edge computing. Future Gener Comput Syst 151:32–44 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5709790","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":394273225,"identity":"a1d260b1-84bc-4ba4-af14-02fa129b8928","order_by":0,"name":"Manojkumar K","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+UlEQVRIie3RsWoCMRjA8YTMcuvXhxAUIbQ0PonLycG5XPoEUlKEc3JXBH0FXaRjSsAp7a0Zla4O17Fbk1NxiudYMP8l3PH9+O4IQqHQfw1kdRD0I9yJ3+TNBM8EAktEPUEnQhoVOT97ao4/P3aPWvWiuSSd7jt7bY6V3TJkPR+h+iVpgVF8+hWThOsUqO5bsk258BGZUYBScaERUTxXQKUlWCg/KQ5HsnTkyZFiX0OM22I/bGVJgh0xdVvModMCPeBrjUftSZ4+bIzdEl/7lyJrf8P2mS80UfCbs4gWg/2uHDIvcZHqLuwNXl7FV8ar2bJmIBQKhe68PxN1ZimUSreFAAAAAElFTkSuQmCC","orcid":"","institution":"Sathyabama Institute of Science and Technology","correspondingAuthor":true,"prefix":"","firstName":"Manojkumar","middleName":"","lastName":"K","suffix":""},{"id":394273228,"identity":"006d7348-126c-4365-a10a-4e2d646a4cc4","order_by":1,"name":"Suji Helen L","email":"","orcid":"","institution":"Sathyabama Institute of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Suji","middleName":"Helen","lastName":"L","suffix":""}],"badges":[],"createdAt":"2024-12-25 07:38:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5709790/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5709790/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":72379767,"identity":"a3ae84f0-26fc-471c-b410-4c197bb569d7","added_by":"auto","created_at":"2024-12-26 08:54:55","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":43208,"visible":true,"origin":"","legend":"\u003cp\u003eArchitecture of the Proposed System\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5709790/v1/b1ca5ae73076cb907af2a7d6.png"},{"id":72380076,"identity":"cfa6c2ff-ade6-4b72-8328-ee082dbed09c","added_by":"auto","created_at":"2024-12-26 09:02:55","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":851538,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Marking the individuals as distinct points (b) Delaunay Triangulation showing distribution of crowd with anomaly detection represented as red coloured zones\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5709790/v1/b2fa9fe381d6c6e7336ff075.png"},{"id":72379771,"identity":"ea6834c1-7526-421b-aeb0-dd7a72cf91e0","added_by":"auto","created_at":"2024-12-26 08:54:55","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":59694,"visible":true,"origin":"","legend":"\u003cp\u003e(a): Performance of Visual Descriptors over UMN Dataset (b) Performance of Visual Descriptors over YouTube Video List\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5709790/v1/c367211f815eb8a36acb7eaf.png"},{"id":72555943,"identity":"4baa62dc-b5f5-484a-810f-b9696a77bbec","added_by":"auto","created_at":"2024-12-29 16:46:34","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1727108,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5709790/v1/6b88d1a0-68b3-4ae7-99df-1fd05352e1fa.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Categorizing Crowd Emotions based on Cross Division Expressions and Anomalies","fulltext":[{"header":"1 INTRODUCTION","content":"\u003cp\u003eIn developing nations, the rapid urbanization has become the regions more crowded and prone to unsafe environments for vulnerable individuals. Despite have numerous measures for video surveillance and monitoring of the crowded environments, the crime rate has been alarmingly high in cities. Major events primarily the religious, sports, and festive seasons have been prone to increasing criminal occurrences in urban cities, according to reports in cities like Delhi, Chennai and Pune [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The chances of detecting the anomalies in a crowded environment become more challenging due to the dynamic features of a crowd and hence need for automated techniques to detect the anomalies becomes a primary concern for governments and other organizations. Since the crowd features are completely dynamic and unpredictable, determining the future actions of every individual becomes more challenging. The recent developments in machine learning algorithms involving deep learning, computer vision and automated techniques have opened up a new research domain named as crowd behaviour analysis [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Detecting the anomalies has been a booming area of research for monitoring the strange and abnormal activities of certain individuals in a crowded environment. As soon as the domain has recognized enough anomalies, public safety can be ensured by preventing the accidents, controlling the flow of individuals in the crowded environments and thereby protecting the innocent individuals in chaotic situations.\u003c/p\u003e \u003cp\u003eThe detection of challenging situations has to be in real-time and thereby alerting the officials in order to prevent further damages to lives and properties. In most of the models, the feature engineering processes play a critical and inevitable role in determining the accuracy of predicting the chaotic situations in a crowded environment. Feature engineering is a process that identifies significant elements that are found to be potential features for representing the crowd behaviour [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. These features clearly discriminate the chaotic movement and behaviour from a normal vs abnormal crowd. Yet, the accuracy of the determination is completely based on the quality of images, surveillance videos, background features along with the specificity of the abnormal event that has to be detected. This research work aims to consider the shortcomings of the existing methods, and contribute to the betterment of quicker detection of abnormal events in the crowds.\u003c/p\u003e \u003cp\u003eThe amount of information surrounding every individual is extremely huge and it takes only a glance for an average individual to process the information and put it to use. The ability to observe, contemplate, and perceive the information obtained from the surroundings has been extended to powerful ensemble algorithms. From the recovered information, statistical data such as mean, median, variance, distribution, relationship between the data and much more are estimated for understanding the situation of the current scenarios. The previous studies on ensemble algorithms have been extensively covered to illustrate the power of visual key descriptors for image and video analytics. The commonly observed attributes of primary importance are the size, orientation, hue, with respect to facial emotions, motions of the individuals and as a crowd, background information along with economic values [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The emotions of the individuals and the crowd depends on the features observed from the backgrounds too. From the existing research works, the ensemble algorithms considered the frequently changing features from a rich feature set obtained from the crowd surveillance videos. The threshold value for each emotion is identified around the mean value for each continuum of feature dimensions. The features may be considered as the immediate change of one emotion to another within a stipulated time, ranging from one to another, small to huge, low intensity to high intensity, thereby guiding the ensemble algorithms to determine the average emotion from emotion dimensions. Unique member identification from a set of emotions can thus be generated by approximating the information collected from the set of features, by recognizing the standalone set of emotions. Despite the measure of change of emotions being estimated in quantity, the stimuli play an important factor in the frequency of changes from happy to sad or to anger with respect to quantity again. All such facial expressions, are further enhanced to improve the quality of detections. Since the proposed work contemplates the cross dimension of various feature sets, every emotion is defined with a distinct emotion observed from the approximation from the feature sets along with the captured stimuli. However, the relationships between the individuals of the crowd were not considered in the ensemble algorithms based on the cross dimensions and categories. Given this scenario, the generalized feature extraction algorithms could not derive the minor overlapping of the feature sets, and hence the need for ensemble algorithms arise. The cross dimensions of the features are hence combined with a feature mapping across various dimensions forming a cross category for better evaluations of emotions.\u003c/p\u003e \u003cp\u003eFrom the previous research work, the models were able to determine the actions and emotions based on low level stimuli from the environments. The observations were made to classify the actions from circled emotions captured from simultaneous frames that are processed in a spatial-temporal sets of features [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The processes ensured that the individuals from the crowded environments are categorized into different classes based on the unique characteristics and thus forming appropriate subsets of feature dimensions. Statistics obtained from each subset depict the summary of various dimensions uniquely assigned to every emotion. Depending on the characteristics of individuals namely the tallness, shortness, fairness, presence of objects and other combinations, multiple cross category feature sets were computed for including monitoring numerous situations accordingly. From the updated cross category feature dimensions, it is observed that the quality of observations improved with considerable accuracy. The feature distribution was formed based on the combinations of different individuals from different dimensions including the category items. Numerically, the peak distributions were observed twice indicating the two-peaks of approximations, in place of a single distribution. The other works extended the previous research works to measure the number of hues collected from the continuous perceptions levels thereby establishing the relationships between various individuals and members of the feature dimension sets. A notable experiment was conducted to identify the hue value and categorize the dimension based on the unique feature set from every cross-category feature sets. Every individual hue property was compared to identify whether it is a new hue property or it exists in other dimension sets. From the investigative results of the previous works, the hue properties were the results of average extraction of individual and unique hue properties from actually derived hues of multiple hue properties. All these factors contributed to the need of ensemble algorithms for deriving multiple feature sets and cross category dimension sets [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. The extent to which the category relationship of set members with high-level features, influences the combination of algorithm remains unclear wherever the facial features are used. While low-level features may differ fundamentally from processing of the high level features through the algorithms, it is challenging to directly apply findings from low-level feature studies to high-level feature scenarios. In everyday life, groups of interacting faces are often diverse in their categorical expressions, both spatially and temporally. For instance, individuals within a group may exhibit varying attitudes toward a specific event, or an individual's expression may shift significantly due to unforeseen circumstances. This leads to scenarios involving crowds with mixed-category emotions presented simultaneously or sequentially. Understanding whether perceivers can form averaged representations from such heterogeneous emotional expressions is essential, as the ability to derive statistical information from a group of emotional faces plays a critical role in daily life and overall well-being.\u003c/p\u003e \u003cp\u003eThis study aimed to investigate whether perceivers could extract average expressions from sets of cross-category facial expressions through two experiments, where facial expressions were presented either spatially or temporally [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e][\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. The focus was on happy and fearful expressions rather than the more commonly studied pair of happy and angry expressions. Happiness and anger are both associated with approach motivation\u0026mdash;happiness with well-being and anger with attack. In contrast, fear belongs to a distinct emotional category and is associated with avoidance motivation, representing the opposite end of the motivational spectrum from happiness. This distinction makes happy and fearful expressions more clearly categorized and widely used in studies on the categorical perception of facial expressions. Faces near the categorical boundary were selected as set members for two main reasons. First, in real-life interactions, individuals often display subtle and ambiguous expressions rather than distinct, prototypical emotions. Second, prior research indicates that increased variance within a set diminishes the ability to derive an average representation, which could obscure potential effects of ensemble coding in cross-category groups.\u003c/p\u003e \u003cp\u003eAdditionally, it is uncommon to encounter groups of individuals simultaneously expressing entirely different facial emotions or an individual rapidly transitioning between extreme emotional categories. While studies on categorical perception of emotional faces support basic emotion theory, ensemble coding of facial expressions aligns with dimensional emotion theory. This perspective suggests that observers perceive facial expressions as part of a continuous spectrum rather than discrete categories.\u003c/p\u003e"},{"header":"2 RELATED WORKS","content":"\u003cp\u003eThe objective of the anomaly detection algorithms is to detect the changes in momentum, velocity and obviously the emotions of individuals in a crowded environment. A significant benefit of the anomaly detection algorithm is to detect the changes immediately and report the abnormal events to the concerned officials. Various methods have been proposed and the primary methods contemplate the anomaly detection based on objects and holistic approaches [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Object based approaches concentrate on the segmentation of individuals into smaller groups, monitoring the trajectories, predicting the trajectories and focus on object-based attributes to extract the probable behaviours of individuals and as a crowd. However, the factors such as occlusion and the presence of multiple target objects are inhibiting the visibility of anomaly detection algorithms. These factors greatly affect the accuracy of predictions and hence the need for better approaches. Holistic approaches, on the other hand, operate on a wider network of individuals where all are interconnected to provide more meaning to the scenes. The low level and mid-level features are extracted for monitoring the crowd behaviours. Holistic approaches concentrated on optical flow fields, that highlighted the low-level and mid-level features for predicting the outcome with better accuracy. An automated method for deriving the optical flow histograms were proposed for monitoring the overall motion of individuals, thereby connecting them in a wide network of people. In chaotic situations such as the stampedes [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], the proposed histograms were fruitful in identifying the potentially critical situations especially in a crowded environment. Another technique suggested the utilization of low to mid-level features in the motion aspect specifically to visualize the magnitude of the crowd along with the direction of dispersions. The segmentation algorithms proposed in the model were enabled to predict the motion of the crowd during chaotic situations by modelling the regions and probable motion of individuals during normal vs abnormal events. A probabilistic model based on Riemannian detection [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] for detecting crowd anomalies was proposed that handled the optical flow with respect to walking, running and dispersion at different velocities in various directions. Comparatively, the performance of models that have taken the regions of interest from the respective frames has been higher than the models that processed the entire video or frames of the videos [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The technique was to consider the particle advection specifically for quicker processing and deriving the prediction outcomes. Such techniques also considered the trajectory information of every individual with respect to spatial-temporal features and changes observed in different time frames. Motion patterns identified from the different video inputs have shown significant similarities in chaotic situations, thereby establishing a remarkable similarity in crowd anomaly detections. A technique named Histogram of Oriented Tracklets (HOT) was suggested based on the motion described by the histograms that illustrated the motion features, magnitude and the direction of individuals. In case of highly populated areas, the patterns were able to track the anomalies better than the conventional methods.\u003c/p\u003e \u003cp\u003eCrowd Anomaly Detection has gained popularity in the research domain in recent years owing to automated surveillance techniques, and numerous techniques have been introduced. The challenges in the current techniques are also accounted for and the advent of machine learning, deep learning and computer vision algorithms has simplified the entire process flow. Various articles contemplated the list of techniques used for monitoring the individual behaviours and specifically segmenting the abnormal events in crowded environments. The techniques contemplated in the literature review have been collectively describing the various crowded environments, anomalies and target objects that can be the reason for the anomalies [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Such techniques emphasized the need for machine learning, deep learning and computer vision algorithms for automated monitoring systems that can be a live-saving measure for crowded environments. Video analysis techniques played a significant role in anomaly detection models for adding more intelligence to the algorithms. Primarily, the intelligent systems for monitoring the anomalies and crowds used the spatial and temporal features along with perturbations, yet the models needed to address the uneven, non-repeating, unique and rare cases of abnormal events. A better approach implementing a Convolutional Neural Network with multiple optimization techniques was designed and delivered to improve the accuracy of crowd anomaly detection. The need for automated surveillance techniques and thus the accuracy of detection and prediction was justified in the article. Crowded environments were processed in a neural network as continuous streams of input video, right from the cameras and the algorithm detected the anomalies with various threshold features. Ability to process the crowded environments from surveillance videos was tested in a technique that enforced the optimized version of multiple Convolutional Neural Networks. Despite the huge number of individuals in the crowded environments, and a diverse range of anomalies, the optimized Convolutional Neural Network ensemble performed considerably better than the previous techniques said in the literature survey. Another significant benefit of this approach was the considerably low computational cost and the efficiency when compared to the traditional techniques, apart from the difficulties faced by the other techniques in handling real-time surveillance videos. The next technique approached the anomalies detection problem with a two-step process, where the objects and individuals were identified using a You Only Look Once (YOLOv5) version 5 model and a model named DeepSORT was implemented for tracking the trajectories of the individuals [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Optical flow, yet again, was helpful in determining the features that highlight the abnormal behaviours and events of the individuals in the crowded environment. The spatial features were derived from the bounding boxes, compared with the threshold features and the abnormalities were determined. Support Vector Machines were the classifiers that compared the features of the captured events, and compared against the threshold features. From the experimental results, the SVM exhibited an Area Under the Curve metric of 88.9% and the same model has outperformed the other conventional approaches of detecting abnormal events [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e][\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. The same model has proven to perform remarkably well while processing the videos of Hajj yatras, indicating the significant performance during the presence of a dense crowd. Moreover, the model has been able to identify seven distinct abnormal events captured during the yatras, even in the huge crowd. Since the further categorization has been narrowed down to seven more distinctive categories, the accuracy of classifying the abnormal events has improved significantly. Combination of multiple individual techniques have showcased the increase in accuracy and area under the curve of classification of abnormal events. Ensemble of state of art machine learning techniques [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] and multiple modalities of classification algorithms, bounding box techniques and feature extraction techniques have been adapting to fit the current landscape of crowd anomaly detection in the recent years [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The following Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e has detailed all the models and techniques discussed in the literature survey that have produced the models for crowd anomaly detections.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cdiv class=\"gridtable\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSummary of Models and Techniques used for Crowd Anomaly Detection and their performance\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003c/colgroup\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConsidered Datasets\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eModels / Techniques Used\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePerformance Metric\u003c/p\u003e \u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOutcome\u003c/p\u003e \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e \u003cp\u003eShangaiTech Violent Flow\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGenerative Adversarial Networks\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eArea Under Curve\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e73.8%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRecurrent Neural Networks, 2D Convolutional Neural Networks\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e93.5%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOptical Flow\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e73.56%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOptical Flow GAN\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e79.6%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eArea Under Curve\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e98.1%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eUCF Crime\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCNN Residual Logn Short Term Memory\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eArea Under Curve\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e70.4%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCNN, Random Forest\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eArea Under Curve\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e88.98%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHajj Yatra\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOptical Flow\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eArea Under Curve\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e88.96%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUMN\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eArea Under Curve\u003c/p\u003e \u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e88.29%\u003c/p\u003e \u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/table\u003e\u003c/div\u003e \u003cp\u003e\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e\n\n\n\n \n\n\u003cp\u003e\u003c/p\u003e \n\n\u003cp\u003e\u003c/p\u003e \u003cp\u003e\u003c/p\u003e \u003cp\u003e\u003c/p\u003e "},{"header":"3 PROPOSED METHODOLOGY","content":"\u003cp\u003eThe proposed model in the research article contemplates the functioning of crowd anomaly detection methodology with two unique techniques namely, the illustration of the crowded environments, their attributes, definition of various behaviours, followed by the dimensionality reduction and classification of behaviours according to the attributes [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. The surveillance video is processed in a series of frames, from which a set of feature sets are derived to form a list of probably trajectories and Delaunay triangles. From the frames, the visual attributes are further categorized into spatial and temporal features primarily being the velocity, density, motion descriptors and other feature attributes indicating the state of the crowd members. These features are critically acclaimed to the important features that define the quality background upon which the classifiers act upon and determine the abnormal events observed in the given environments [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. The second part of the methodology involves the processing of the feature sets into a set of histograms after aggregation for computational purposes. This section applies the process of dimensionality reduction through the Principal Component Analysis and autoencoders enabling the classification process. Classification is a typical element of neural networks which categorizes and lists the captured events into two primary classes known to be the normal and abnormal events. The following Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e illustrates the functionality of the proposed model and the following sections explain the processes in detail.\u003c/p\u003e\u003ch3\u003e3.1 Crowd Behaviour Analysis\u003c/h3\u003e\u003cp\u003eThe processes commence with the definition of visual descriptors that explain the characteristics of individuals and the entire crowd. Typically for an anomaly detection system, the input videos are processed directly from the surveillance videos captured in real-time. Given the installation of surveillance cameras in major cities, homes and crowded public places, the availability of surveillance videos from places has increased in recent years [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. The frames are thoroughly analysed for regions of interest that may hold the potential of analysis followed by the definition of the feature points. Pixels of specific regions of interest with clear and concise events of abnormality are defined from such regions of interest and thus the other regions are removed from the processing of neural networks. The commonly applied techniques for identifying such events from the frames are FAST, SIFT and AKAZE [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], that are known for the quickness and accuracy of detecting such feature sets from subsequent images. Every frame is analysed through these renowned algorithms to narrow down the features of interest and following them in a series of images thereby ensuring that the abnormal event is prolonged for a specific duration. Such feature sets usually depict the characteristics of objects, individuals and the surroundings detected in a surveillance video. The identical feature sets in subsequent frames describe the valuable information that has to be processed for features tracking and predicting the trajectories of every individual in the crowd. In order to add more sense to the predictions, the objects are considered as well, to detect the abnormal events well in advance. The spatial features are better derived from the Delaunay triangulation strategy, where the spatial elements are highlighted and differentiated between the subsequent frames. Such features are further defined as the visual descriptors which are classified into individual and combined features accordingly [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e].\u003c/p\u003e\u003ch3\u003e3.1.1 Feature Extraction and Tracking\u003c/h3\u003e\u003cp\u003eIn the proposed approach, the Features from Accelerated Segment Test (FAST) technique was applied to identify the pixels of interest from the frames of significant importance and the features were compared with the continuous frames to match for similarities [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. The variations of intensity around the circled region of interest explains the continuous changes in the trajectories and on the other hand, the similar activities in the regions of interest in the neighbouring frames indicate the capture of abnormal events. In order to reduce the computational complexity, the corner criteria technique is applied to identify the predetermined areas of interest on a specific frame. As soon as a specific feature is identified using FAST algorithm, Scale Invariant Feature Transform (SIFT) algorithm is applied to derive a 16x16 neighbourhood matrix around the detected feature thereby deriving a 128-bin value. The equation represents the Difference of Gaussians (DoG) [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], which approximates the Laplacian of Gaussian (LoG) for edge detection in image processing. This difference isolates features that vary in size between the scales, highlighting image details at a specific range of scales. The resulting difference is then convolved (*) with the input image I(x,y) to detect edges or features using the following expression 1, where x, y are the source data to which the Difference of Gaussians filter is applied to detect features. Convolution integrates the filter with the image to detect the desired features at each pixel. The scales 𝑘𝜎 and 𝜎 indicate the difference between the two functions.\u003c/p\u003e\u003cp\u003eD(x,y,σ)=[G(x,y,kσ) − G(x,y,σ)]∗I(x,y) (1)\u003c/p\u003e\u003cp\u003eIn this approach, the feature extraction capabilities of SIFT and FAST along with AKAZE were individually analysed, highlighting their unique contributions to anomaly detection in sparse feature tracking. SIFT is well-known for its ability to remain robust under variations in scale, rotation, and lighting conditions. This method is particularly effective in identifying distinct key points within crowd images, allowing it to capture intricate patterns that signal deviations from typical crowd behaviour. As a result, it is highly suitable for detecting anomalies of different sizes. FAST, on the other hand, is optimized for quick corner detection, enabling the rapid identification of key features in crowd images [\u003cspan additionalcitationids=\"CR27\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e–\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. While it lacks the scale and rotation invariance of SIFT and FAST, its high speed makes it an excellent choice for real-time anomaly detection applications, where quick responses are essential. The efficient detection of key points by FAST enhances its utility in sparse feature tracking for recognizing unusual crowd behaviours. Sparse feature tracking is a widely utilized technique in computer vision for tracking a subset of key features across video frames. Unlike dense tracking methods, which monitor every single pixel, this approach zeroes in on distinctive features within the frames. This focus on unique features such as corners or structural elements enables it to handle challenges like occlusion (where parts of the crowd are hidden) and dynamic changes in crowd behaviours more effectively. Sparse tracking excels by identifying and following these prominent key points [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], ensuring robust performance even when parts of the crowd are obscured or when movement patterns shift unpredictably. One notable advantage of sparse feature tracking is its ability to re-detect and associate features over time. This ensures the system adapts to evolving crowd dynamics, allowing for accurate long-term tracking. By prioritizing features that are simple to identify and track, our approach maintains its effectiveness across various scenarios [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eThe Accelerated-KAZE (AKAZE) algorithm extends the original KAZE algorithm by using a computationally efficient framework known as Fast Explicit Diffusion (FED) to create its non-linear scale spaces. Built upon non-linear diffusion filtering, AKAZE utilizes the determinant of the Hessian matrix for feature detection. To enhance rotation invariance, it employs Scharr filters. The maximum responses from these detectors pinpoint specific feature point locations, forming the basis for AKAZE’s strong and distinctive feature detection capabilities. The AKAZE descriptor [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e] leverages the Modified Local Difference Binary (MLDB) algorithm, renowned for its power and efficiency. Due to the non-linear nature of AKAZE’s scale spaces, the algorithm exhibits invariance to scale, rotation, and limited affine transformations. Furthermore, the distinctiveness of AKAZE’s features is maintained and even enhanced as they are scaled up or down.\u003c/p\u003e\u003cp\u003eOur methodology capitalizes on these local features and incorporates the Lucas–Kanade optical flow algorithm. This algorithm is particularly adept at handling the challenges posed by unrestricted optical flow, making it ideal for applications like crowd anomaly detection. The Lucas–Kanade method stands out due to its precision, robustness, and adaptability to complex environments [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. By emphasizing sparse feature tracking, we substantially reduce computational overhead while maintaining high accuracy, perfect for real-time applications. The Lucas–Kanade algorithm enhances its effectiveness in detecting anomalous behaviours by capturing subtle motion variations within crowded environments. Ensuring temporal coherence in feature tracking stabilizes the system and minimizes false positives, thereby increasing the reliability of anomaly detection mechanisms. This approach operates on the assumption that neighbouring pixels within a small, localized region share consistent optical flow values. Instead of analysing every pixel in the frame, it calculates optical flow based on these groups. This method provides several benefits, such as faster computations and the efficient generation of training data. Mathematically as expressed in Eq.\u0026nbsp;2, the optical flow constraint for a set of pixels moving at the same velocity can be expressed by a specific equation, ensuring a cohesive and efficient analysis of movement within the crowd [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. By leveraging these principles, our methodology strikes a balance between computational efficiency and accuracy, making it a powerful tool for real-time crowd monitoring and anomaly detection.\u003c/p\u003e\u003cp\u003eI\u003csub\u003ex\u003c/sub\u003e(x\u003csub\u003e1\u003c/sub\u003e, y\u003csub\u003e1\u003c/sub\u003e) ・ v\u003csub\u003ex\u003c/sub\u003e + I\u003csub\u003ey\u003c/sub\u003e(x\u003csub\u003e1\u003c/sub\u003e, y\u003csub\u003e1\u003c/sub\u003e) ・ v\u003csub\u003ey\u003c/sub\u003e = − I\u003csub\u003et\u003c/sub\u003e(x\u003csub\u003e1\u003c/sub\u003e, y\u003csub\u003e1\u003c/sub\u003e)\u003c/p\u003e\u003cp\u003eI\u003csub\u003ex\u003c/sub\u003e(x\u003csub\u003e2\u003c/sub\u003e, y\u003csub\u003e2\u003c/sub\u003e) ・ v\u003csub\u003ex\u003c/sub\u003e + I\u003csub\u003ey\u003c/sub\u003e(x\u003csub\u003e2\u003c/sub\u003e, y\u003csub\u003e2\u003c/sub\u003e) ・ v\u003csub\u003ey\u003c/sub\u003e = − It(x\u003csub\u003e2\u003c/sub\u003e, y\u003csub\u003e2\u003c/sub\u003e)\u003c/p\u003e\u003cp\u003e.. .\u003c/p\u003e\u003cp\u003eI\u003csub\u003ex\u003c/sub\u003e(x\u003csub\u003en\u003c/sub\u003e, y\u003csub\u003en\u003c/sub\u003e) ・ v\u003csub\u003ex\u003c/sub\u003e + I\u003csub\u003ey\u003c/sub\u003e(x\u003csub\u003en\u003c/sub\u003e, y\u003csub\u003en\u003c/sub\u003e) ・ v\u003csub\u003ey\u003c/sub\u003e = − I\u003csub\u003et\u003c/sub\u003e(x\u003csub\u003en\u003c/sub\u003e, y\u003csub\u003en\u003c/sub\u003e) (2)\u003c/p\u003e\u003cp\u003eThis efficient and adaptive approach to sparse optical flow makes it a promising solution for crowd anomaly detection in diverse surveillance and monitoring scenarios [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e\u003ch3\u003e3.1.2 Trajectory Detection\u003c/h3\u003e\u003cp\u003eThe Lucas Kanade method may not function in case of object detection especially in motion and the trajectories may not be detected effectively. This raises the requirement to include the gradient transformation for considering the neighbouring pixels of the respective frames to detect the object motions. The proposed approach introduces a novel approach for considering the pixels from regions of interest in form of a pyramidal format. This technique contemplates the techniques of down-sampling, passing through a low-pass filtering technique and applying a factor of 2. The purpose of optical flow inclusion is justified when the low-quality images are processed first [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], followed by the high-quality frames from the same videos, in order to increase the optical flow accuracy. Once the spatial features are computed, the significant features are forwarded to the next stage of processing as tracklets. The series of frames, where the objects are defined to be the primary features of consideration, are transformed into a graph with mapped tracklets. Delaunay Triangulation graph technique [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e] processes the features in omnidirectional nodes in the neighbouring frames in order to trace the local, spatial and temporal features or tracklets. The number of nodes depends on the tracklets and the node connections are represented as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{ϵ}^{n}\\)\u003c/span\u003e\u003c/span\u003e and the relationships between the tracklets are represented by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{g}^{n}\\left({\\vartheta\\:}^{n},{ϵ}^{n},{F}^{n}\\right).\\:\\)\u003c/span\u003e\u003c/span\u003eThe number of triplet features are represented as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{F}^{n}\\)\u003c/span\u003e\u003c/span\u003e. Temporal features are effectively described with respect to the topographical features over the different states of time without affecting the shape of the triangle in a graph eliminating the noise and occlusion factors. A graph containing the tracklets derived from the previous stages, the local tracklets are identified as cliques and neighbouring nodes are identified as seed points, which forms the tracking points for objects and individuals in a crowded environment. The cliques are represented according to the following Eq.\u0026nbsp;\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:C\\left({V}_{i}^{k}\\right)=\\:\\left\\{{V}_{i}^{k}\\right\\}\\:\\cup\\:\\:\\left\\{{V}_{j}^{k},\\:\\forall\\:\\:\\left({V}_{i}^{k},\\:{V}_{j}^{k}\\right)\\:\\in\\:\\:{ϵ}^{n}\\right\\}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003cp\u003eThe connections between the cliques or the tracklets are connected through the short-term or long-term connections representing the probable trajectories. In terms of spatial features, the cliques are connected on varying terms increasing the dynamic ability of the crowd. With a more diversified crowd, the number of cliques with temporal and spatial features are comparatively higher than insignificant cliques.\u003c/p\u003e\u003ch3\u003e3.1.3 Individual vs Crowd Behaviour Visual Descriptors Identification\u003c/h3\u003e\u003cp\u003eOnce the visual descriptors are defined and mapped in a graph, the characteristics of objects and individuals are further analysed for classification. Visual descriptors indicate the semantic information about the crowd participants, typically the spatial and temporal aspects. Information captured from the input videos are highly dynamic and critical for the further analysis [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. The proposed system carefully analyses the individual and collective features, thereby contemplating the entire set of features from the surveillance videos. Such visual descriptors are further classified into individual and entire crowd-based features namely the collective features. Individual behaviours are processed to segment the participants of the crowd, understanding their activities by observing the dynamic properties such as direction of the flow and velocity of individuals. The direction of motions, describes the individuals with respect to normal and extreme conditions in varying tracklets motions [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. A complete structure of the tracklets, neighbouring nodes, are represented by the F segment as described in the following Eq.\u0026nbsp;4.\u003c/p\u003e\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\left\\{{S}_{i}^{n},\\:{S}_{i}^{n-{\\tau\\:}_{2}},\\dots\\:.\\:{S}_{i}^{n-\\left(F-1\\right){\\tau\\:}_{2}}\\right\\}\\:\\left(4\\right)$$\u003c/div\u003e\u003c/div\u003e\u003cp\u003eThe changes in the directions of the individuals are observed by monitoring the angular variations of every trajectory and tracklets. The following Eq.\u0026nbsp;\u003cspan refid=\"Equ2\" class=\"InternalRef\"\u003e5\u003c/span\u003e provides the calculation for predicting the change of direction in a given trajectory.\u003c/p\u003e\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\:{D}^{var\\:}\\left({V}_{i}^{k}\\right)=\\:\\frac{1}{F}\\:.\\:\\sum\\:_{0}^{F-2}{d}_{\\theta\\:}\\left({S}_{i}^{n-f{\\tau\\:}_{2}},\\:{S}_{i}^{n-(f+1){\\tau\\:}_{2}}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\theta\\:\\)\u003c/span\u003e\u003c/span\u003e represents the angular variations, F is represented by vectors of the graph. On the other critical element for determining the individual characteristics, velocity of motion vectors in a specific direction explains the seriousness of the situation occurred in the crowded environment [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. The accuracy of the visual descriptors depends on the considered frame in the history of frames. The motion vectors suitably describe the information about the motion of individuals using the following Eq.\u0026nbsp;(6). Euclidean distance is the measure between the tracklets in any direction, preferably between the current node and origin, instead of measuring the sum of all the tracklets between the nodes.\u003c/p\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:{D}^{velocity}\\left({V}_{i}^{k}\\right)=\\:\\frac{1}{{\\tau\\:}_{1}}\\:.\\:‖\\underset{{V}_{i}^{k-\\tau\\:1}{V}_{i}^{k}}{\\to\\:}‖\\:\\left(6\\right)$$\u003c/div\u003e\u003c/div\u003e\u003cp\u003eIn a crowded environment, it is extremely important to derive the collective characteristics of all the participants and the collective behaviours in this section explains the need for collective visual descriptors. The characteristics of the crowded collective features are derived from the stability, collectiveness, density and uniformity [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. The concept of stability in crowd analysis reflects the degree of consistency in the crowd's topological structure over time. It evaluates how individuals within a crowd maintain their proximity to the same neighbours as time goes on. By examining the stability property, we can glean valuable insights into persistent patterns and relationships within the crowd, thereby enhancing our understanding of its dynamics and behaviours.\u003c/p\u003e\u003cp\u003eThe collectiveness property pertains to how pedestrians move as a cohesive group. This property is measured by calculating each individual's directional deviation from the overall movement of the group. Traditionally, coherent motion has been assessed using predefined collective transitions. However, in this approach, cliques are utilized for the local computation of this descriptor, offering a nuanced alternative. Conflict, an important property in crowd analysis, captures interactions among individuals, especially when they are in close proximity [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Similar to the computation of the collectiveness descriptor, the conflict property is also determined locally.\u003c/p\u003e\u003cp\u003eThe local density descriptor focuses specifically on the spatial distribution aspect of the model. Unlike previous interactive descriptors, it emphasizes the spatial arrangement of individuals within the scene. This descriptor captures a key characteristic of crowd behaviours: the distribution of individuals. An approximate measure of local density can be obtained by evaluating the proximity of nearby features. This is based on the observation that when nearby features converge, it signifies a higher probability of a larger crowd forming in that area. The uniformity descriptor assesses the coherence of the spatial distribution of regional features. It indicates whether a group has a tendency to cluster together in a uniform manner or to fragment into smaller subgroups, reflecting non-uniform behaviours.\u003c/p\u003e\u003ch2\u003e3.1.4 Dimensionality Reduction and Classification\u003c/h2\u003e\u003cp\u003eThe previous section explained the list of descriptors that worked upon the features with spatial and temporal features that were predominant in identifying the abnormal events in the crowded environments. Dimensionality reduction [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e] is a renowned technique for processing the potential list of features by reducing the rich feature sets and processing only the required critical information. The proposed approach reduced the dimensionalities using two remarkable techniques namely principal component analysis and autoencoding. These two approaches are common in computer vision applications and have been proven to reduce the computational cost associated with processing images and videos. The first technique employed was principal component analysis (PCA). PCA aims to transform the original features into a new set of uncorrelated variables, known as principal components, while preserving as much variance as possible. By projecting the data onto these components, we effectively reduce the information into a lower-dimensional space.\u003c/p\u003e\u003cp\u003eThe detection of anomalies within crowd dynamics is a critical challenge in visual crowd analysis. This paper presents a comprehensive approach that leverages dimensionality reduction techniques and neural network architectures to effectively identify unusual crowd behaviours. The process begins with the application of principal component analysis to reduce the dimensionality of the input data. By capturing the most significant features, PCA helps to prepare the data for the subsequent classification stage. The second approach involves using autoencoders, a type of neural network designed to learn efficient representations of the input data. Autoencoders consist of an encoder network that compresses the input into a latent space representation and a decoder network that reconstructs the original input from this representation. By training the autoencoder to minimize the reconstruction error, it learns to capture the most important features of the data in the latent space. After applying dimensionality reduction using PCA and autoencoders, we proceed to the classification stage. In this step, we utilize the reduced-dimensional feature set to train neural network classifiers. Specifically, we employ neural networks like the multi-layer perceptron. Neural networks are well-suited to handle non-linear problems and extract intricate patterns from the input data.\u003c/p\u003e\u003cp\u003eThe adaptability of neural networks is especially useful in detecting anomalies in crowd behaviours. They excel at uncovering subtle relationships within the data and identifying unusual crowd dynamics that might be overlooked by traditional methods. Due to their hierarchical architectures, neural networks can capture both detailed and higher-level representations, providing a comprehensive understanding of crowd behaviours. Throughout our study, we evaluated neural networks with varying numbers of hidden layers to determine the optimal configuration for our analysis. Through systematic testing and performance comparisons, we identified the most effective approach for crowd anomaly detection (Favarelli \u0026amp; Giorgetti, 2020). This rigorous approach ensured that our methodology was robust and capable of addressing the complexities of detecting anomalies within crowd dynamics.\u003c/p\u003e"},{"header":"4 RESULTS AND DISCUSSIONS","content":"\u003cp\u003eThe proposed approach exemplifies the performance of crowd anomaly detection through a series of investigations with state of art datasets known to have captured crowd activities. The datasets possess various crowd activities and is a standard dataset to be used in this specific research domain. University of Minnesota (UMN) along with other datasets that have captured violent and unsettling events have been used for measuring the performance of the proposed approach. Various scenes in the 11 videos of the datasets depict the normal and abnormal (violent) behaviours in subsequent frames, thereby having scenes of people running in a single direction, different directions, moving from the same point of origin with high velocity, and other atypical scenes of abnormal nature. YouTube is also the repository that contains numerous videos of unsettling crowded environments with different tension among the participants, and other violent scenes. Surveillance videos are curated into a list of videos containing a balanced number of normal videos and violent videos. From the UMN datasets, the proposed approach has delivered the visual descriptors for sensing the different behaviours and the YouTube videos were used to measure the accuracy of predicting the abnormal events in challenging situations. The following Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e illustrates the spatial features captured by the Delaunay triangulation technique which explains the distribution of crowd using the following metrics. Each individual in the crowd is marked as a distinct point. Triangles are formed such that no point lies inside the circumcircle of any triangle. The edges of these triangles represent the proximity of individuals to one another, capturing the spatial relationships within the crowd.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe spatial features between individuals and cluster of people will be tracked using the visual descriptors identified by the feature extraction techniques of the proposed model. As soon as the features are marked and identified to be potentially significant features, the Delaunay Triangulation technique will mark the presence of individuals as distinct points, a group of such distinct points marks a crowd with normal activities. Upon the presence of abnormal activities, the events are marked as red, with varying velocity, the Delaunay triangulation will mark the edges in green coloured lines representing different clusters of people/participants. The cohesion among the group participants will be lesser in case of an abnormal event, thereby dispersing the crowd in different directions. The clusters will be less crowded owing to the events and the Delaunay triangulation clearly differentiates the normal vs abnormal crowds as shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e(a) and 2(b) respectively. When the crowd is found to be agitated due to an abnormal event, the distance between the nodes/distinctive points is found to be longer than the nodes with normal events. The length between the nodes will be consistent in a normal crowd and uniformity is a critical element. Length of the green lines will be longer and scattered when the participants are dispersed or scattered in different directions when a chaotic environment is found in the crowd. The following Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e has listed the performance measures in terms of accuracy and area under the curve (AUC) metrics based on the quality of feature extraction techniques applied in the proposed model. The configuration of the neural networks is tested for three different scenarios where a normal neural network is applied without any dimensionality reduction techniques, and an autoencoding technique is applied in the next configuration where the dimensions are reduced to 128 bits. The last variant is the performance of the proposed method where the contrast features are reduced based on Principal Component Analysis (PCA) where the variance value is considered to be 95. For testing purposes, a 5-fold cross-validation technique is applied for classifying the behaviours.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of different techniques for detecting chaotic scenes in input videos\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eTechniques\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eNeural Network\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003e128-Bits Dimensionality Reduction NN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c7\" namest=\"c6\"\u003e \u003cp\u003ePCA \u0026ndash; NN\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSIFT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.803\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.805\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.799\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.802\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.848\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.846\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFAST\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.848\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.853\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.844\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.847\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.844\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.851\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eA-KAZE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.858\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.862\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.865\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.890\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.923\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.898\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIt is observed that SIFT feature extraction technique has performed better than the other two techniques in terms of crowded environments and in terms of dimensionality reduction, the PCA based neural network has shown the highest accuracy and area under the curve. The spatial information and crowd features were best described with a SIFT based neural network, configured with principal component analysis-based dimensionality reduction. From the classification results tabulated in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, the videos taken from the YouTube sources were processed better than the inputs from UMN datasets. Given that the challenges in the YouTube videos were present, the proposed model has shown promising results as highlighted in the given table.\u003c/p\u003e \u003cp\u003eThe following Figs.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e illustrates the results captured across the three methods and showcases the classification accuracy of the three techniques. Descriptors obtained through the processing of SIFT, FAST and AKAZE upon the violent videos and UMN dataset, crowd anomaly detection was more balanced. From the inputs of UMN dataset, the videos were more structured and the scenarios were limited, posing a clear-cut approach for training the models. The events in the UMN dataset were more controlled and hence the possessed lesser challenging environments. On the other hand, the videos from the YouTube repository were highly dynamic and more challenging. The quality of the videos was comparatively lesser on the YouTube, and hence added more constraints and longer processing time for understanding the scenes better. Conversely, the reduced performance on the YouTube dataset could be due to issues such as poor video quality, occlusion, and varied surveillance scenarios. Addressing these specific challenges within each dataset is crucial to enhance the effectiveness of crowd anomaly detection methods.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe proposed model achieved an AUC of 89.67% and 97.50% for scenes with normal activities and abnormal activities, respectively when the input from UMN dataset was processed. The input videos from the UMN dataset were processed, where the accuracy and AUC performance parameters are tabulated in the below Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. From the comparative analysis, our approach has shown superior performance relative to other methods as listed below. Specifically, our approach achieved 99.5%, 96.5%, and 99% for normal and abnormal activities of the UMN dataset, respectively. Moreover, our approach shows a more pronounced advantage on the YouTube videos, where the proposed approach has delivered 89.67% AUC and 88.5% accuracy in detecting the abnormal activities. From the investigative results, consistent performance of our approach across both datasets indicates its high effectiveness in classification tasks, leveraging unique features and techniques that contribute to its superior performance. The following Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e lists down the comparative results of classification accuracy with other state of art techniques and models.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of different Models for classification\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModels\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOptical Flow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e84%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e87.3%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSparse Reconstruction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e90.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e88.1%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVisual Descriptors\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e88%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e89.1%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGAN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e87.65%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e87.2%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProposed Approach\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e89.67%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e88.5%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e"},{"header":"5 CONCLUSION","content":"\u003cp\u003eThis study introduces an innovative model by leveraging the combination of visual descriptors, feature extraction and neural networks for detecting the abnormal events occurring in a crowded environment. The proposed method's effectiveness, demonstrated through the investigative results on the renowned UMN and crowd activities datasets on YouTube, highlights its potential in accurately detecting abnormal crowd behaviours. SIFT and AKAZE descriptors along with the neural networks, along with the impressive performance of the neural network configurations, underscores the robustness of our approach. Consequently, the proposed work makes a significant contribution to the evolving field, laying a solid foundation for future advancements in detecting the abnormal events in a crowded environment. The enhancements to the conventional neural networks with better feature descriptors and architectural updates to the deep learning systems improve the accuracy of detecting abnormal events. The patterns of abnormal behaviours are processed collectively to provide meaningful insights to the models for future detection. The proposed model has been tested against the real time videos curated from YouTube and UMN datasets for applicability in real time environments. The model has been tested recursively with different scenarios for addressing the various complex situations and other factors limiting the processing ability of the proposed model. In order to enhance the real time implications, the proposed model has to incorporate multimodal inputs from various sources to be more adaptable and promising.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eWe hereby declare that this manuscript is the result of our collaborative effort. Manojkumar K contributed to the conceptualization, methodology, and writing of the manuscript. Suji Helen provided significant contributions to data analysis, interpretation, and critical revisions of the manuscript. Both authors have read and approved the final version of the manuscript and agree to its submission to the Iranian Journal of Science and Technology.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eManoj Kumar K, Sujihelen L (2022) Recognising Actions with Segmentation and Prediction Techniques in ROI based Deep Learning Framework. Math Stat Eng Appl 71(4):4072\u0026ndash;4090. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.17762/msea.v71i4.977\u003c/span\u003e\u003cspan address=\"10.17762/msea.v71i4.977\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSujihelen L (2021) Behavioural Analysis For Prospects In Crowd Emotion Sensing: A Survey, \u003cem\u003eThird International Conference on Inventive Research in Computing Applications (ICIRCA)\u003c/em\u003e, Coimbatore, India, 2021, pp. 735\u0026ndash;743. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ICIRCA51532.2021.9544607\u003c/span\u003e\u003cspan address=\"10.1109/ICIRCA51532.2021.9544607\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHalboob W, Altaheri H, Derhab A, Almuhtadi J (2024) Crowd Management Intelligence Framework: Umrah Use Case, in \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 12, pp. 6752\u0026ndash;6767. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2024.3350188\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2024.3350188\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu D, Liu W, Yuan X, Jiang Y (2023) Conscious and Unconscious Processing of Ensemble Statistics Oppositely Modulate Perceptual Decision-Making. Am Psychol 78:346\u0026ndash;357\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eList of Human Stampedes and Crushes (2022) [online] Available: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://en.wikipedia.org/wiki/List_of_human_stampedes_and_crushes\u003c/span\u003e\u003cspan address=\"https://en.wikipedia.org/wiki/List_of_human_stampedes_and_crushes\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAljuaid H, Akhter I, Alsufyani N, Shorfuzzaman M, Alarfaj M, Alnowaiser K, Jalal A, Park J (2023) Postures anomaly tracking and prediction learning model over crowd data analytics. PeerJ Comput Sci. 9, e1355\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Y, Luo X, Zhou Z (Aug. 2024) Contrasting Estimation of Pattern Prototypes for Anomaly Detection in Urban Crowd Flow. IEEE Trans Intell Transp Syst 25(8):10231\u0026ndash;10245. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TITS.2024.3355143\u003c/span\u003e\u003cspan address=\"10.1109/TITS.2024.3355143\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAl-Shaery AM et al (2024) Open Dataset for Predicting Pilgrim Activities for Crowd Management During Hajj Using Wearable Sensors, in \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 12, pp. 72828\u0026ndash;72846. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2024.3402230\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2024.3402230\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo L, Li Y, Yin H, Xie S, Hu R, Cai W (2023) Crowd-level abnormal behavior detection via multi-scale motion consistency learning, \u003cem\u003eProc. AAAI Conf. Artif. Intell.\u003c/em\u003e, pp. 8984\u0026ndash;8992\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo L, Xie S, Yin H, Peng C, Ong Y-S (2024) Detecting and Quantifying Crowd-Level Abnormal Behaviors in Crowd Events. IEEE Trans Inf Forensics Secur 19:6810\u0026ndash;6823. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TIFS.2024.3423388\u003c/span\u003e\u003cspan address=\"10.1109/TIFS.2024.3423388\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao X-C, Chen W-N, Guo X-Q, Zhong J, Hu X-M (2023) Crowd management through optimal layout of fences: An ant colony approach based on crowd simulation, \u003cem\u003eIEEE Trans. Intell. Transp. Syst.\u003c/em\u003e, vol. 24, no. 9, pp. 9137\u0026ndash;9149, Sep\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbdullah F, Abdelhaq M, Alsaqour R, Alatiyyah MH, Alnowaiser K, Alotaibi SS et al (2023) ,Context aware crowd tracking and anomaly detection via deep learning and social force model. IEEE Access 11:75884\u0026ndash;75898\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDing X, He F, Lin Z, Wang Y, Guo H, Huang Y (2021) Crowd density estimation using fusion of multi-layer features, \u003cem\u003eIEEE Trans. Intell. Transp. Syst.\u003c/em\u003e, vol. 22, no. 8, pp. 4776\u0026ndash;4787, Aug\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNoor TH (2023) Behavior analysis-based IoT services for crowd management, \u003cem\u003eComput. J.\u003c/em\u003e, vol. 66, no. 9, pp. 2208\u0026ndash;2219, Sep\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlsubai S et al (2024) Design of Artificial Intelligence Driven Crowd Density Analysis for Sustainable Smart Cities, in \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 12, pp. 121983\u0026ndash;121993. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2024.3390049\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2024.3390049\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo L, Li Y, Yin H, Xie S, Hu R, Cai W (2023) Crowd-level abnormal behavior detection via multi-scale motion consistency learning, \u003cem\u003eProc. AAAI Conf. Artif. Intell.\u003c/em\u003e, pp. 8984\u0026ndash;8992\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao R et al (2022) Oct., Dynamic crowd accident-risk assessment based on internal energy and information entropy for large-scale crowd flow considering COVID-19 epidemic, \u003cem\u003eIEEE Trans. Intell. Transp. Syst.\u003c/em\u003e, vol. 23, no. 10, pp. 17466\u0026ndash;17478\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang T, Wang C, Zhou T, Cai Z, Wu K, Hou B (2022) Identification of anomalous behavioral patterns in crowd scenes. Comput Mater Continua 71(1):925\u0026ndash;939\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYadav S, Gulia P, Gill NS, Chatterjee JM (2022) A real-time crowd monitoring and management system for social distance classification and healthcare using deep learning, \u003cem\u003eJ. Healthcare Eng.\u003c/em\u003e, vol. pp. 1\u0026ndash;11, Apr. 2022\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLong J, Liang W, Li K-C, Wei Y, Marino MD (Feb. 2023) A regularized cross-layer ladder network for intrusion detection in industrial Internet of Things. IEEE Trans Ind Informat 19(2):1747\u0026ndash;1755\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Xie Z, Li B, Mohiuddin M (2022) The impacts of in situ urbanization on housing mobility and employment of local residents in China, \u003cem\u003eSustainability\u003c/em\u003e, vol. 14, no. 15, p. 9058, Jul\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu G, Wang S, Cai Z, Liu X, Zhu E, Yin J (2023) Video anomaly detection via visual cloze tests. IEEE Trans Inf Forensics Secur 18:4955\u0026ndash;4969\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang M, Li T, Yu Y, Li Y, Hui P, Zheng Y (2022) Urban anomaly analytics: Description detection and prediction, \u003cem\u003eIEEE Trans. Big Data\u003c/em\u003e, vol. 8, no. 3, pp. 809\u0026ndash;826, Jun\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLalit R, Purwar RK (2022) Crowd abnormality detection using optical flow and GLCM-based texture features, \u003cem\u003eJ. Inf. Technol. Res.\u003c/em\u003e, vol. 15, no. 1, pp. 1\u0026ndash;15, Jun\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen TN, Zeadally S (May 2022) Mobile crowd-sensing applications: Data redundancies challenges and solutions. ACM Trans Internet Technol 22(2):1\u0026ndash;15\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang S, Yang Y, Liang W, Sandor VKA, Xie G, Choo KR (2023) MKSS: An effective multi-authority keyword search scheme for edge-cloud collaboration, \u003cem\u003eJ. Syst. Archit.\u003c/em\u003e, vol. 144\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMart\u0026iacute; P, Serrano-Estrada L, Nolasco-Cirugeda A, Baeza JL (2022) Revisiting the spatial definition of neighborhood boundaries: Functional clusters versus administrative neighborhoods, \u003cem\u003eJ. Urban Technol.\u003c/em\u003e, vol. 29, no. 3, pp. 73\u0026ndash;94, Jul\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen C et al (2022) Jun., Comprehensive regularization in a bi-directional predictive network for video anomaly detection, \u003cem\u003eProc. AAAI Conf. Artif. Intell.\u003c/em\u003e, vol. 36, no. 1, pp. 230\u0026ndash;238\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeng L, Lian D, Huang Z, Chen E (2022) Graph convolutional adversarial networks for spatiotemporal anomaly detection, \u003cem\u003eIEEE Trans. Neural Netw. Learn. Syst.\u003c/em\u003e, vol. 33, no. 6, pp. 2416\u0026ndash;2428, Jun\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang T, Wang C, Zhou T, Cai Z, Wu K, Hou B (2022) Identification of anomalous behavioral patterns in crowd scenes. Comput Mater Continua 71(1):925\u0026ndash;939\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu Y, Liu X, Li X, Li M, Li Y (2023) Participants recruitment for coverage maximization by mobility predicting in mobile crowd sensing, \u003cem\u003eChina Commun.\u003c/em\u003e, vol. 20, no. 8, pp. 163\u0026ndash;176, Aug\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen Y, Deng A (2022) IEEE Access 10:93513\u0026ndash;93524Using POI data and Baidu migration big data to modify nighttime light data to identify urban and rural area\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCao C, Lu Y, Zhang Y (2024) Context recovery and knowledge retrieval: A novel two-stream framework for video anomaly detection. IEEE Trans Image Process 33:1810\u0026ndash;1825\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu CH et al (2021) Apr., Modeling citywide crowd flows using attentive convolutional LSTM, \u003cem\u003eProc. IEEE 37th Int. Conf. Data Eng. (ICDE)\u003c/em\u003e, pp. 217\u0026ndash;228\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang J, Zhang X (2023) Multi-task allocation in mobile crowd sensing with mobility prediction, \u003cem\u003eIEEE Trans. Mobile Comput.\u003c/em\u003e, vol. 22, no. 2, pp. 1081\u0026ndash;1094, Feb\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang H, Zeng S, Li Y, Jin D (2021) Predictability and prediction of human mobility based on application-collected location data, \u003cem\u003eIEEE Trans. Mobile Comput.\u003c/em\u003e, vol. 20, no. 7, pp. 2457\u0026ndash;2472, Jul\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu T, Zhang C, Lam K-M, Kong J (2023) Decouple and resolve: Transformer-based models for online anomaly detection from weakly labeled videos. IEEE Trans Inf Forensics Secur 18:15\u0026ndash;28\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWoo G, Liu C, Sahoo D, Kumar A, Hoi S (2022) CoST: Contrastive learning of disentangled seasonal-trend representations for time series forecasting, \u003cem\u003eProc. Int. Conf. Learn. Represent.\u003c/em\u003e, pp. 1\u0026ndash;18\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu X, Liu J (2022) A truthful double auction mechanism for multi-resource allocation in crowd sensing systems, \u003cem\u003eIEEE Trans. Serv. Comput.\u003c/em\u003e, vol. 15, no. 5, pp. 2579\u0026ndash;2590, Sep./Oct\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu P et al (2024) VadCLIP: Adapting vision-language models for weakly supervised video anomaly detection, \u003cem\u003eProc. AAAI Conf. Artif. Intell.\u003c/em\u003e, pp. 6074\u0026ndash;6082\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang Y, Zhang C, Zhou T, Wen Q, Sun L (2023) DCdetector: Dual attention contrastive representation learning for time series anomaly detection, \u003cem\u003eProc. 29th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining (KDD)\u003c/em\u003e, pp. 3033\u0026ndash;3045\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang S, He J, Liang W, Li K (2024) A secure and verifiable multimedia data search scheme for cloud-assisted edge computing. Future Gener Comput Syst 151:32\u0026ndash;44\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"crowd emotion sensing, anomaly detection, spatial, temporal, behaviour, cross division","lastPublishedDoi":"10.21203/rs.3.rs-5709790/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5709790/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe crowd emotion sensing is a critical element in surveillance and management of the crowd in different environments. With exploding populations, and developing nations, the crowd in urban cities mandate state of art surveillance methodologies involving continuous monitoring and reporting of criminal activities. The research article presents a novel technique to compute the spatial and temporal features obtained from the crowd environments and combine the novelty of neural networks for detecting the emotions of crowds with better accuracy and swiftness. The features are obtained from the continuous feed of surveillance videos typically categorized into the common features of human beings namely anger, sadness, disgust, surprise, fear, happiness and obviously neutrality. Such features are extracted after careful background separation which are typically difficult in crowded environments, using techniques namely SIFT, and FAST termed to be the visual descriptors. Once the features are extracted, spatial and temporal features are classified into individual and combined features as defined in the cross-division environment in order to portray the crowd dynamics and characteristics. Cross division environment computes the necessary features for identifying the anomalies in the crowded situations in a neural network, after a series of operations such as dimensionality reduction, and principal component analysis. From the semantic information, crowd behaviours are detected based on interactive features in a dynamic environment and the proposed technique has demonstrated effective results in terms of 98.9% accuracy in detecting especially violence in crowd datasets collected from UMN.\u003c/p\u003e","manuscriptTitle":"Categorizing Crowd Emotions based on Cross Division Expressions and Anomalies","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-12-26 08:54:50","doi":"10.21203/rs.3.rs-5709790/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"02cf03fd-acf5-4640-93df-69682cb211ba","owner":[],"postedDate":"December 26th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-12-29T16:38:26+00:00","versionOfRecord":[],"versionCreatedAt":"2024-12-26 08:54:50","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5709790","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5709790","identity":"rs-5709790","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00