<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">50435</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.050435</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Abnormal Action Recognition with Lightweight Pose Estimation Network in Electric Power Training Scene</article-title>
<alt-title alt-title-type="left-running-head">Abnormal Action Recognition with Lightweight Pose Estimation Network in Electric Power Training Scene</alt-title>
<alt-title alt-title-type="right-running-head">Abnormal Action Recognition with Lightweight Pose Estimation Network in Electric Power Training Scene</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Cai</surname><given-names>Yunfeng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Qin</surname><given-names>Ran</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Tang</surname><given-names>Jin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Long</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Bi</surname><given-names>Xiaotian</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yang</surname><given-names>Qing</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>yangq@njit.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>State Grid Jiangsu Electric Power Co., Ltd. Research Institute</institution>, <addr-line>Nanjing, 211103</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Computer Engineering, Nanjing Institute of Technology</institution>, <addr-line>Nanjing, 211167</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Qing Yang. Email: <email>yangq@njit.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>20</day>
<month>6</month>
<year>2024</year></pub-date>
<volume>79</volume>
<issue>3</issue>
<fpage>4979</fpage>
<lpage>4994</lpage>
<history>
<date date-type="received">
<day>06</day>
<month>2</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>5</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Cai et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Cai et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_50435.pdf"></self-uri>
<abstract>
<p>Electric power training is essential for ensuring the safety and reliability of the system. In this study, we introduce a novel Abnormal Action Recognition (AAR) system that utilizes a Lightweight Pose Estimation Network (LPEN) to efficiently and effectively detect abnormal fall-down and trespass incidents in electric power training scenarios. The LPEN network, comprising three stages&#x2014;MobileNet, Initial Stage, and Refinement Stage&#x2014;is employed to swiftly extract image features, detect human key points, and refine them for accurate analysis. Subsequently, a Pose-aware Action Analysis Module (PAAM) captures the positional coordinates of human skeletal points in each frame. Finally, an Abnormal Action Inference Module (AAIM) evaluates whether abnormal fall-down or unauthorized trespass behavior is occurring. For fall-down recognition, three criteria&#x2014;falling speed, main angles of skeletal points, and the person&#x2019;s bounding box&#x2014;are considered. To identify unauthorized trespass, emphasis is placed on the position of the ankles. Extensive experiments validate the effectiveness and efficiency of the proposed system in ensuring the safety and reliability of electric power training.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Abnormal action recognition</kwd>
<kwd>action recognition</kwd>
<kwd>lightweight pose estimation</kwd>
<kwd>electric power training</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Natural Science Foundation of Jiangsu Province</funding-source>
<award-id>BK20230696</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Computer vision is a specialized field within Artificial Intelligence (AI) that focuses on enabling computers to interpret and extract information from images and videos. It has been applied to various fields including healthcare, transportation, industrial manufacturing, and electric power systems. The electric power system is one of the fundamental infrastructures that support the functioning of modern society. It provides a stable supply of electricity to households, businesses, medical institutions, schools, and other institutions, supporting various aspects of life and industrial activities. With the growing demand for electricity in society, higher requirements have been placed on the safety and reliability of the utilization of electric power energy. The application of computer vision technology to electric power systems could enhance safety, efficiency, and reliability in various aspects. By using computer vision technology to monitor power equipment, such as transformers, switchgear, and cables, potential faults or damages can be detected promptly [<xref ref-type="bibr" rid="ref-1">1</xref>]. Utilizing unmanned aerial vehicles equipped with computer vision systems, conducting regular inspections of power lines enables the rapid detection and pinpointing of potential issues, such as line breaks or foreign object attachments [<xref ref-type="bibr" rid="ref-2">2</xref>]. In the data centers of the power system, computer vision technology can be exploited to monitor the operational status, temperature, and energy efficiency of equipment, aiding in enhancing operational efficiency [<xref ref-type="bibr" rid="ref-3">3</xref>]. It is also implemented to monitor the safety conditions of substations including foreign object trespass detection, fire detection, as well as the real-time surveillance of personnel while manipulating electrical power equipment.</p>
<p>Electric power training is essential to ensure the safety and reliable operation of the system, which enables employees to acquire the operational and maintenance skills necessary for electrical power equipment, playing a pivotal role in ensuring the regular functioning of equipment and timely maintenance. Electric power training typically includes theoretical knowledge sessions and practical hands-on training with electrical equipment, often in high-voltage and complex work environments. Therefore, operational safety is crucial in the electric power training scenario, particularly during the practical hands-on training part. Virtual Reality (VR) and Augmented Reality (AR) technologies have been employed in some practical hands-on training processes, where the trainer can undergo operational training for power equipment in a virtual environment [<xref ref-type="bibr" rid="ref-4">4</xref>], simulating real-world scenarios to enhance their proficiency and safety in actual work settings. However, some of the practical hands-on training still needs to be conducted in real-world work environments. Computer vision technology can be employed to monitor the actions of trainers in real time [<xref ref-type="bibr" rid="ref-5">5</xref>], providing immediate feedback to maintain order and safety during electric power training Abnormal fall-down is one of the most serious actions in electrical power training scenarios, particularly when dealing with complex or high-voltage electrical equipment. For safety training purposes, there are restricted zones where trainers are not permitted to enter, especially when power equipment is in operation. However, due to a shortage of teachers compared to trainers, there may be still abnormal trespass during the training process. Thus, it is of great significance to accurately and effectively detect abnormal falls-down and trespass using computer vision technology to ensure safety in electrical power training.</p>
<p>To detect abnormal actions of fall-down and trespass in electric power training scenarios, we provide a novel abnormal action recognition system that utilizes the Lightweight Pose Estimation Network (LPEN). In specific, video frames are evenly extracted and input into the proposed LPEN network, which includes MobileNet, Initial Stage, and Refinement Stage for fast extraction image features, rapid detection of human key points, and further refinement of human key points, respectively. Then a Pose-aware Action Analysis Module (PAAM) is used to obtain the positional coordinates of human skeletal points in the frame. An Abnormal Action Inference Module (AAIM) is next employed to assess whether there is a current occurrence of abnormal fall-down or unauthorized trespass behavior. The assessment of abnormal fall-down considers three criteria: Falling speed, main angles of skeletal points, and the bounding box of humans. For unauthorized trespass, the major focus is on the position of the ankles. Finally, we experimentally validate the performance of the proposed Abnormal Action Recognition System (AARS) with respect to real captured video during the scene of electrical power training.</p>
<p>In summary, the main contributions of this article are as follows:
<list list-type="bullet">
<list-item>
<p>We propose an Abnormal Action Recognition (AAR) system equipped with a lightweight pose estimation network to recognize the fall-down action and trespass action of people in electric power training scenarios.</p></list-item>
<list-item>
<p>We introduce a new lightweight pose estimation network, named LPEN, to realize effective and efficient performance in terms of pose estimation.</p></list-item>
<list-item>
<p>We design an Abnormal Action Inference Module (AAIM) to automatically detect the state of fall-down, as well as to automatically recognize the trespass action.</p></list-item>
</list></p>
<p>The rest of the paper is organized as follows. In <xref ref-type="sec" rid="s2">Section 2</xref>, we briefly introduce the previous works of 2D pose estimation, along with a short review of OpenPose. 
Then, the overview of the proposed approach, including technical details, is described in <xref ref-type="sec" rid="s3">Section 3</xref>. The experimental evaluations and analysis are presented in <xref ref-type="sec" rid="s4">Section 4</xref>, and finally, we conclude the paper in <xref ref-type="sec" rid="s5">Section 5</xref>.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<p>In this section, we review related works on pose estimation and abnormal action recognition.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Pose Estimation</title>
<p>Human pose estimation aims to accurately infer the positions of key joints of the human body from images or videos, thereby building a representation of the human body [<xref ref-type="bibr" rid="ref-6">6</xref>]. It involves detecting and locating the positions of key joints such as the head, shoulders, elbows, wrists, knees, and ankles. The spatial relationships of these key joints are critical for accurate pose estimation. The input data for human body pose estimation can be either two-dimensional or three-dimensional. Consequently, pose estimation methods are divided into two-dimensional and three-dimensional categories [<xref ref-type="bibr" rid="ref-7">7</xref>]. Since this paper focuses on electrical power training scenarios, the images captured by the sensor camera are in a two-dimensional format. Therefore, this paper primarily discusses two-dimensional human body pose estimation methods and their subsequent application in human body action recognition.</p>
<p>During the early stages, the predominant approaches of human pose estimation are normally the combination of manually designed feature images [<xref ref-type="bibr" rid="ref-8">8</xref>] with graph-structured models [<xref ref-type="bibr" rid="ref-9">9</xref>]. Grounding on this traditional framework, researchers have continuously worked to enhance the accuracy of feature descriptions and the efficiency of searching for body parts. However, due to the high flexibility of human poses, challenges arise in performance when faced with real-world scenarios. Simultaneously, there is a growing demand for more detailed feature descriptions of body parts to improve the accuracy of human pose estimation. With the development of deep learning technologies, especially the tremendous success of Convolutional Neural Networks (CNN) [<xref ref-type="bibr" rid="ref-10">10</xref>], a new dawn has emerged for addressing the aforementioned issues. Pose estimation methods based on deep learning [<xref ref-type="bibr" rid="ref-11">11</xref>] involve constructing neural network models to fit large amounts of training data, implicitly learning the mapping relationship from input images to the coordinates of human key points. In comparison to traditional methods that rely on manually designed representations, deep convolutional neural networks can automatically learn feature representations from data, thus avoiding the drawbacks associated with manual features.</p>
<p>Pose estimation in videos can be divided into single-person and multi-person types, depending on the number of individuals present. The primary objective of single-person pose estimation is to identify the key points of a specific individual appearing in a captured image. In scenarios where multiple individuals are present in an image, it becomes necessary to segment the image into patches, ensuring that each patch contains only one person. This segmentation can be achieved using either an upper-body detector [<xref ref-type="bibr" rid="ref-12">12</xref>] or a full-body detector [<xref ref-type="bibr" rid="ref-13">13</xref>]. Single-person pose estimation methods can be categorized into two types based on the training data: Key point regression-based approaches and heatmap-based approaches [<xref ref-type="bibr" rid="ref-14">14</xref>]. Keypoint regression-based methods, also known as direct regression, directly capture the locations of key points from feature maps learned by an end-to-end framework. Toshev et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] proposed a cascaded Deep Neural Network (DNN) regressor known as DeepPose for pose estimation. This model uses a 7-layered generic convolutional DNN to regress the location of each body joint based on the full image input. DeepPose has shown better performance compared to traditional methods, paving the way for further advancements in learning-based pose estimation techniques. However, such methods, where the numerical location regression is directly derived from the end-to-end regressor, tend to result in a loss of spatial information of key points, which limits the model&#x2019;s spatial generalization ability. Subsequently, heatmap-based approaches are proposed to address this limitation. Tompson et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] introduced a hybrid architecture for heatmap-based pose estimation, which combines a Deep Convolutional Network with a Markov Random Field. This method leverages the structural relationships between human key points and optimizes the prediction results using a Markov Random Field. As a result, it became one of the most advanced techniques for human pose estimation at the time. By transforming the problem of human pose estimation from coordinate regression to a detection problem based on heat map regression, this method maximizes the preservation of spatial information of key point coordinates. Consequently, it significantly enhances the spatial generalization capability of the learned pose estimation model and improves its accuracy. Wei et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed Convolutional Pose Machines (CPMs) for the task of articulated pose estimation by designing a sequential architecture composed of convolutional networks that directly operate on belief maps from previous stages. CPMs introduce intermediate supervision periodically through the network, thereby replenishing back-propagated gradients and conditioning the learning procedure. Newell et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] proposed a Stacked Hourglass Network (SHN) that captures and consolidates information across different scales of an image. Based on this model, researchers have further developed methods for pose estimation, with a focus on improving accuracy by considering the structure of the human body. Chu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] proposed an end-to-end framework for human pose estimation that incorporates convolutional neural networks with a multi-context attention mechanism. SHNs are also adopted to generate attention maps from features at multiple resolutions with various semantics. Luo et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] developed a technique called Scale-Adaptive Heatmap Regression (SAHR) to modify the standard deviation for each keypoint. This approach is more tolerant of various human scales and labeling ambiguities, but it aggravates the imbalance between fore-background samples. Huang et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed a principled Unbiased Data Processing (UDP) strategy. This method is equipped with unit length-based measurement and employs a combination of classification and regression for encoding-decoding processes. Additionally, there has been a focus on developing lightweight models, which are crucial for the practical implementation of single-person pose estimation applications [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
<p>Multi-person pose estimation, which deals with scenes containing several individuals, presents a more complex challenge than single-person estimation, compromising the top-down approach and bottom-up approaches [<xref ref-type="bibr" rid="ref-23">23</xref>]. The top-down approach transforms the original multi-person human pose estimation task into several single-person estimations. It first identifies each human body in the scene and then estimates the coordinates of key points for each identified individual. Chen et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] presented a Cascaded Pyramid Network (CPN) for multi-person pose estimation, which includes two main components: GlobalNet and RefineNet. GlobalNet is responsible for localizing fundamental key points such as eyes and hands, while RefineNet handles key points that cannot be accurately estimated by integrating both global and local features. Xiao et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed a simple yet effective baseline scheme, which employs Faster R-CNN as the body boundary detector and ResNet as the backbone. Rodrigues et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] presented a pose estimation method utilizing Time of Flight (ToF) cameras. DeepCut [<xref ref-type="bibr" rid="ref-27">27</xref>] and DeeperCut [<xref ref-type="bibr" rid="ref-28">28</xref>] are the earliest proposed bottom-up approaches for pose estimation. They initially detect all the key points of the human body from the image. These key points are then treated as nodes of an undirected graph, with the keypoint correlations serving as the weights between graph nodes. Finally, instance discrimination is modeled as an integer linear programming problem. However, this kind of problem belongs to the NP problems, and implementing integer linear programming on a complete graph incurs extremely high computational complexity. Newell et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] introduced the Associative Embedding method, which assigns a token value to each keypoint by generating a heatmap. This method then uses clustering to group key points with similar token values, ensuring clear distinction between individuals. Cao et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] developed an open-source real-time system, called OpenPose, for multi-person 2D pose detection, including body, foot, hand, and facial key points. OpenPose introduces Part Affinity Fields (PAFs) to achieve fast keypoint connections, which are a set of 2D vector fields used to encode the position and orientation of limbs in the image domain. Similarly, Kreiss et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] designed the PifPaf network to predict the Part Intensity Field (PIF) that represents the positions of body key points, and the Part Association Field that represents the strength of associations between body key points. Finally, a greedy strategy is employed for instance matching. In contrast to methods that focus on partitioning and matching key points, some works are concerned with more accurately predicting instance-independent key points in the image. Cheng et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] proposed a HigherHRNet network to address the scale variations challenges in multi-person pose estimation. Wang et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] demonstrated that a high-resolution branch is unnecessary for a low-computation-region model through a progressive shrinking experiment. They proposed a fusion deconvolution head to eliminate redundant details and improve the performance of the bottom-up pose estimation model by utilizing large convolutional kernels. Although the bottom-up approach is generally faster than the top-down approach, requiring only one pose estimation step, it has its limitations. Notably, the network in bottom-up methods cannot directly obtain features from the frames, resulting in a lower average resolution per person during training compared to top-down methods, when using the same network and GPUs [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Abnormal Action Recognition</title>
<p>The key to identifying abnormal actions is to understand the definition of an exception, which may vary depending on the monitoring scenario. Additionally, an appropriate anomaly detection method should be selected by analyzing the characteristics of the data currently obtained, and the detection results are explained at last. At present, the most used methods for detecting anomalies in distributions include the deviation-based detection algorithm [<xref ref-type="bibr" rid="ref-35">35</xref>], the distance-based detection algorithm [<xref ref-type="bibr" rid="ref-36">36</xref>], etc. The deviation-based detection method compares the main features of the data, with a large deviation if the data in a certain section differs from other data objects. However, this method is not applicable to complex objects with multiple attributes since the requirement of obtaining the main features of the data first. The challenge of the distance-based detection algorithm lies in selecting the appropriate distance function. The basic distance function may not be sufficient for calculating distances of multi-dimensional features, due to the increase in feature dimension. Therefore, it is important to select an appropriate distance function based on the actual application scenario. The distribution-based detection method assumes that the data conforms to a certain probability distribution model and uses inconsistency to judge whether the data matches to determine the isolated points. The density-based detection algorithm utilizes the local anomaly factor to represent the anomaly degree of the data, which primarily relies on genuine anomalies and the correlation of local nearest neighbor densities of the objects to be detected.</p>
<p>The rapid growth of trajectory data holds significant potential for mining individual behavior patterns. Compared with the individual movement of a person at a certain moment, the pedestrian&#x2019;s trajectory can be represented in time and space. Clustering is an effective method to mine trajectory patterns. Based on human trajectory clustering, trajectory clusters are obtained, and behavior patterns of specific people are mined. Kang et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] used dynamic hierarchical clustering to describe a certain type of action pattern by using the central trajectory, and measured the similarity between the detection trajectory and the central trajectory to determine whether the anomaly was found. Lee et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] proposed a trajectory segmentation detection framework and the trajectory clustering algorithm TRACLUS, which introduces the standard MDL (Minimum Description Length) widely adopted in information compression to extract the velocity feature points of the trajectory. Liu et al. [<xref ref-type="bibr" rid="ref-39">39</xref>] proposed a density-based Trajectory Outlier Detection method DBTOD (Density-based Trajectory Outlier Detection) based on TRADO, which can detect more anomaly points.</p>
<p>Some researchers combine the global and segmented difference measurement methods to test the abnormal trajectory. For instance, Wang et al. [<xref ref-type="bibr" rid="ref-40">40</xref>] reconstructed the trajectory and represented it as a symbolic sequence. Atluri et al. [<xref ref-type="bibr" rid="ref-41">41</xref>] used clustering to identify spatio-temporal anomalies by checking time-domain and space-domain information. Ullah et al. [<xref ref-type="bibr" rid="ref-42">42</xref>] proposed an LSTM network based on an attention mechanism, incorporating Channel Attention and Spatial Attention modules. These modules enhance the most useful information in videos. Farooq et al. [<xref ref-type="bibr" rid="ref-43">43</xref>] proposed a method based on motion shapes and deep learning to detect anomalous behaviors in high-density crowds, specifically focusing on crowd dispersal behaviors. Morris et al. [<xref ref-type="bibr" rid="ref-44">44</xref>] evaluated different similarity measures and clustering methods to find out their advantages and disadvantages in locus clustering. Current research indicates that pedestrian movement is often complex. By extracting the inherent features of the moving target, such as size, color, and posture, and combining them with the motion state information of the pedestrian&#x2019;s speed and direction, we can obtain a measurement method that represents richer trajectory information. This can then be used to identify and track the target, providing significant assistance in acquiring motion trajectory and constructing a normal trajectory model.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>The Proposed System</title>
<p>To detect abnormal actions such as fall-down and trespass in electric power training scenarios, we have developed a new abnormal action recognition system that uses the Lightweight Pose Estimation Network (LPEN). The overall framework of the proposed technology is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The framework comprises four main components: Surveillance in an Electric Power Training Scenario for frame capture; a Lightweight Pose Estimation Network (LPEN) for human body key points detection; a Pose-aware Action Analysis Module (PAAM) for obtaining the main factors utilized in abnormal action detection; and an Abnormal Action Inference Module (AAIM) for the detection of abnormal fall-down and trespass.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The framework of the proposed Abnormal Action Recognition (AAR) system. This framework consists of four modules, namely Surveillance in Electric Power Training Scenario, Lightweight Pose Estimation Network (LPEN), Pose-aware Action Analysis Module (PAAM), and Abnormal Action Inference Modules (AAIM)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_50435-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Lightweight Pose Estimation Network</title>
<p>An LPEN (Lightweight Pose Estimation Network) is employed to detect the key points of the human body. This network is constructed based on the OpenPose framework, with a lighter MobileNet V1 [<xref ref-type="bibr" rid="ref-45">45</xref>] replacing the VGG-19 backbone. This modification aims to maintain high accuracy while making the model more lightweight, improving recognition efficiency, and lowering hardware processing requirements. The proposed LPEN not only replaces its backbone network but also simplifies the multi-stage structure of OpenPose, including only an Initial Stage and a Refinement Stage. The architecture of LPEN is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The architecture of the Lightweight Pose Estimation Network (LPEN). Given the video sequence within the human action as the input, the proposed LPEN can automatically obtain the human pose with the 2D skeleton positions</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_50435-fig-2.tif"/>
</fig>
<p>Specifically, color images of size <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>h</mml:mtext></mml:mrow></mml:math></inline-formula> are analyzed by the pre-trained MobileNet V1 network to produce a set of feature maps <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:math></inline-formula>. The set of feature maps is input to the Initial Stage where the network outputs a set of Part Affinity Fields (PAFs): <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the CNNs used for inference at this stage. In the subsequent Refinement Stage, the initial PAFs predictions <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and the original image features <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:math></inline-formula> are combined to produce more refined PAFs predictions: <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the CNNs used in this stage. Iteratively, the above process starts from the latest PAFs prediction to repeatedly detect confidence maps: <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03C1;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03C1;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03C1;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the CNNs used for inference at stage <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:math></inline-formula>.</p>
<p>As the confidence maps are based on the latest and most refined PAFs predictions, the difference between stages is minimal. To guide the network to predict PAFs in the first branch and confidence maps, a loss function is applied at the end of each stage. Specifically, two loss functions <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msubsup><mml:mrow><mml:mtext>f</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msubsup><mml:mrow><mml:mtext>f</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> for the PAFs branch and the confidence map branch at stage <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow></mml:math></inline-formula> are defined as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mi>W</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>&#x22C5;</mml:mo><mml:mo>&#x2225;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msubsup><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></disp-formula>
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>J</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mi>W</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>&#x22C5;</mml:mo><mml:mo>&#x2225;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msubsup><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>J</mml:mi></mml:math></inline-formula> is the number of confidence maps for different body parts, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>C</mml:mi></mml:math></inline-formula> is the count of vector fields corresponding to pairs of these parts, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the ground truth of PAFs, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msubsup><mml:mi>S</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the ground truth of the confidence maps, and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>W</mml:mi></mml:math></inline-formula> is a binary mask that is zero when pixel <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>p</mml:mi></mml:math></inline-formula> lacks annotation. The final objective of the whole network is to minimize the sum of loss functions across all layers.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>f</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>Overall, images captured by the monitoring system are first processed by a pre-trained MobileNet to extract image features, then rapidly detect key points in the Initial Stage, and further optimized in the Refinement Stage to achieve more accurate pose estimation.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Pose-Aware Action Analysis Module</title>
<p>The Pose-Aware Action Analysis Module (PAAM) is used to obtain the positional coordinates of human skeletal points in the image. Through these obtained coordinates, we could calculate some critical parameters such as the speed of the body&#x2019;s center of gravity descent <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>fall</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, the angle deviation of the body trunk from the vertical axis <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow></mml:math></inline-formula>, the ratio of the body&#x2019;s boundary height to width <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mrow><mml:mtext>ratio</mml:mtext></mml:mrow></mml:math></inline-formula>, as well as the positions of the left and right ankle <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>ankles</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>ankles</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. These parameters contribute to providing a basis for the judgment of the Abnormal State Inference Module (AAIM) in the subsequent analysis.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Abnormal Action Inference Module</title>
<p>The Abnormal Action Inference Module (AAIM) consists of two modules: The Fall-down Inference Module and the Trespass Inference Module. The former is designed to detect abnormal actions related to fall-down, while the latter is responsible for detecting trespass.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Fall-Down Inference Module</title>
<p>The Fall-down Inference Module aims to determine whether a person has fallen. Fall-down is a momentary action accompanied by rapid changes in body posture. We utilize the descent speed of the body&#x2019;s center of gravity, the angle of deviation between the body trunk and the vertical axis of the frame, and the height-to-width ratio of the body boundary as key indicators to assess the occurrence of a falling. Specifically, we represent the body&#x2019;s center of gravity <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cg</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>cg</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> using the midpoint between the left hip <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lHip</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lHip</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and right hip <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>rHip</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>rHip</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The formula for calculating the coordinates of this midpoint is as follows:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:mfrac><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:math></disp-formula></p>
<p>The variation in the position of this midpoint is utilized to estimate the descent speed. This speed can be instantaneous or the average speed within a specific time window. In this system, we choose thelatter one to estimate the descent speed. The descent speed <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>fall</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> can be calculated using the following formula:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2248;</mml:mo><mml:mfrac><mml:msqrt><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the position of body center in the first frame and <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the position of body center after <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>t</mml:mi></mml:math></inline-formula> seconds (after 5 frames in our work).</p>
<p>The descent speed of the center of gravity is a crucial indicator for determining a fall, as fall-down action typically involves a rapid increase in the center of gravity&#x2019;s speed. In addition, we also analyze the tilt angle of the human body trunk relative to the vertical axis <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow></mml:math></inline-formula>. In this work, the line connecting the neck and the center of gravity represents the human body trunk. Thus <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow></mml:math></inline-formula> can be formulated as:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>&#x03B8;</mml:mi><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the position of the neck.</p>
<p>Typically, when the human body trunk deviates from the vertical axis by 15&#x2013;30 degrees, it may indicate a risk of falling. Therefore, it is also an important auxiliary condition for determining whether a fall has occurred. To reduce the possibility of misjudgment in special cases such as bending or squatting, we have added a third criterion: The height-to-width ratio of the body bounding box:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mi>h</mml:mi><mml:mi>w</mml:mi></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>h</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>w</mml:mi></mml:math></inline-formula> are the height and width of the body bounding box, respectively. When a person is upright, this ratio tends to be relatively high. Conversely, when a person falls, the ratio decreases due to the more horizontal orientation of the body. Therefore, it can also serve as a measure to determine a fall. To assess a fall event more accurately, we choose to comprehensively consider three core parameters. For example, even if someone&#x2019;s descent speed is not particularly fast, the system can still identify a fall if there are significant changes in trunk tilt angle or height-to-width ratio. To achieve this, we assign weights to each indicator based on experimental data and prior experience, ensuring the influence of each condition in the final decision. We first standardize each parameter:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">f</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">l</mml:mi><mml:mi mathvariant="bold-italic">l</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mrow><mml:mo>,</mml:mo></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi mathvariant="bold-italic">v</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow></mml:msub><mml:mo mathvariant="bold">,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow></mml:msub><mml:mo mathvariant="bold">,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the maximum value of each parameter. Thus we would obtain a score to estimate falling:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Finally, the score obtained above is compared with the predefined threshold <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. If the score surpasses this threshold, the system determines that a falling has occurred, recording the timestamp of the falling. The formula is as follows:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>I</mml:mi><mml:mi>s</mml:mi><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:mi>S</mml:mi><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The analytical approach provides a more holistic and flexible framework for assessing the risk of fall-down, significantly enhancing its practical value in safety scenarios within the context of power training.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Trespass Inference Module</title>
<p>The Trespass Inference Module (TIM) is utilized to assess whether personnel have entered prohibited areas. The system continuously tracks the ankle positions of individuals in the video and compares them with predefined boundaries of restricted zones. A restricted zone is defined as a polygon or rectangular area <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow></mml:math></inline-formula> with boundaries determined by a series of coordinate points <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mtext>boundary</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. If the ankle coordinates are within the restricted zone, the system deems it an abnormal trespass and issues a warning. To reduce false alarms, we consider the temporal variation in ankle positions <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>pass</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. A time threshold <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is presented to determine whether the ankles have remained within the restricted zone for a duration exceeding the threshold. For instance, brief contact with the edge of the restricted zone would not trigger an immediate alert, differentiating between actual trespass and incidental or unintentional touching the crossings. Meanwhile, to mitigate the impact on the accuracy of the pose estimation algorithm, we have appropriately expanded the edges of the restricted zone to eliminate potential jitter issues. Let the expanded region as <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msup><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, the width of expansion as <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow></mml:math></inline-formula>, and the new restricted zone <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msup><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is computed as follows:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt><mml:mo>&#x2264;</mml:mo><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mi>P</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The final judgment formula for the Trespass Inference Module is:</p>
<p><disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>I</mml:mi><mml:mi>s</mml:mi><mml:mi>P</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the poison of the ankles acquired in the LPEN.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<p>To well validate the performance of the proposed system, we conduct validation experiments on the collected data.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset</title>
<p>We collect 300 original videos from the surveillance in the electric power training scenario, encompassing a broad spectrum of situational occurrences. First, each video is customized and sampled as one video clip containing the fall-down, non-fall-down, trespass, and non-trespass motions. The videos were recorded indoors by eight personnel, varying in age, gender, and body type, to ensure effective motion recognition across diverse individuals. These video clips were categorized into left and right viewing angles, with lengths ranging from 19 to 57 s, and a uniform resolution of 1280 &#x00D7; 720. Second, we employ 10 students to annotate the video clip based on the semantic motion by choosing one label, namely fall-down, non-fall-down, trespass, or non-trespass. Finally, we obtained 85 video clips with the label fall-down, 65 video clips with the label of non-fall-down, 80 video clips with the label of trespass, and 70 video clips with the label of non-trespass. The training and testing set are divided as 2/3 for training, and 1/3 for testing, which are detailed in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>The statistics of the dataset</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Types</th>
<th align="center" colspan="4">Category</th>
</tr>
<tr>
<th>Fall-down</th>
<th>Non-fall-down</th>
<th>Trespass</th>
<th>Non-trespass</th>
</tr>
</thead>
<tbody>
<tr>
<td>Training set</td>
<td>57</td>
<td>43</td>
<td>53</td>
<td>47</td>
</tr>
<tr>
<td>Testing set</td>
<td>28</td>
<td>22</td>
<td>27</td>
<td>23</td>
</tr>
<tr>
<td>Dataset</td>
<td>85</td>
<td>65</td>
<td>80</td>
<td>70</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Result and Analysis</title>
<p>To illustrate the superior performance of the proposed method, we compare the proposed system and the OpenPose-based system. Here, for fair comparison, the OpenPosed-based system is instead of the proposed LPEN and embedded into the proposed system. The recognition accuracy obtained by these two methods on the collected dataset is listed in <xref ref-type="table" rid="table-2">Table 2</xref>. We can see that the proposed method achieves the highest accuracy on both the recognition tasks of fall-down and trespass. It is noted that the accuracy of the proposed system in terms of trespass recognition is higher than that in terms of fall-down recognition. This is because the motion of trespass is more obvious than the motion of fall-down in the visual space.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>The comparison of recognition accuracy between the proposed method and the related method</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Method</th>
<th align="center" colspan="3">Recognition accuracy (%)</th>
</tr>
<tr>
<th>Fall-down</th>
<th>Trespass</th>
<th>Overall</th>
</tr>
</thead>
<tbody>
<tr>
<td>OpenPose-based system</td>
<td>95.2</td>
<td>97.5</td>
<td>96.35</td>
</tr>
<tr>
<td>The proposed system</td>
<td>97.6</td>
<td>98.7</td>
<td>98.15</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>LPEN aims to estimate the positions of all human skeletons. To validate the estimation accuracy of the proposed LPEN in terms of the pose estimation, we adopt the PCK (Percentage of Correct Key points) evaluation to test the performance of LPEN by comparing it with OpenPose. <xref ref-type="table" rid="table-3">Table 3</xref> lists the comparison between LPEN and OpenPose. It can be seen that the proposed LPEN achieves the best estimation accuracy, namely gaining an improvement of 3.724% compared with OpenPose.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The accuracy (%) comparison between LPEN and OpenPose</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Method</th>
<th align="center" colspan="5">Body</th>
<th rowspan="2">Mean</th>
</tr>
<tr>
<th>Neck</th>
<th>Left hip</th>
<th>Right hip</th>
<th>Left ankle</th>
<th>Right ankle</th>
</tr>
</thead>
<tbody>
<tr>
<td>OpenPose</td>
<td>89.021</td>
<td>86.634</td>
<td>89.021</td>
<td>77.326</td>
<td>78.758</td>
<td>84.152</td>
</tr>
<tr>
<td>LPEN</td>
<td>94.033</td>
<td>88.544</td>
<td>89.976</td>
<td>83.532</td>
<td>83.293</td>
<td>87.876</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As mentioned before, the proposed abnormal action recognition system is not only effective for recognizing abnormal actions but also is efficient in terms of computation. This belongs to the proposed lightweight LPEN, as the key module in the proposed abnormal action recognition system. Thus, we conduct the computation comparison between the proposed LPEN and the OpenPose model in terms of the GFLOPs, Parameters, Fps, Start-up time, and Memory Consumption. The comparison results are listed in <xref ref-type="table" rid="table-4">Table 4</xref>. We can see that the proposed LPEN has lower GFLOPs, a smaller number of parameters, a shorter start-up time, and a lower Memory Consumption, compared with those of OpenPose. Meanwhile, for processing the same videos, the proposed LPEN can reach 23.59 Fps, which is higher than the 20.76 Fps of OpenPose. Overall, through the computation comparison, the proposed method shows better efficiency than the current representation model, namely OpenPose. In summary, it is evident that the utilization of our Lightweight Pose Estimation Network (LPEN) is more suitable for deployment in practical electric power training scenarios.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>The comparison of computation between the proposed LPEN and the related model. The notion &#x2193; denotes the smaller value is better, while the notion &#x2191; denotes the larger value is better</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Type</th>
<th>OpenPose</th>
<th>LPEN</th>
</tr>
</thead>
<tbody>
<tr>
<td>GFLOPs&#x2193;</td>
<td>136.1</td>
<td>9.0</td>
</tr>
<tr>
<td>Parameters (total)&#x2193;</td>
<td>25.94 M</td>
<td>4.1 M</td>
</tr>
<tr>
<td>Fps (average) &#x2191;</td>
<td>20.76</td>
<td>23.59</td>
</tr>
<tr>
<td>Start-up time (average)&#x2193;</td>
<td>0.858 s</td>
<td>0.092 s</td>
</tr>
<tr>
<td>Memory consumption (average) &#x2193;</td>
<td>232.5 MB</td>
<td>171.4 MB</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, we also conduct a qualitative analysis to illustrate the effectiveness of the proposed system. We report some results of abnormal action recognition obtained by the proposed system. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the interface of the proposed system. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the state of fall-down recognition obtained by the proposed system. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows the state of trespass recognition obtained by the proposed system. We can see that the proposed system can accurately recognize the state of fall-down action and the state of trespass action.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The interface of the proposed abnormal action recognition system</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_50435-fig-3.tif"/>
</fig><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The recognition result of fall-down action obtained by the proposed system</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_50435-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The recognition result of trespass action obtained by the proposed system</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_50435-fig-5.tif"/>
</fig>
<p>In addition to the aforementioned successful results, we also report some false results in terms of fall-down action and trespass action obtained by the proposed system, as shown in <xref ref-type="fig" rid="fig-6">Figs. 6</xref> and <xref ref-type="fig" rid="fig-7">7</xref>. The findings indicate that the system exhibits a higher accuracy rate in recognizing actions when the majority of the monitored individual&#x2019;s body is within the surveillance area. However, in some cases where only a minimal portion of the body is exposed to the surveillance field of view, the system may produce erroneous judgments due to the limitations of the pose estimation algorithm. Although such occurrences are relatively infrequent, they can be mitigated by deploying multiple cameras throughout the entire project area, thereby reducing the likelihood of such misjudgments.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The false result of fall-down action obtained by the proposed system</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_50435-fig-6.tif"/>
</fig><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>The false result of trespass action obtained by the proposed system</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_50435-fig-7.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>Electric power training is an indispensable component to ensure the safety and reliable operation of the system. To address the problem of abnormal action recognition in electric power training scenarios, we present a novel Abnormal Action Recognition (AAR) system embedded with a new Lightweight Pose Estimation Network (LPEN) for effectively and efficiently recognizing the fall-down action and the trespass action. The AAR system consists of four modules, namely Surveillance in Electric Power Training Scenario, Lightweight Pose Estimation Network (LPEN), Pose-aware Action Analysis Module (PAAM), and Abnormal Action Inference Modules (AAIM). LPEN, which includes MobileNet, Initial Stage, and Refinement Stage, is designed to capture the skeleton positions of humans. PAAM aims to acquire the positional coordinates of human skeletal points in the frame. AAIM aims to assess whether there is a current occurrence of abnormal fall-down or unauthorized trespass behavior in the current frame. In the assessment of abnormal fall-down, three aspects are considered as criteria: Falling speed, major angles of the skeletal points, and the bounding box of the human. For the unauthorized trespass, the position of the ankles is the main focus. Extensive experiments are conducted to show the effectiveness and efficiency of the proposed system.</p>
</sec>
</body>
<back>
<ack>
<p>Many thanks to Yang Zhu, Bo Jiao and other volunteers during data collection.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work has been supportted by Natural Science Foundation of Jiangsu Province (No. BK20230696).</p>
</sec>
<sec><title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Study conception and design: Yunfeng Cai, Qing Yang; data collection: Long Zhang; analysis and interpretation of results: Jin Tang, Xiaotian Bi; draft manuscript preparation: Yunfeng Cai, Ran Qin, Qing Yang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>Due to the nature of this research and the privacy information of volunteers, participants of this study did not agree for their data to be shared publicly.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Alsumaidaee</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Yaw</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Koh</surname></string-name></person-group>, &#x201C;<article-title>Review of medium-voltage switchgear fault detection in a condition-based monitoring system by using deep learning</article-title>,&#x201D; <source>Energies</source>, vol. <volume>15</volume>, no. <issue>18</issue>, pp. <fpage>6762</fpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.3390/en15186762</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Mi</surname></string-name></person-group>, &#x201C;<article-title>A review of foreign object detection (FOD) for inductive power transfersystems</article-title>,&#x201D; <source>eTransportation</source>, vol. <volume>15</volume>, no. <issue>11</issue>, pp. <fpage>102</fpage>&#x2013;<lpage>116</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1016/j.etran.2019.04.002</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Khalaj</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Scherer</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Siriwardana</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Halgamuge</surname></string-name></person-group>, &#x201C;<article-title>Increasing the thermal efficiency of an operational data center using cold aisle containment</article-title>,&#x201D; <conf-name>presented at the ICIAfS</conf-name>, <publisher-loc>Colombo, Sri Lanka</publisher-loc>, <year>Dec. 22&#x2013;24, 2014</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. K.</given-names> <surname>Tam</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Badra</surname></string-name>, <string-name><given-names>R. J.</given-names> <surname>Marceau</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Marin</surname></string-name>, and <string-name><given-names>A. S.</given-names> <surname>Malowany</surname></string-name></person-group>, &#x201C;<article-title>A Web-based virtual environment for operator training</article-title>,&#x201D; <source>IEEE</source>, vol. <volume>14</volume>, no. <issue>3</issue>, pp. <fpage>802</fpage>&#x2013;<lpage>808</lpage>, <year>1999</year>. doi: <pub-id pub-id-type="doi">10.1109/59.780889</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. Z.</given-names> <surname>Dong</surname></string-name> and <string-name><given-names>F. N.</given-names> <surname>Catbas</surname></string-name></person-group>, &#x201C;<article-title>A review of computer vision-based structural health monitoring at local and global levels</article-title>,&#x201D; <source>Struct. Health Monit.</source>, vol. <volume>20</volume>, no. <issue>2</issue>, pp. <fpage>692</fpage>&#x2013;<lpage>743</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1177/1475921720935585</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Andriluka</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Pishchulin</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Gehler</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Schiele</surname></string-name></person-group>, &#x201C;<article-title>2D human pose estimation: New benchmark and state of the art analysis</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Columbus, OH, USA</publisher-loc>, <year>Jun. 23&#x2013;28, 2014</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. M.</given-names> <surname>Haralick</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Joo</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhuang</surname></string-name>, <string-name><given-names>V. G.</given-names> <surname>Vaidya</surname></string-name> and <string-name><given-names>M. B.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Pose estimation from corresponding point data</article-title>,&#x201D; <source>IEEE Trans. Syst.</source>, vol. <volume>19</volume>, no. <issue>6</issue>, pp. <fpage>1426</fpage>&#x2013;<lpage>1446</lpage>, <year>1989</year>. doi: <pub-id pub-id-type="doi">10.1016/B978-0-12-266719-0.50006-3</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Erol</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Bebis</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Nicolescu</surname></string-name>, and <string-name><given-names>R. D.</given-names> <surname>Boyle</surname></string-name></person-group>, &#x201C;<article-title>Vision-based hand pose estimation: A review</article-title>,&#x201D; <source>Comput. Vis. Image Underst.</source>, vol. <volume>108</volume>, no. <issue>1&#x2013;2</issue>, pp. <fpage>52</fpage>&#x2013;<lpage>73</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Long</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Ding</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Wen</surname></string-name></person-group>, &#x201C;<article-title>Graph-PCNN: Two stage human pose estimation with graph pose refinement</article-title>,&#x201D; <conf-name>presented at the ECCV</conf-name>, <publisher-loc>Glasgow, Scotland</publisher-loc>, <year>Aug. 23&#x2013;28, 2020</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Peng</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>A survey of convolutional neural networks: Analysis, applications, and prospects</article-title>,&#x201D; <source>IEEE Trans. Neural Netw. Learn. Syst.</source>, vol. <volume>33</volume>, no. <issue>12</issue>, pp. <fpage>6999</fpage>&#x2013;<lpage>7019</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TNNLS.2021.3084827</pub-id>; <pub-id pub-id-type="pmid">34111009</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Zhu</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Shen</surname></string-name></person-group>, &#x201C;<article-title>Deep learning-based human pose estimation: A survey</article-title>,&#x201D; <source>ACM Comput. Surv.</source>, vol. <volume>56</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>37</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1145/1122445.1122456</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A. S.</given-names> <surname>Micilotta</surname></string-name>, <string-name><given-names>E. J.</given-names> <surname>Ong</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Bowden</surname></string-name></person-group>, &#x201C;<article-title>Real-time upper body detection and 3D pose estimation in monoscopic images</article-title>,&#x201D; <conf-name>presented at the ECCV</conf-name>, <publisher-loc>Graz, Austria</publisher-loc>, <year>May 7&#x2013;13, 2006</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Hidalgo</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Single-network whole-body pose estimation</article-title>,&#x201D; <conf-name>presented at the ICCV</conf-name>, <publisher-loc>Seoul, South Korea</publisher-loc>, <year>Oct. 27&#x2013;Nov. 2, 2019</year>, pp. <fpage>6982</fpage>&#x2013;<lpage>6991</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Nibali</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>He</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Morgan</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Prendergast</surname></string-name></person-group>, &#x201C;<article-title>3D human pose estimation with 2D marginal heatmaps</article-title>,&#x201D; <conf-name>presented at the WACV</conf-name>, <publisher-loc>Wailea, HI, USA</publisher-loc>, <year>Jan. 7&#x2013;11, 2019</year>, pp. <fpage>1477</fpage>&#x2013;<lpage>1485</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Toshev</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Szegedy</surname></string-name></person-group>, &#x201C;<article-title>DeepPose: Human pose estimation via deep neural networks</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Columbus, OH, USA</publisher-loc>, <year>Jun. 23&#x2013;28, 2014</year>, pp. <fpage>1653</fpage>&#x2013;<lpage>1660</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J. J.</given-names> <surname>Tompson</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jain</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>LeCun</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Bregler</surname></string-name></person-group>, &#x201C;<article-title>Joint training of a convolutional network and a graphical model for human pose estimation</article-title>,&#x201D; <conf-name>presented at the NIPS</conf-name>, <publisher-loc>Montreal, Canada</publisher-loc>, <year>Dec. 1&#x2013;5, 2014</year>, pp. <fpage>1799</fpage>&#x2013;<lpage>1807</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. E.</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Ramakrishna</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Kanade</surname></string-name></person-group>, &#x201C;<article-title>Convolutional pose machines</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>Jun. 26&#x2013;Jul. 1, 2016</year>, pp. <fpage>4727</fpage>&#x2013;<lpage>4732</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Newell</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Yang</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Deng</surname></string-name></person-group>, &#x201C;<article-title>Stacked hourglass networks for human pose estimation</article-title>,&#x201D; <conf-name>presented at the ECCV</conf-name>, <publisher-loc>Amsterdam, Netherlands</publisher-loc>, <year>Oct. 11&#x2013;14, 2016</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Chu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Ouyang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>A. L.</given-names> <surname>Yuille</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Multi-context attention for human pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Honolulu, HI, USA</publisher-loc>, <year>Jul. 21&#x2013;26, 2017</year>, pp. <fpage>1831</fpage>&#x2013;<lpage>1840</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>Rethinking the heatmap regression for bottom-up human pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>,  <year>Jun. 19&#x2013;25, 2021</year>, pp. <fpage>13264</fpage>&#x2013;<lpage>13273</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Guo</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>The devil is in the details: Delving into unbiased data processing for human pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, <year>Jun. 14&#x2013;19, 2020</year>, pp. <fpage>5700</fpage>&#x2013;<lpage>5709</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Groos</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Ramampiaro</surname></string-name>, and <string-name><given-names>E. A. F.</given-names> <surname>Ihlen</surname></string-name></person-group>, &#x201C;<article-title>EfficientPose: Scalable single-person pose estimation</article-title>,&#x201D; <source>Appl. Intell.</source>, vol. <volume>51</volume>, no. <issue>4</issue>, pp. <fpage>2518</fpage>&#x2013;<lpage>2533</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1007/s10489-020-01918-7</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H. S.</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>Y. W.</given-names> <surname>Tai</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Lu</surname></string-name></person-group>, &#x201C;<article-title>RMPE: Regional multi-person pose estimation</article-title>,&#x201D; <conf-name>presented at the ICCV</conf-name>, <publisher-loc>Venice, Italy</publisher-loc>, <year>Oct. 22&#x2013;29, 2017</year>, pp. <fpage>2334</fpage>&#x2013;<lpage>2343</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Peng</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Cascaded pyramid network for multi-person pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Salt Lake City, UT, USA</publisher-loc>, <year>Jun. 18&#x2013;22, 2018</year>, pp. <fpage>7103</fpage>&#x2013;<lpage>7112</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wu</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Wei</surname></string-name></person-group>, &#x201C;<article-title>Simple baselines for human pose estimation and tracking</article-title>,&#x201D; <conf-name>presented at the ECCV</conf-name>, <publisher-loc>Munich, Germany</publisher-loc>, <year>Sep. 8&#x2013;14, 2018</year>, pp. <fpage>466</fpage>&#x2013;<lpage>481</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Rodrigues</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Top-down human pose estimation with depth images and domain adaptation</article-title>,&#x201D; <conf-name>presented at the VISIGRAPP</conf-name>, <publisher-loc>Prague, Czech Republic</publisher-loc>, <year>Feb. 25&#x2013;27, 2019</year>, pp. <fpage>281</fpage>&#x2013;<lpage>288</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Pishchulin</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Insafutdinov</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Tang</surname></string-name></person-group>, &#x201C;<article-title>DeepCut: Joint subset partition and labeling for multi-person pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>Jun. 26&#x2013;Jul. 1, 2016</year>, pp. <fpage>4929</fpage>&#x2013;<lpage>4937</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Insafutdinov</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Pishchulin</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Andres</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Andriluka</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Schiele</surname></string-name></person-group>, &#x201C;<article-title>DeeperCut: A deeper, stronger, and faster multi-person pose estimation model</article-title>,&#x201D; <conf-name>presented at the ECCV</conf-name>, <publisher-loc>Amsterdam, Netherlands</publisher-loc>, <year>Oct. 11&#x2013;14, 2016</year>, pp. <fpage>34</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Newell</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Huang</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Deng</surname></string-name></person-group>, &#x201C;<article-title>Associative embedding: End-to-end learning for joint detection and grouping</article-title>,&#x201D; <conf-name>presented at the NIPS</conf-name>, <publisher-loc>Long Beach, CA, USA</publisher-loc>, <year>Dec. 4&#x2013;9, 2017</year>, pp. <fpage>2277</fpage>&#x2013;<lpage>2287</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Cao</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Simon</surname></string-name>, <string-name><given-names>S. E.</given-names> <surname>Wei</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Sheikh</surname></string-name></person-group>, &#x201C;<article-title>Realtime multi-person 2D pose estimation using part affinity fields</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Honolulu, HI, USA</publisher-loc>, <year>Jul. 21&#x2013;26, 2017</year>, pp. <fpage>7291</fpage>&#x2013;<lpage>7299</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kreiss</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bertoni</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Alahi</surname></string-name></person-group>, &#x201C;<article-title>PIFPAF: Composite fields for human pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Long Beach, CA, USA</publisher-loc>, <year>Jun. 16&#x2013;20, 2019</year>, pp. <fpage>11977</fpage>&#x2013;<lpage>11986</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>T. S.</given-names> <surname>Huang</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>HigherHRNet: Scale-aware representation learning for bottom-up human pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, <year>Jun. 13&#x2013;19, 2020</year>, pp. <fpage>5386</fpage>&#x2013;<lpage>5395</lpage>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Cai</surname></string-name>, <string-name><given-names>W. M.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Han</surname></string-name></person-group>, &#x201C;<article-title>LitePose: Efficient architecture design for 2D human pose estimation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <year>Jun. 19&#x2013;25, 2022</year>, pp. <fpage>13126</fpage>&#x2013;<lpage>13136</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Dang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yin</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Zheng</surname></string-name></person-group>, &#x201C;<article-title>Deep learning based 2D human pose estimation: A survey</article-title>,&#x201D; <source>Tsinghua Sci. Technol.</source>, vol. <volume>24</volume>, no. <issue>6</issue>, pp. <fpage>663</fpage>&#x2013;<lpage>676</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1145/1122445.1122456</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C. C.</given-names> <surname>Aggarwal</surname></string-name> and <string-name><given-names>P. S.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>Outlier detection for high dimensional data</article-title>,&#x201D; <conf-name>presented at the SIGMOD Conf.</conf-name>, <publisher-loc>Santa Barbara, CA, USA</publisher-loc>, <year>May 21&#x2013;24, 2001</year>, pp. <fpage>37</fpage>&#x2013;<lpage>46</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ramaswamy</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Rastogi</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Shim</surname></string-name></person-group>, &#x201C;<article-title>Efficient algorithms for mining outliers from large data sets</article-title>,&#x201D; <conf-name>presented at the SIGMOD Conf.</conf-name>, <publisher-loc>Dallas, TX, USA</publisher-loc>, <year>May 16&#x2013;18, 2000</year>, pp. <fpage>427</fpage>&#x2013;<lpage>438</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Kang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Xing</surname></string-name></person-group>, &#x201C;<article-title>Motion pattern study and analysis from video monitoring trajectory</article-title>,&#x201D; <source>IEICE Trans. Inf. Syst.</source>, vol. <volume>97</volume>, no. <issue>6</issue>, pp. <fpage>1574</fpage>&#x2013;<lpage>1582</lpage>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.1587/transinf.E97.D.1574</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J. G.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Han</surname></string-name>, and <string-name><given-names>K. Y.</given-names> <surname>Whang</surname></string-name></person-group>, &#x201C;<article-title>Trajectory clustering: A partition-and-group framework</article-title>,&#x201D; <conf-name>presented at the SIGMOD Conf.</conf-name>, <publisher-loc>Beijing, China</publisher-loc>, <year>Jun. 12&#x2013;14, 2007</year>, pp. <fpage>593</fpage>&#x2013;<lpage>604</lpage>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Pi</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>Density-based trajectory outlier detection algorithm</article-title>,&#x201D; <source>J. Syst. Eng. Electron.</source>, vol. <volume>24</volume>, no. <issue>2</issue>, pp. <fpage>335</fpage>&#x2013;<lpage>340</lpage>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1109/jsee:2013.00042</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Schmid</surname></string-name></person-group>, &#x201C;<article-title>Action recognition with improved trajectories</article-title>,&#x201D; <conf-name>presented at the ICCV</conf-name>, <publisher-loc>Sydney, Australia</publisher-loc>, <year>Dec. 1&#x2013;8, 2013</year>, pp. <fpage>3551</fpage>&#x2013;<lpage>3558</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Atluri</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Karpatne</surname></string-name>, and <string-name><given-names>V.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Spatio-temporal data mining: A survey of problems and methods</article-title>,&#x201D; <source>ACM Comput. Surv.</source>, vol. <volume>51</volume>, no. <issue>4</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>41</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1145/3161602</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Ullah</surname></string-name>, <string-name><given-names>M. M.</given-names> <surname>Yamin</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Mohammed</surname></string-name>, <string-name><given-names>S. D.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Ullah</surname></string-name> and <string-name><given-names>F. A.</given-names> <surname>Cheikh</surname></string-name></person-group>, &#x201C;<article-title>Attention-based LSTM network for action recognition in sports</article-title>,&#x201D; <source>Electron. Imaging</source>, vol. <volume>33</volume>, no. <issue>4</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.2352/ISSN.2470-1173.2021.6.IRIACV-302</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. U.</given-names> <surname>Farooq</surname></string-name>, <string-name><given-names>M. N. M.</given-names> <surname>Saad</surname></string-name>, and <string-name><given-names>S. D.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Motion-shape-based deep learning approach for divergence behavior detection in high-density crowd</article-title>,&#x201D; <source>Vis. Comput.</source>, vol. <volume>38</volume>, no. <issue>5</issue>, pp. <fpage>1553</fpage>&#x2013;<lpage>1577</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s00371-021-02088-4</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Morris</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Trivedi</surname></string-name></person-group>, &#x201C;<article-title>Learning trajectory patterns by clustering: Experimental studies and comparative evaluation</article-title>,&#x201D; <conf-name>presented at the CVPR</conf-name>, <publisher-loc>Miami, FL, USA</publisher-loc>, <year>Jun. 20&#x2013;25, 2009</year>, pp. <fpage>312</fpage>&#x2013;<lpage>319</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Sinha</surname></string-name> and <string-name><given-names>M.</given-names> <surname>El-Sharkawy</surname></string-name></person-group>, &#x201C;<article-title>Thin MobileNet: An enhanced MobileNet architecture</article-title>,&#x201D; <conf-name>presented at the UEMCON</conf-name>, <publisher-loc>NY, USA</publisher-loc>, <year>Oct. 10&#x2013;12, 2019</year>, pp. <fpage>0280</fpage>&#x2013;<lpage>0285</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>