<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">27676</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2023.027676</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Survey on Artificial Intelligence in Posture Recognition</article-title>
<alt-title alt-title-type="left-running-head">A Survey on Artificial Intelligence in Posture Recognition</alt-title>
<alt-title alt-title-type="right-running-head">A Survey on Artificial Intelligence in Posture Recognition</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Jiang</surname><given-names>Xiaoyan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Hu</surname><given-names>Zuojin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Shuihua</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Zhang</surname><given-names>Yudong</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>yudongzhang@ieee.org</email></contrib>
<aff id="aff-1"><label>1</label><institution>School of Mathematics and Information Science, Nanjing Normal University of Special Education</institution>, <addr-line>Nanjing, 210038</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Computing and Mathematical Sciences, University of Leicester</institution>, <addr-line>Leicester, LE1 7RH</addr-line>, <country>UK</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yudong Zhang. Email: <email>yudongzhang@ieee.org</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic"><year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>21</day><month>4</month><year>2023</year>
</pub-date>
<volume>137</volume>
<issue>1</issue>
<fpage>35</fpage>
<lpage>82</lpage>
<history>
<date date-type="received"><day>08</day><month>11</month><year>2022</year></date>
<date date-type="accepted"><day>05</day><month>1</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Jiang et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Jiang et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_27676.pdf"></self-uri>
<abstract>
<p>Over the years, the continuous development of new technology has promoted research in the field of posture recognition and also made the application field of posture recognition have been greatly expanded. The purpose of this paper is to introduce the latest methods of posture recognition and review the various techniques and algorithms of posture recognition in recent years, such as scale-invariant feature transform, histogram of oriented gradients, support vector machine (SVM), Gaussian mixture model, dynamic time warping, hidden Markov model (HMM), lightweight network, convolutional neural network (CNN). We also investigate improved methods of CNN, such as stacked hourglass networks, multi-stage pose estimation networks, convolutional pose machines, and high-resolution nets. The general process and datasets of posture recognition are analyzed and summarized, and several improved CNN methods and three main recognition techniques are compared. In addition, the applications of advanced neural networks in posture recognition, such as transfer learning, ensemble learning, graph neural networks, and explainable deep neural networks, are introduced. It was found that CNN has achieved great success in posture recognition and is favored by researchers. Still, a more in-depth research is needed in feature extraction, information fusion, and other aspects. Among classification methods, HMM and SVM are the most widely used, and lightweight network gradually attracts the attention of researchers. In addition, due to the lack of 3D benchmark data sets, data generation is a critical research direction.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Posture recognition</kwd>
<kwd>artificial intelligence</kwd>
<kwd>machine learning</kwd>
<kwd>deep neural network</kwd>
<kwd>deep learning</kwd>
<kwd>transfer learning</kwd>
<kwd>feature extraction</kwd>
<kwd>classification</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In recent years, posture recognition has been a research hotspot in computer vision and artificial intelligence (AI) [<xref ref-type="bibr" rid="ref-1">1</xref>], which analyzes the original information of the target object captured by a sensor device or camera through a series of algorithms to obtain the posture. Human body posture recognition has broad market prospects in many application fields, such as behavior recognition, gait analysis, games, animation, augmented reality, rehabilitation testing, sports science, etc. [<xref ref-type="bibr" rid="ref-2">2</xref>]. AI-based posture recognition has also attracted more and more attention from researchers. We retrieved literature on AI-based posture recognition every year from 2000 to 2022, and the number of them showed an increasing trend, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The number of posture recognition papers published (2000&#x2013;2022)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-1.tif"/>
</fig>
<p>Although human posture recognition has become the leading research direction in the field of posture recognition, there are also many studies on animal posture recognition, such as birds [<xref ref-type="bibr" rid="ref-3">3</xref>], pigs [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>], and cattle [<xref ref-type="bibr" rid="ref-6">6</xref>]. With the rise of artificial intelligence, more and more scholars are interested in the research of posture recognition.</p>
<p>According to the input image type, we generally divide posture recognition algorithms into two categories: algorithms based on RGB images and algorithms based on depth images. The RGB image-based recognition algorithm utilizes the contour features of the human body. For example, the edge of the human body can be described through the histogram of oriented gradients (HOG). The depth-based image algorithm mainly uses the image&#x2019;s gray value to represent the target&#x2019;s spatial position and contour. The latter is not disturbed by light, color, shadow, and clothing, but it has higher requirements for information image acquisition equipment [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>The existing posture recognition methods can be summarized into two methods. One is based on the traditional machine learning method, and the other is based on the deep neural network method. In the posture recognition method based on traditional machine learning, the traditional image segmentation algorithm is introduced to realize the segmentation of an image or action video. Then machine learning methods are used for classification, such as support vector machines (SVM), Gaussian mixture model (GMM), and hidden Markov models (HMM). The disadvantage of this method is that the representation ability of these features is limited, representative semantic information is challenging to extract from complex content, and step-by-step recognition lacks good real-time performance.</p>
<p>In the recognition method based on deep learning, the low level-feature information of the image is combined with the deep neural network to estimate and recognize the posture at a higher level. Compared with traditional machine learning algorithms, target detection networks based on deep neural networks often have stronger adaptability and can achieve higher recognition speed and accuracy [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>We conducted a systematic review based on the Preferred Reporting Items for Systematic Review and Meta-Analysis (PRISMA). Through Google scholar, Elsevier, and Springer Link, we searched the papers on the application of artificial intelligence to posture recognition. According to the title and content, we eliminated irrelevant and duplicate papers, and finally, the review included 188 papers. The PRISMA chart is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The PRISMA chart of the article selection process for this review</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-2.tif"/>
</fig>
</sec>
<sec id="s2">
<label>2</label>
<title>Main Recognition Techniques</title>
<sec id="s2_1">
<label>2.1</label>
<title>Sensor&#x2011;Based Recognition</title>
<p>The sensor-based posture recognition requires the target to wear a variety of sensors or optical symbols and collect the action information of the target object based on this. The research on sensor-based human posture recognition algorithms started earlier. As early as the 1950s, some people used gravity sensors to recognize human posture [<xref ref-type="bibr" rid="ref-10">10</xref>]. In daily human posture recognition research, sensors have been used to distinguish standing, walking, running, sitting, and other stable human posture [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>The common classification methods of posture recognition sensors are as follows. According to the position of the sensor, it can be divided into lower limbs, waist, arm, neck, wrist, etc. Sensors can also be classified according to the number of sensors, which can be divided into single-sensor and multi-sensor. Compared with the method of single-sensor signal processing, the multi-sensor system can obtain more information about the measured target and environment effectively [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<p>Whether or not the sensor is installed on the user can be divided into wearable and fixed sensors. Wearable devices are a representative example of sensor-based human activity recognition (HAR) [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. The sensor&#x2019;s type of data output can be divided into array time domain signal, image matrix data, vector data, or strap-down matrix data. Common wearable sensors include inertial sensors (such as accelerometers and gyroscopes), physiological sensors (such as EEG, ECG, GSR, EMG), pressure sensors (such as FSR, bending sensors, barometric pressure sensors, textile-based capacitive pressure sensors), vision wearable sensors (such as WVS), flexible sensors [<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p>To avoid physical discomfort and system instability caused by workers on construction sites wearing invasive sensors or attaching multiple sensors to the body, Antwi-Afari et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] utilized the network based on deep learning as well as wearable insole sensor data to automatically identify and classify various postures presented by workers during construction. Hong et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] designed a system using multi-sensors and a collaborative AI-IoT-based approach and proposed multi-pose recognition (MPR) and cascade-adaboosting-cart (CACT) posture recognition algorithms to further improve the effect of human posture recognition.</p>
<p>Fan et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed a squeezed convolutional gated attention (SCGA) model to recognize basketball shooting postures fused by various sensors. Sardar et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed a mobile sensor-based human physical activity recognition platform for COVID-19-related physical activity recognition, such as hand washing, hand disinfection, nose-eye contact, and handshake, as well as contact tracing, to minimize the spread of COVID-19.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Vision&#x2011;Based Recognition</title>
<p>The vision-based method extracts the information of the key node and skeleton by analyzing the position of each joint point of the target object in the image data. In vision-based methods, cameras are usually used to obtain images or videos that require posture recognition and can be used in a non-contact environment. Therefore, this method does not affect the comfort of motion and has low acquisition costs.</p>
<p>Obtaining human skeleton keypoints from two-dimensional (2D) images or depth images through posture estimation is the basis of vision-based posture recognition. There are inherent limitations when 2D images are used to model three-dimensional (3D) postures, so RGB-D-based methods are ineffective in practical applications. In addition to RGB images and depth maps, skeletons have become a widely used data modality for posture recognition, where skeleton data are used to construct high-level features that characterize 3D configurations of postures [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
<p>The general process of vision-based posture recognition includes the following: image data acquisition, preprocessing, feature extraction, and feature classification, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Procedure of vision-based posture recognition</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-3.tif"/>
</fig>
<p>Currently, video-based methods mainly use deep neural networks to learn relevant features from video images for posture recognition directly. For example, WMS Abedi et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] used convolutional neural networks to identify and classify different categories of human poses (such as sitting, lying, and standing) in the available frames. Tome et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] fused the probabilistic information of 3D human posture with the multi-stage CNN architecture to achieve 3D posture estimation of the original images. Fang et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] designed a visual teleoperation framework based on a deep neural network structure and posture mapping method. They applied a multi-level network structure to increase the flexibility of visual teleoperation network training and use. Kumar et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] used the integration of six independent deep neural architectures based on genetic algorithms to improve the driver&#x2019;s performance on the distraction classification problem to assist the existing driver-to-pose recognition technology. Mehrizi et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] proposed a computer vision-based label-free motion capture method that combines the discriminative method of posture estimation with morphological constraints to improve the accuracy and robustness of posture estimation.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>RF-Based Recognition</title>
<p>In some specific posture recognition situations, the target object cannot wear the sensing device, and radio frequency (RF)-based technology can solve this problem. Due to their non-contact nature, various radio frequency-based technologies are finding applications in human activity recognition. Yao et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] used radio frequency identification (RFID) technology to build a posture recognition system to identify the posture of the elderly, who do not need to wear equipment at this time. Yao et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] used RFID and machine learning algorithms to decipher signal fluctuations to identify activities. Liu et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] proposed a sleep monitoring system based on passive RFID tags, which combined hierarchical recognition, image processing, and polynomial fitting to identify body posture through changes caused by backscattered signals from tags.</p>
<p>Radio frequency signals are extremely sensitive to environmental changes, and changes caused by human movements or activities can be easily captured. Radio frequency signals are absorbed, reflected, and scattered by the body, which will cause changes in the signals. Human activities will cause different changes in the radio frequency signal so that human activities can be identified by analyzing the changes in the signals. The most typically used radio frequency technologies are radar, WiFi, and RFID [<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>].</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Traditional Machine Learning&#x2011;Based Approach</title>
<sec id="s3_1">
<label>3.1</label>
<title>Preprocessing</title>
<p>Image preprocessing is the basis of posture recognition, which can directly affect the extraction of feature points and the result of posture classification, thus affecting the recognition rate of posture. The main tasks in the preprocessing stage are denoising, human skeleton keypoint detection, scale, gray level normalization, and image segmentation.</p>
<p>The keypoint detection of the human skeleton mainly detects the keypoint information such as human joints and facial features. The output is the skeletal feature of the human body, which is the primary part of posture recognition and behavior analysis, mainly used for segmentation and alignment.</p>
<p>The normalization of scale and gray level should first ensure the effective extraction of key features of the human body and then process the color information and size of the image to reduce the amount of computation.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Feature Extraction</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Figure Format Histogram of Oriented Gradients</title>
<p>Histogram of oriented gradient (HOG) constitutes features by calculating and statistical histogram of gradient direction in local image regions [<xref ref-type="bibr" rid="ref-33">33</xref>], which describes the entire image region and reflects strong description ability and robustness. HOG classifier is generally combined with the SVM classifier in image recognition, especially in human detection, which has achieved great success. HOG can describe objects&#x2019; appearance features and the shape of local gradient distribution [<xref ref-type="bibr" rid="ref-34">34</xref>]. HOG feature extraction steps are as follows in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The flowchart of HOG feature extraction</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-4.tif"/>
</fig>
<p>(1) The color space of input images is normalized by gamma correction to reduce the influence of light factors and suppress noise interference. Gamma compression is shown in the following formula:</p>
<p><disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, the value of gamma has three conditions:</p>
<p>(i) When gamma is equal to 1, the output value is equal to the input value, and only the original image will be displayed.</p>
<p>(ii) When gamma is greater than 1, the dynamic range of the low gray value region of the input image becomes smaller, and the contrast of the low gray value region of the image is reduced. In the area of high gray value, as the dynamic range increases, the contrast in the area of high gray value of the image will be correspondingly enhanced. Eventually, the overall gray value of the image will be darkened.</p>
<p>(iii) When gamma is less than 1, the dynamic range of the low gray value region of the input image becomes larger, and the contrast of the low gray value region of the image is enhanced. In the area of high gray value, if the dynamic range becomes smaller, the contrast in the area of high gray value will decrease accordingly, thus brightening the overall gray level of the image.</p>
<p>(2) The horizontal and vertical gradient values and the gradient direction values of each pixel in the image can be calculated by the following formula:</p>
<p><disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the horizontal gradient, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the vertical gradient, and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the pixel value at pixel point (<italic>a</italic>, <italic>b</italic>).</p>
<p>The gradient amplitude and orientation at the pixel point are shown in <xref ref-type="disp-formula" rid="eqn-4">Eqs. (4)</xref> and <xref ref-type="disp-formula" rid="eqn-5">(5)</xref>, respectively:</p>
<p><disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>arctan</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>(3) The gradient orientation histogram is constructed for each cell unit to provide the corresponding code for the local image region. At the same time, the image of the human posture and appearance are kept weak sensitivity.</p>
<p>(4) Every few cell units are formed into large blocks, and the gradient intensity is normalized to realize the compression of illumination, shadow, and edge.</p>
<p>(5) All overlapping blocks in the detection window are collected for HOG features and combined into the final feature vector for classification.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Scale&#x2011;Invariant Feature Transform</title>
<p>Scale Invariant Feature Transform (SIFT) is an algorithm that maps images to local feature vector sets based on computer vision technology. The essence is to find the keypoints or feature points in different scale-spaces and then calculate the direction of the keypoints [<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>Therefore, SIFT features do not vary with the changes in image rotation, scaling, and brightness and are almost immune to illumination, affine transformation, and noise [<xref ref-type="bibr" rid="ref-36">36</xref>]. Yang et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] used SIFT feature extraction to study writing posture and achieved good results. The main steps of SIFT algorithm are as follows in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The flowchart of SIFT algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-5.tif"/>
</fig>
<p>(1) Extreme value detection of scale space</p>
<p>Images over all scale spaces are searched, and Gaussian differential functions are used to identify potential points of interest that are not affected by scale and selection. This can be done efficiently by using the &#x201C;scale space&#x201D; function as follows:</p>
<p><disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:msup><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mfrac><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a Gaussian kernel function, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the space coordinate, <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> refers to the scale space factor, and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> refers to the Gaussian scale space of the image. The purpose of establishing the scale space is to detect the feature points that exist on different scales. The Gaussian Laplacian operator (LoG) is a good operator for detecting feature points, but its computation is extremely large, so the Gaussian difference (DoG) is usually used to approximate LoG.</p>
<p><disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <italic>k</italic> is the scaling factor of two adjacent Gaussian scale spaces.</p>
<p>(2) Localization of feature points</p>
<p>At this stage, we need to remove the points that do not meet the criteria from the list of keypoints. The points that do not meet the requirements are mainly low-contrast feature points and unstable edge response points.</p>
<p>(3) Feature orientation assignment</p>
<p>One or more directions should be assigned to each keypoint location according to the local gradient direction of the image to achieve rotation invariance. To ensure the invariance of these features, scholars perform all subsequent operations on the orientation, scale, and position of the keypoints. After finding the feature point, the scale of the feature point and its scale image can be obtained:</p>
<p><disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msqrt><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>arctan</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denote the magnitude and orientation of the gradient at each point <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, respectively. After calculating the gradient direction, the gradient orientation and amplitude of the pixel near the feature point are calculated by the histogram.</p>
<p>In the histogram, the horizontal axis represents the intersection angle of the gradient orientation, the vertical axis represents the sum of the gradient amplitudes corresponding to the gradient orientation, and the orientation corresponding to the peak value is the primary orientation of the feature points.</p>
<p>(4) Generate a feature description</p>
<p>After the above operation, the feature point descriptor must be generated, containing the feature points and the pixels around them. In general, the generation of feature descriptors consists of the following steps: (i) To achieve rotation invariance, the main orientation of rotation is corrected. (ii) Generate descriptors and form 128-dimensional feature vectors. (iii) Normalize the feature vector length to remove illumination&#x2019;s influence.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Dynamic Time Warping</title>
<p>In time series analysis, dynamic time warping (DTW) is introduced to compare the similarity or distance between two arrays or time series of different lengths. DTW was initially used in speech recognition and is now widely used in posture recognition [<xref ref-type="bibr" rid="ref-38">38</xref>&#x2013;<xref ref-type="bibr" rid="ref-41">41</xref>].</p>
<p>Suppose there are two sequences denoted by <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. Here, <italic>m</italic> and <italic>n</italic> are the lengths of the two sequences, respectively. When <italic>m</italic> is equal to <italic>n</italic>, the Euclidean distance (<xref ref-type="disp-formula" rid="eqn-11">formula (11)</xref>) can be directly used to calculate the distance <italic>d</italic> between the two sequences.</p>
<p><disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>When <italic>m</italic> is not equal to <italic>n</italic>, DTW is introduced to regularize the sequence to make it matches. To align the two sequences, construct a matrix grid of <italic>m</italic> &#x00D7; <italic>n</italic>. The elements in the matrix (<italic>a</italic>, <italic>b</italic>) are the distance between <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, that is, the similarity between each point in sequence <italic>P</italic> and each point in sequence <italic>Q</italic>. The smaller the distance, the higher the similarity, and the shortest path from the start to the end. This path is called the &#x201C;warping path&#x201D; and is denoted by <italic>W</italic>. The <italic>l</italic>th element of <italic>W</italic> is defined as:</p>
<p><disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>W</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mi>L</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>m</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1.</mml:mn></mml:math></disp-formula></p>
<p>This path needs to satisfy the following constraints [<xref ref-type="bibr" rid="ref-42">42</xref>]:</p>
<p>(1) The order of each sequence part cannot be changed, and the selected path starts at the bottom left corner of the matrix and ends at the top right corner. The boundary conditions must be met, as shown in <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>:</p>
<p><disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>(2) Ensure that each coordinate in the two sequences appears in the warping path <italic>W</italic>, so a point can only be aligned with its neighboring point.</p>
<p>(3) The points above the warping path <italic>W</italic> must be monotonically progressed over time.</p>
<p>Therefore, only three directions to choose the path to each grid point. Assuming the path has already passed through grid point (<italic>a</italic>, <italic>b</italic>), the location of the next grid point to pass through can only be one of three cases: (<italic>a</italic> &#x002B; 1, b), (<italic>a</italic>, <italic>b</italic> &#x002B; 1), and (<italic>a</italic> &#x002B; 1, <italic>b</italic> &#x002B; 1), as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. We can solve the value of DTW according to <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Diagram of the path search direction</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-6.tif"/>
</fig>
<p><disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>D</mml:mi><mml:mi>T</mml:mi><mml:mi>W</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>Q</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mfrac><mml:msqrt><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:msqrt><mml:mi>L</mml:mi></mml:mfrac><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>Other Feature Extraction Approaches</title>
<p>In addition to the above two feature extraction methods (HOG [<xref ref-type="bibr" rid="ref-43">43</xref>&#x2013;<xref ref-type="bibr" rid="ref-46">46</xref>], SIFT [<xref ref-type="bibr" rid="ref-47">47</xref>&#x2013;<xref ref-type="bibr" rid="ref-49">49</xref>], DTW [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]) for posture recognition, several feature extraction methods are widely used in posture recognition, such as Hu moment invariant (HMI) [<xref ref-type="bibr" rid="ref-51">51</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>], Fourier descriptors (FD) [<xref ref-type="bibr" rid="ref-53">53</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>], nonparametric weighted feature extraction (NWFE) [<xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>], gray-level co-occurrence matrix (GLCM) [<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>].</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Feature Reduction</title>
<p>After feature extraction is completed, feature dimension reduction is needed when the dimension is too high to improve the speed and efficiency of calculation and decision-making. Principal component analysis (PCA) [<xref ref-type="bibr" rid="ref-59">59</xref>] and linear discriminant analysis (LDA) [<xref ref-type="bibr" rid="ref-60">60</xref>] are the most commonly used dimensionality reduction methods.</p>
<p>PCA aims to try to recombine the numerous original indicators with certain correlations into a new set of unrelated comprehensive indicators and then replace the original ones [<xref ref-type="bibr" rid="ref-61">61</xref>]. It is an unsupervised dimensionality reduction algorithm. LDA is a supervised linear dimensionality reduction algorithm. Unlike PCA, LDA maintains data information and makes dimensionality reduction data as easy to distinguish as possible [<xref ref-type="bibr" rid="ref-62">62</xref>].</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Classification</title>
<sec id="s3_4_1">
<label>3.4.1</label>
<title>SVM</title>
<p>Corinna Cortes et al. first proposed the support vector machine (SVM) to find the optimal solution from two types of different sample data [<xref ref-type="bibr" rid="ref-63">63</xref>]. There may be multiple partition hyperplanes for the sample space to separate the two training samples. SVM is used to find the best hyperplane to separate the training samples.</p>
<p>Therefore, the main idea of the support vector machine is to establish a decision hyperplane and realize the division of two different types of samples by obtaining the maximum distance between two types of samples closest to the plane on both sides of the plane [<xref ref-type="bibr" rid="ref-9">9</xref>], as shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. Here, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the support vector, and all of which are divided into two categories by the hyperplane.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Schematic diagram of linear support vector machine</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-7.tif"/>
</fig>
<p>The model trained by SVM is only related to the support vector, so the algorithm&#x2019;s complexity is mainly affected by the number of support vectors. Vectors and labels can define the training samples in the two-dimensional feature space. The <italic>N</italic> training samples in the m-dimensional feature space are defined as:</p>
<p><disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the <italic>i</italic>th vector of the sample space, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the category of the <italic>i</italic>th sample. If the training sample is linearly separable, we describe the hyperplane by the following equation [<xref ref-type="bibr" rid="ref-64">64</xref>]:</p>
<p><disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:mi>X</mml:mi><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></disp-formula>where <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:mo>;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> is the normal vector of the hyperplane, which determines the direction of the hyperplane. <italic>X</italic> <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:mo>;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> is the training samples. <italic>T</italic> is the transpose, <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>b</mml:mi></mml:math></inline-formula> refers to the biases, which determine the distance between the hyperplane and the origin of the space. Once the normal vector <italic>w</italic> and the biases <italic>b</italic> are determined, a partition hyperplane can be uniquely determined. The distance <italic>d</italic> from the vector <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to the hyperplane can be calculated by the following formula:</p>
<p><disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:msqrt><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:msqrt></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:mi>X</mml:mi><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>w</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>We assume that the hyperplane can classify the training samples correctly so that the following relation holds [<xref ref-type="bibr" rid="ref-65">65</xref>]:</p>
<p><disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, we define the category label of the points on and above the plane <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> as &#x201C;&#x002B;1&#x201D;, and the category label of the points on and below the plane <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> as &#x201C;&#x2212;1&#x201D;. It can be obtained that the distance <italic>d</italic> between the plane <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> is</p>
<p><disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>2</mml:mn><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>w</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, the distance <italic>d</italic> is the sum of the distances of the two outlier support vectors to the hyperplane and is called the margin. We need to find the segmentation hyperplane with the maximum marginal value, that is, the parameters <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>w</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>b</mml:mi></mml:math></inline-formula> (<xref ref-type="disp-formula" rid="eqn-20">Eq. (20)</xref>), satisfying the constraint conditions to maximize <italic>d</italic>.</p>
<p>In practice, the samples are often linearly inseparable, so it is necessary to transform the nonlinear separability into linear separability. In support vector machines, the kernel function can map samples from low-dimensional to high-dimensional space so that SVM can deal with nonlinear problems. In other words, the kernel function extends linear SVM to nonlinear SVM, which makes SVM more universal.</p>
<p>Different kernel functions correspond to different mapping methods. The SVM algorithm was initially used to deal with binary classification problems and extended on this basis. It can also deal with multiple classification problems and regression problems.</p>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>GMM</title>
<p>The Gaussian mixture model (GMM) uses the Gaussian probability density functions (normal distribution curves) to quantify the variable distribution accurately and decomposes the distribution of variables into several statistical models based on Gaussian probability density functions (normal distribution curves). Theoretically, suppose the number of Gaussian models fused by a GMM is enough, and the weights between them are set reasonably enough. In that case, the GMM can fit samples with any arbitrary distribution.</p>
<p>Suppose that the Gaussian mixture model consists of <italic>M</italic> Gaussian models, and each Gaussian is called a &#x201C;Component&#x201D;, the probability density function of GMM is as follows [<xref ref-type="bibr" rid="ref-66">66</xref>,<xref ref-type="bibr" rid="ref-67">67</xref>]:</p>
<p><disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A3;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <italic>x</italic> denotes a D-dimensional feature vector, <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the probability density function of the <italic>m</italic>th Gaussian model, which can be seen as the probability of <italic>x</italic> produced by the <italic>m</italic>th Gaussian model after selection, as shown in the following formula:</p>
<p><disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03BC;</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi>D</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mfrac><mml:mn>1</mml:mn><mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mfrac><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the weight of the <italic>m</italic>th Gaussian model, that is, the prior probability of choosing the <italic>m</italic>th Gaussian model, and satisfies <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. &#x03A3; represents the covariance of each component, and <italic>&#x03BC;</italic> represents the average value of each component. Solving the GMM model is essentially to solve these three parameters. The EM algorithm is usually used to solve this problem, which includes expectation-step (E-step) and maximization-step (M-step), as shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>The solving process of GMM based on the EM algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-8.tif"/>
</fig>
<p>(1) E-step</p>
<p>First, estimate the probability that each component generates the data. Here, we mark the probability of data <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> generated by the <italic>m</italic>th component as <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, as shown in <xref ref-type="disp-formula" rid="eqn-24">formula (24)</xref>.</p>
<p><disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>(2) M-step</p>
<p>Next, iteratively solve the parameter values according to the calculation results of the previous step.</p>
<p><disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-26"><label>(26)</label><mml:math id="mml-eqn-26" display="block"><mml:msub><mml:mi mathvariant="normal">&#x03A3;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-27"><label>(27)</label><mml:math id="mml-eqn-27" display="block"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mi>N</mml:mi></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mi>&#x03B3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, repeat the above E-M steps until the value of the log-likelihood function (<xref ref-type="disp-formula" rid="eqn-23">formula (23)</xref>) no longer changes significantly.</p>
<p><disp-formula id="eqn-28"><label>(28)</label><mml:math id="mml-eqn-28" display="block"><mml:mi>l</mml:mi><mml:mi>n</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03C0;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x03A3;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mi>l</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A3;</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_4_3">
<label>3.4.3</label>
<title>HMM</title>
<p>As we all know, the hidden Markov model (HMM) is a classic machine learning model which has proved its value in language recognition, natural language processing, pattern recognition, and other fields [<xref ref-type="bibr" rid="ref-68">68</xref>,<xref ref-type="bibr" rid="ref-69">69</xref>]. This model describes the process of generating a random sequence of unobservable states from a hidden Markov chain and then generating the observed random sequence from each state. Among them, the transition between the states and the observation sequence and the state sequence have a certain probability relationship [<xref ref-type="bibr" rid="ref-70">70</xref>]. The hidden Markov model is mainly used to model the above process.</p>
<p>We assume that <italic>M</italic> and <italic>N</italic> represent the set of all possible hidden states and the set of all possible observed states, respectively. Then <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>M</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>N</mml:mi></mml:math></inline-formula> are expressed as follows:</p>
<p><disp-formula id="eqn-29"><label>(29)</label><mml:math id="mml-eqn-29" display="block"><mml:mi>M</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <italic>P</italic> and <italic>Q</italic> are the number of possible hidden states and the number of possible observed states, respectively, which are not necessarily equal.</p>
<p>In a sequence of length <italic>T</italic>, <italic>U</italic> and <italic>V</italic> correspond to the state and observation sequences, respectively, as follows:</p>
<p><disp-formula id="eqn-30"><label>(30)</label><mml:math id="mml-eqn-30" display="block"><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where the subscript of each element represents the moment. That is, the state sequence and the observation sequence elements are successively related. Any hidden state <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>M</mml:mi></mml:math></inline-formula> and any observed state <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi></mml:math></inline-formula>. Therefore, the graph model structure of the above hidden Markov model is shown in the following <xref ref-type="fig" rid="fig-9">Fig. 9</xref>.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Graph model structure of hidden Markov model</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-9.tif"/>
</fig>
<p>To facilitate the solution, assume that the hidden state at any moment is only related to its previous hidden state. The hidden state at time <italic>t</italic> is <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the hidden state at time <italic>t</italic>&#x002B;1 is <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, then the transition probability of HMM state <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> from time <italic>t</italic> to time <italic>t</italic>&#x002B;1 can be obtained as follows:</p>
<p><disp-formula id="eqn-31"><label>(31)</label><mml:math id="mml-eqn-31" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Thus, the state transition matrix <italic>A</italic> can be obtained:</p>
<p><disp-formula id="eqn-32"><label>(32)</label><mml:math id="mml-eqn-32" display="block"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Assuming that the observed state at any moment is only related to the hidden state at the current moment when the hidden state at time <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>t</mml:mi></mml:math></inline-formula> is <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the corresponding observed state is <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, then the probability <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> generated by the observed state <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> at this time satisfies the following equation under the hidden state <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p><disp-formula id="eqn-33"><label>(33)</label><mml:math id="mml-eqn-33" display="block"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>In this way, <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> can form the probability matrix <italic>B</italic> generated by the observed state.</p>
<p><disp-formula id="eqn-34"><label>(34)</label><mml:math id="mml-eqn-34" display="block"><mml:mi>B</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>Q</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>In addition, we define the probability distribution <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mrow><mml:mi mathvariant="normal">&#x03A0;</mml:mi></mml:mrow></mml:math></inline-formula> of hidden states at time <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> as follows:</p>
<p><disp-formula id="eqn-35"><label>(35)</label><mml:math id="mml-eqn-35" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x03A0;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mrow><mml:mi mathvariant="normal">&#x03A0;</mml:mi></mml:mrow></mml:math></inline-formula> is an n-dimensional vector with each element representing the probability of being in a certain state at time <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. In this way, the initial probability distribution of hidden states <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mrow><mml:mi mathvariant="normal">&#x03A0;</mml:mi></mml:mrow></mml:math></inline-formula>, the state transition probability matrix <italic>A</italic>, and the observed state probability matrix <italic>B</italic> can determine the HMM model, which can be expressed as follows [<xref ref-type="bibr" rid="ref-71">71</xref>]:</p>
<p><disp-formula id="eqn-36"><label>(36)</label><mml:math id="mml-eqn-36" display="block"><mml:mi>&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A0;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mrow><mml:mi mathvariant="normal">&#x03A0;</mml:mi></mml:mrow></mml:math></inline-formula> and <italic>A</italic> determines the sequence of states, and <italic>B</italic> determines the sequence of observations.</p>
</sec>
<sec id="s3_4_4">
<label>3.4.4</label>
<title>Other Classification Approaches</title>
<p>In addition to the above classification algorithms (SVM [<xref ref-type="bibr" rid="ref-72">72</xref>&#x2013;<xref ref-type="bibr" rid="ref-75">75</xref>,<xref ref-type="bibr" rid="ref-43">43</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>], GMM [<xref ref-type="bibr" rid="ref-76">76</xref>,<xref ref-type="bibr" rid="ref-77">77</xref>], HMM [<xref ref-type="bibr" rid="ref-70">70</xref>,<xref ref-type="bibr" rid="ref-69">69</xref>,<xref ref-type="bibr" rid="ref-78">78</xref>]), some other classification algorithms are used for posture recognition, such as k-nearest neighbor (k-NN) [<xref ref-type="bibr" rid="ref-79">79</xref>], random forest (RF) [<xref ref-type="bibr" rid="ref-80">80</xref>&#x2013;<xref ref-type="bibr" rid="ref-82">82</xref>], Bayesian classification algorithm [<xref ref-type="bibr" rid="ref-83">83</xref>], decision tree (DT) [<xref ref-type="bibr" rid="ref-72">72</xref>,<xref ref-type="bibr" rid="ref-84">84</xref>,<xref ref-type="bibr" rid="ref-85">85</xref>], linear discriminant analysis [<xref ref-type="bibr" rid="ref-86">86</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>], na&#x00EF;ve Bayes (NB) [<xref ref-type="bibr" rid="ref-72">72</xref>,<xref ref-type="bibr" rid="ref-87">87</xref>], etc.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Deep Neural Network&#x2011;Based Approach</title>
<p>Deep learning mainly uses neural network models, such as convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), transfer learning, attention model, and long short-term memory (LSTM), as parameter structures to optimize machine learning algorithms.</p>
<p>This method is an end-to-end learning method, which does not require manual operation, but relies on the algorithm to automatically extract features, starting directly from the original input data, and automatically completes feature extraction and model learning through a hierarchical network [<xref ref-type="bibr" rid="ref-17">17</xref>]. In recent years, it has been widely used in many fields and achieved remarkable results, such as image recognition, intelligent monitoring, text recognition, semantic analysis, and other fields. Human posture recognition based on deep learning can quickly fit the human posture information in the sample label so as to generate a model with posture analysis ability.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Posture Estimation</title>
<p>Regarding network architecture, deep learning-based posture estimation is divided into a single-stage approach and a multi-stage approach. The usual difficulty of single-stage networks lies in the subsequent feature fusion work, and multi-stage networks generally repeat and superimpose a small network structure.</p>
<p>Since the number and position of people in the image are unknown in advance, multi-body posture estimation is more difficult than single-body posture estimation, which is usually divided into two ideas: top-down and bottom-up. The former is first to incorporate person detectors, then estimate each part, and finally calculate the pose of each person. The latter is to detect all parts in the image, the parts of each person, and then use a certain algorithm to associate/group the parts belonging to different people. The algorithms mainly include CPM [<xref ref-type="bibr" rid="ref-88">88</xref>], stacked hourglass networks [<xref ref-type="bibr" rid="ref-89">89</xref>], and MSPN [<xref ref-type="bibr" rid="ref-90">90</xref>]. Single-stage approaches are all Top-down, such as CPN [<xref ref-type="bibr" rid="ref-91">91</xref>] and simple baselines [<xref ref-type="bibr" rid="ref-92">92</xref>].</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Convolutional Neural Networks</title>
<p>In posture recognition, convolutional neural networks (CNN) have achieved good results. In the system designed by Yan et al. [<xref ref-type="bibr" rid="ref-93">93</xref>], CNN is used to learn and predict the preset driving posture automatically. Wang [<xref ref-type="bibr" rid="ref-94">94</xref>] used CNN to design a human posture recognition model for sports training. CNN has also been successfully used for capture posture detection [<xref ref-type="bibr" rid="ref-95">95</xref>]. Rani et al. [<xref ref-type="bibr" rid="ref-96">96</xref>] adopted the lightweight network of convolution neural network-long short-term memory (CNN-LSTM) for classical dance pose estimation and classification. Zhu et al. [<xref ref-type="bibr" rid="ref-97">97</xref>] proposed a two-flow RGB-D faster R-CNN algorithm to achieve automatic posture recognition of sows, which applied the feature level fusion strategy.</p>
<p>The neurons in each layer of convolutional neural networks are arranged in three dimensions (width, height, and depth). It should be noted that depth here refers to the number of layers of the network. The convolutional neural network is mainly composed of the input layer, convolutional layer (CL), ReLU layer, pooling layer (PL), and fully connected layer (FCL). A simple diagram of CNN is shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>A simple diagram of CNN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-10.tif"/>
</fig>
<p>The core layer of the convolutional neural network is the convolutional layer, which is composed of several convolution units. The important purposes of dimension reduction and feature extraction are achieved through convolution operation. In the first layer of the convolution layer, only some low-level features, such as edges, lines, and angles, can be extracted. In contrast, more complex posture features need to be extracted from more layers of iteration.</p>
<p>The pooling layer is sandwiched between continuous convolution layers to compress the amount of data and parameters, improve identification efficiency and effectively control overfitting. A pooling layer is actually a nonlinear form of drop sampling.</p>
<p>Generally, the full connection layer is in the last few layers and is used to make the final identification judgment. Their activation can be matrix multiplication, and then the deviation is added.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Improved Convolutional Neural Networks</title>
<p>Since the sparse network structure of the traditional CNN cannot retain the high efficiency of dense computation of a fully connected network, and the classification results are inaccurate, or the convergence speed is slow due to the low utilization of convolutional features in the experimental process, so many researchers have carried out various optimization of the CNN algorithm.</p>
<p>For example, by using batch normalization (BN), the distribution of input values of any neuron in each layer of the neural network is forced to return to the normal distribution with a mean of 0 and a variance of 1 (or other), so that the activated input values fall in the sensitive area of the input, thus avoiding the vanishing gradient [<xref ref-type="bibr" rid="ref-98">98</xref>].</p>
<p>Deep residual networks address network degradation using residual learning with identity connections [<xref ref-type="bibr" rid="ref-99">99</xref>]. CNN-LSTM provides solutions to complex problems with large amounts of data [<xref ref-type="bibr" rid="ref-96">96</xref>]. Since target tracking methods based on traditional CNN and correlation filters are usually limited to feature extraction with scale invariance, multi-scale spatio-temporal residual network (MSST-ResNet) can be used to realize multi-scale feature and spatio-temporal interaction between the flows of spatial and time [<xref ref-type="bibr" rid="ref-100">100</xref>], which is also regarded as an extension of residual network architecture. Bounding box regression and labeling from raw images via faster R-CNN showed high reliability [<xref ref-type="bibr" rid="ref-101">101</xref>]. In human posture recognition, many networks based on CNN have emerged (such as stacked hourglass networks, MSPN, CPM, and HRNet [<xref ref-type="bibr" rid="ref-102">102</xref>]).</p>
<p>Stacked hourglass networks show good performance in human posture estimation based on successive pooling and upsampling steps to capture and integrate information at all image scales. The network is combined with intermediate supervision for bottom-up, top-down repetitive processing [<xref ref-type="bibr" rid="ref-89">89</xref>]. The stacked hourglass model is formed by concatenating hourglass modules, each consisting of many residual units, pooling layers, and upsampling layers [<xref ref-type="bibr" rid="ref-103">103</xref>], so it is able to capture all information at each scale and combine these features to output pixel-level predictions.</p>
<p>In the study by Alejandro Newell et al. [<xref ref-type="bibr" rid="ref-89">89</xref>] using a single pipeline with skip layers to preserve spatial information at each resolution, the topology of the hourglass is symmetric. That is, for each layer that exists downward, there is an upper-level corresponding to it. After reaching the output resolution of the network, the final network prediction is completed by two successive rounds of <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolution.</p>
<p>The output of the network is a set of heat maps, and for a given heat map, the probability of a joint occurring at each pixel will be predicted. The remaining modules are used as much as possible in the stacked hourglass network, and local and global features are integrated by each hourglass module, which is further understood in subsequent bottom-up and top-down processing phases. The hourglass modules do not share the weight with each other. The filters are all less than or equal to 3 &#x00D7; 3, and the bottleneck limits the total number of parameters per layer, thus reducing the overall memory usage [<xref ref-type="bibr" rid="ref-89">89</xref>].</p>
<p>Li et al. [<xref ref-type="bibr" rid="ref-90">90</xref>] first introduced the multi-stage pose estimation network (MSPN), which adopted ResNet-based global net as a single-stage module and used a cross-stage feature aggregation strategy, that is, two independent information streams are introduced from the downsampling unit and upsampling unit of the previous stage to the downsampling process of the current stage for each scale, and <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolution is added to each stream for feature aggregation to alleviate the problem of information loss during repeated upsampling and downsampling of multi-stage networks.</p>
<p>Furthermore, feature aggregation can be regarded as an extended residual design that helps solve the vanishing gradient. The multi-stage pose estimation network is designed as a multi-branch supervision method from coarse to fine. Different Gaussian kernel sizes are used at different stages, and the closer the stage kernel-size is to the input, the larger the stage kernel-size will be. Multi-scale supervision is introduced to perform intermediate supervision with four different scales at each stage, resulting in a large amount of contextual information at different levels to help localize challenging poses.</p>
<p>Wei et al. [<xref ref-type="bibr" rid="ref-88">88</xref>] introduced the first pose estimation model based on deep learning, which is called the convolutional pose machine (CPM). CPM combines the advantages of a deep convolutional architecture with a pose machine framework consisting of a series of convolutional networks. In other words, the pose machine&#x2019;s prediction and image feature calculation modules are replaced by deep convolutional architecture, which allows the image and context features to be directly learned from the data to represent these networks. The convolutional architecture is fully differentiable, and all stages of the CPM can be trained end-to-end. In this way, the problem of structured prediction in computer vision can be solved without inferring the graphical model. Furthermore, the method of intermediate supervision is also used to solve the gradient disappearance problem in the cascade model training process.</p>
<p>The high-resolution network (HRNet) was proposed by Sun et al. [<xref ref-type="bibr" rid="ref-102">102</xref>], showing superior performance in human body pose estimation. This network will connect sub-networks from high resolution to low resolution in parallel to maintain high-resolution expression. Furthermore, the predicted heatmaps are more accurate by performing repeated multi-scale fusions to obtain high-resolution features with low-resolution representations of the same depth and similar levels.</p>
<p>We have introduced several typical CNN-based posture recognition algorithms above, which all have their own characteristics, and the summary is shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of several improved CNN algorithms</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead valign="top">
<tr>
<th>Improved CNN</th>
<th>Description</th>
<th>Dataset</th>
<th>Performances</th>
<th>Characteristics</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td>Newell et al. [<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
<td>Stacked hourglass networks</td>
<td>FLIC</td>
<td>PCK@0.2: elbow: 99.0%, Wrist: 97%</td>
<td rowspan="2">The bottom-up and top-down structures are repeatedly used in the network architecture, using intermediate supervised learning, and the network converges quickly. The mechanism is simple and can handle diverse and challenging pose sets. Heavy shielding and close contact with multiple people will lead to ambiguity or even overlap.</td>
</tr>
<tr>
<td/>
<td/>
<td>MPII</td>
<td>PCKh@0.5: 90.9% (Total)</td>
</tr>
<tr>
<td>Li et al. [<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
<td>MSPN</td>
<td>COCO</td>
<td>Single model: 76.1 AP, ensemble model: 78.1 AP</td>
<td rowspan="2">Multi-stage pipeline with a single-stage module, supervision from coarse to fine. The Cross-stage feature aggregation strategy is used to reduce the information loss and realize the multi-person posture estimation.</td>
</tr>
<tr>
<td/>
<td/>
<td>MPII</td>
<td>PCKh@0.5: 92.6% (Mean)</td>
</tr>
<tr>
<td>Wei et al. [<xref ref-type="bibr" rid="ref-88">88</xref>]</td>
<td>CPM</td>
<td>MPII</td>
<td>PCKh@0.5: 87.95% (Total)</td>
<td rowspan="2">Predict the long-term dependencies between variables in a structured task with an implicit model. The accuracy of part location is improved. Close multi-person processing and a single end-to-end architecture are less efficient.</td>
</tr>
<tr>
<td/>
<td/>
<td>LSP</td>
<td>PCK: 84.32%</td>
</tr>
<tr>
<td/>
<td/>
<td>FLIC</td>
<td>PCK@0.2:elbow: 97.59%, Wrist: 95.03%</td>
</tr>
<tr>
<td>Sun et al. [<xref ref-type="bibr" rid="ref-102">102</xref>]</td>
<td>HRNet</td>
<td>COCO</td>
<td>HRNet-W48: 75.5 AP, HRNet-W48 &#x002B; extra data: 77.0 AP</td>
<td rowspan="2">The whole process is represented by high resolution, and the multi-resolution representation is repeatedly fused to present a reliable high-resolution representation.</td>
</tr>
<tr>
<td/>
<td/>
<td>MPII</td>
<td>PCKh@0:5: Single-scale testing: 90.3% (Total)<break/>Multi-scale testing: 90.8% (Total)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Lightweight Network</title>
<p>The practice proves that a large number of convolutional neural network models have a significant effect on posture recognition. However, with the increasing complexity of convolutional neural network models, the number of layers of the model will gradually deepen accordingly, resulting in an increasing number of parameters, which will require more computing resources. Moreover, with the support of Internet of Things (IoT) technology and smart terminals, such as mobile phones and embedded devices, there is an increasing demand for porting human posture recognition networks to resource-constrained platforms [<xref ref-type="bibr" rid="ref-104">104</xref>]. Therefore, lightweight research on the convolutional neural network model is gradually carried out. The emerging lightweight network models mainly include Squeeze Net [<xref ref-type="bibr" rid="ref-105">105</xref>], Mobile Net [<xref ref-type="bibr" rid="ref-106">106</xref>], Shuffle Net [<xref ref-type="bibr" rid="ref-107">107</xref>], Xception [<xref ref-type="bibr" rid="ref-108">108</xref>], and Shuffle Net V2 [<xref ref-type="bibr" rid="ref-109">109</xref>].</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Spatial Separable Convolutions</title>
<p>Spatially separable convolution (SSC) mainly refers to splitting or transforming the convolution kernel, then performing convolution calculations separately, which mainly deals with the two spatial dimensions of image width and height and the convolution kernel. A spatially separable convolution splits a kernel into two smaller kernels.</p>
<p>For example, before a <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> convolution core is split, nine times multiplication is required to complete a convolution. After being split into a <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> convolution core, three times multiplication is required for each convolution, and a total of 6 multiplications for the combination of the two convolutions can achieve the same effect as before [<xref ref-type="bibr" rid="ref-110">110</xref>], as shown in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>. The cost of multiplication is reduced, so the computational complexity is reduced, and the network can run faster. It should be noted that not all convolution kernels can be split into two smaller ones.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>An example of spatial separable convolutions</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-11.tif"/>
</fig>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Depthwise Separable Convolution</title>
<p>In depthwise separable convolution (DSC), one convolution kernel can also be split into two small convolution kernels, but different from spatially separable convolution, depthwise separable convolution can be applied to those convolution kernels that cannot be split, and then perform two calculations for these two convolution kernels: depthwise convolution and pointwise convolution, which greatly reduces the amount of computation in the convolution process.</p>
<p>Depthwise convolution is a channel-to-channel convolution operation that establishes a <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula> convolution kernel for each channel of input data. A convolution kernel convolves a channel, and a channel is convolved only by a convolution kernel. In this process, the number of generated feature mapping channels is exactly equal to the number of input channels [<xref ref-type="bibr" rid="ref-111">111</xref>].</p>
<p>Pointwise convolution operations are very similar to regular convolution operations. A <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolution kernel is implemented on every channel completed by depthwise convolution. The size of the pointwise convolution kernel is <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula>, where <italic>L</italic> is the number of channels on the upper layer. The mapping in the previous step is weighted in the depth direction by the convolution operation to generate a new feature map pointwise convolution.</p>
<p>The spatial dimension can be processed by depthwise separable convolution, and the matrix can also be divided by the depth of the convolution kernel. It is to segment the channels of the convolution kernel instead of directly decomposing the matrix.</p>
</sec>
<sec id="s4_4_3">
<label>4.4.3</label>
<title>Feature Pyramid Networks</title>
<p>The feature pyramid network (FPN) is designed according to the concept of a feature pyramid. Instead of the feature extractor of detection models (such as faster R-CNN), FPN generates multi-layer feature maps and pays attention to both the texture features of the shallow network and semantic features of the deep network when extracting features.</p>
<p>FPN includes three parts: bottom-up path, top-down path, and lateral connection [<xref ref-type="bibr" rid="ref-112">112</xref>], as shown in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>. The bottom-up path calculation is a feature hierarchy composed of feature maps of multiple scales, which is the traditional convolutional network to achieve feature extraction. With the deepening of the convolution network, the spatial resolution decreases, and the spatial information is lost, but the semantic value of the network layer increases correspondingly and is more detected.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Schematic diagram of FPN network structure</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-12.tif"/>
</fig>
<p>The top-down path builds higher-resolution layers based on semantically richer layers. These features are then augmented by horizontal connections using the features in the bottom-up path [<xref ref-type="bibr" rid="ref-112">112</xref>]. The feature maps of the same spatial size of the bottom-up path and the top-down path are merged by each horizontal connection.</p>
</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Batch Normalization</title>
<p>For a neural network, the parameters will be continuously updated with the gradient descent, which will cause changes in the data distribution of internal nodes, that is, the internal covariance translation phenomenon. In this case, the above problems can be solved by batch normalization (BN), and the speed of model training and the performance of network generalization can be significantly improved [<xref ref-type="bibr" rid="ref-113">113</xref>].</p>
<p>The main idea of BN is that any layer in the network can be normalized, and the normalized feature graph can be re-scaled and shifted to make the data meet or approximate the Gaussian form of distribution. Batch normalization can reparameterize almost any deep network, addressing the situation where the data distribution in the middle layers changes during training [<xref ref-type="bibr" rid="ref-114">114</xref>]. Like the convolution layer, activation function layer, pooling layer, and fully connected layer, batch normalization is also a network layer. The forward transmission process of the BN network layer is shown in <xref ref-type="disp-formula" rid="eqn-37">Eqs. (37)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-40">(40)</xref> and <xref ref-type="fig" rid="fig-13">Fig. 13</xref>.</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>The forward transmission process of the BN network layer</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-13.tif"/>
</fig>
<p><disp-formula id="eqn-37"><label>(37)</label><mml:math id="mml-eqn-37" display="block"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-38"><label>(38)</label><mml:math id="mml-eqn-38" display="block"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mo>=</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-39"><label>(39)</label><mml:math id="mml-eqn-39" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msqrt><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:msqrt></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-40"><label>(40)</label><mml:math id="mml-eqn-40" display="block"><mml:mi>B</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> refers to mini-batch mean, <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msubsup><mml:mrow><mml:mi mathvariant="normal">&#x03C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> refers to mini-batch variance, <italic>n</italic> refers to the mini-batch size, and <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the normalization process. We define &#x03C4; as a very small value to prevent the denominator from being zero. To maintain the expressiveness of the model, we introduce two learning parameters <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> refers to the scale factor, and <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> refers to the shift factor.</p>
<p>In convolutional neural networks, batch normalization occurs after the convolution computation and before the activation function is applied. If the convolution calculation outputs multiple channels, the outputs of these channels should be batch normalized separately, and each channel has the independent scale and shift parameters, which are all scalars.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Deep Residual Network</title>
<p>In conventional neural networks, the continuous increase of network depth will lead to the gradual increase of accuracy until saturation and then rapid decline, resulting in the difficulty of deep network training, that is, network degradation, which may be caused by the model being too large and the convergence speed too slow. The degradation problem can be solved by the deep residual network (DRN) [<xref ref-type="bibr" rid="ref-115">115</xref>]. The network layer can be made very deep through this residual network structure, and the final classification effect is also very good.</p>
<p>In the residual network structure, for a neural network with a stacked-layer structure, assuming the input is <italic>x</italic>, <italic>H</italic>(<italic>x</italic>) denotes the learned feature, and the residual that can be learned is expected to be denoted as <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi></mml:math></inline-formula>, so the original learned feature obtained as <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>x</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> can be implemented by a feedforward neural network with &#x201C;shortcut connections&#x201D; [<xref ref-type="bibr" rid="ref-115">115</xref>], as shown in <xref ref-type="fig" rid="fig-14">Fig. 14</xref>. When the residual <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is equal to 0, only the identity mapping is completed by the stacking layer, and the goal of the later learning is to approximate the residual result to 0 so that with the deepening of the network, the network performance will not be degraded.</p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>The basic residual block</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-14.tif"/>
</fig>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Dropout Technology</title>
<p>In the deep neural network model, if the number of neural network layers is too large, the training samples are few, or the training time is too long, it will lead to the phenomenon of overfitting [<xref ref-type="bibr" rid="ref-116">116</xref>]. Dropout technology can be used to reduce overfitting to prevent complex co-adaptation to training data [<xref ref-type="bibr" rid="ref-117">117</xref>].</p>
<p>In the neural network using dropout technology, a batch of units is randomly selected and temporarily removed from the network at each iteration in the training stage, keeping these units out of forward inference and backward propagation [<xref ref-type="bibr" rid="ref-116">116</xref>].</p>
<p>It should be noted that instead of simply discarding the outputs of some neural units, we need to change the values of the remaining outputs to ensure that the expectations of the outputs before and after discarding remain unchanged. In general, a fixed probability <italic>p</italic> that each cell retains can be selected using the validation set, and the probability <italic>p</italic> is often set to 0.5. Still, the optimal retention probability is usually closer to 1 for input cells. In networks with dropout, the generalization errors of various classification problems can be significantly reduced using the approximate averaging method.</p>
<p>Suppose we want to train such a neural network, as shown in <xref ref-type="fig" rid="fig-15">Fig. 15a</xref>. After Dropout is applied to the neural network, the training process is mainly the following:
<list list-type="simple">
<list-item><label>(i)</label><p>Randomly delete half of the hidden neurons in the network. Note that these deleted neurons are only temporarily deleted, not permanently deleted, and the input and output neurons remain unchanged, as shown in <xref ref-type="fig" rid="fig-15">Fig. 15b</xref>.</p></list-item>
<list-item><label>(ii)</label><p>The input is then propagated forward along the modified network, and the loss result is propagated back along the modified network. After this procedure was performed on a small group of training samples, the parameters of the neurons that were not deleted were updated according to the random gradient descent (SGD) method.</p></list-item>
<list-item><label>(iii)</label><p>Restore the deleted neuron. At this time, the deleted neuron parameters keep the results before deletion, while the non-deleted neuron parameters have been updated. The above process is repeated continuously.</p></list-item>
</list></p>
<fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>A neural network using dropout technology</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-15.tif"/>
</fig>
</sec>
<sec id="s4_8">
<label>4.8</label>
<title>Advanced Activation Functions</title>
<p>In a neural network, an important purpose of using multi-layer convolution is to use the size of different convolution kernels to extract image features at different convolution kernel scales. The convolution algorithm is composed of a mass of multiplications and additions, so the convolution algorithm is also linear and can be considered a linear weighting operation through the convolution kernel. The convolutional neural network composed of many convolution algorithms will degenerate into a simple linear model without introducing nonlinear factors, making the multi-layer convolution meaningless.</p>
<p>Therefore, adding a nonlinear function after the convolution of each layer of the neural network can complete the linear isolation of the two convolution layers and ensure that each convolution layer completes its own convolution task. Currently, the common activation functions mainly include sigmoid, tanh, rectified linear unit (ReLU), etc. Compared with the traditional activation functions of neural networks, such as sigmoid and tanh, RELU has the following advantages: (i) When the input of the ReLU function is positive, the gradient saturation will not occur in the network. (ii) Since the ReLU function has only a linear relationship, its calculation speed is faster than sigmoid and tanh. The definition of the ReLU function is shown in <xref ref-type="disp-formula" rid="eqn-41">Eq. (41)</xref> [<xref ref-type="bibr" rid="ref-118">118</xref>]:</p>
<p><disp-formula id="eqn-41"><label>(41)</label><mml:math id="mml-eqn-41" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the input in the <italic>i</italic>th channel. There are many variants of the ReLU function, such as parametric ReLU, leaky ReLU, random ReLU, etc. Each activation function has advantages in one or several specific deep learning networks.</p>
<p>Leaky ReLU (LReLU) is similar to ReLU, except that the input is less than 0. In the ReLU function, all negative values are zero, and the outputs are non-negative. In contrast, in the Leaky ReLU, all negative values are assigned a non-zero slope with a negative value and a small gradient [<xref ref-type="bibr" rid="ref-119">119</xref>]. The Leaky ReLU activation function can avoid zero gradients, which is defined as follows:</p>
<p><disp-formula id="eqn-42"><label>(42)</label><mml:math id="mml-eqn-42" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">L</mml:mi><mml:mi mathvariant="italic">R</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">L</mml:mi><mml:mi mathvariant="italic">U</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a fixed parameter, usually with a value of 0.01. In the process of backpropagation, the gradient can also be calculated for the part of the Leaky ReLU activation function input less than zero, which can avoid the problem of gradient direction aliasing.</p>
<p>Parametric ReLU (PReLU) adaptively learns to rectify the parameters of linear units and is able to improve classification accuracy at a negligible extra computational cost [<xref ref-type="bibr" rid="ref-120">120</xref>]. The definition of the PReLU function is shown in <xref ref-type="disp-formula" rid="eqn-43">Eq. (43)</xref>:</p>
<p><disp-formula id="eqn-43"><label>(43)</label><mml:math id="mml-eqn-43" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">P</mml:mi><mml:mi mathvariant="italic">R</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">L</mml:mi><mml:mi mathvariant="italic">U</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is responsible for controlling the slope of the negative semi-axis, and the activation functions of different channels can be different. When the value of <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is 0, PReLU can be regarded as ReLU. If the value of <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is small and fixed, then PReLU can be considered Leaky ReLU.</p>
<p>Randomized ReLU (RReLU) can be understood as a variant of Leaky ReLU. The definition of RReLU function is shown as follows:</p>
<p><disp-formula id="eqn-44"><label>(44)</label><mml:math id="mml-eqn-44" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">R</mml:mi><mml:mi mathvariant="italic">R</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">L</mml:mi><mml:mi mathvariant="italic">U</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the input of the <italic>i</italic>th channel in the <italic>j</italic>th example, <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a random value drawn from a uniform distribution <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>u</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p><disp-formula id="eqn-45"><label>(45)</label><mml:math id="mml-eqn-45" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>u</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>u</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>and</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>u</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The diagrams of ReLU, LReLU, PReLU, and RReLU are shown in the following <xref ref-type="fig" rid="fig-16">Fig. 16</xref> [<xref ref-type="bibr" rid="ref-121">121</xref>].</p>
<fig id="fig-16">
<label>Figure 16</label>
<caption>
<title>The diagrams of ReLU, LReLU, PReLU, and RReLU</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-16.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Advanced Neural Networks</title>
<p>In order to improve the performance of the system, some advanced neural networks are studied in the field of posture recognition, such as transfer learning, ensemble learning, graph neural networks, explainable deep neural networks, etc.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Transfer Learning</title>
<p>Transfer learning (TL) refers to the transfer of the trained model parameters to the new model to help the new model training [<xref ref-type="bibr" rid="ref-122">122</xref>]. Transfer learning technology has been used in posture recognition. Hu et al. [<xref ref-type="bibr" rid="ref-123">123</xref>] used transfer learning in their sleep posture system, and the system accuracy and real-time processing speed were much higher than the standard training-test method. Ogundokun et al. [<xref ref-type="bibr" rid="ref-124">124</xref>] applied the transfer learning algorithm with hyperparameter optimization (HPO) to human posture detection. The experiments show that the algorithm is superior to the algorithm using image enhancement in terms of training loss and verification accuracy, but the system&#x2019;s complexity increased after the algorithm was used. Long et al. [<xref ref-type="bibr" rid="ref-125">125</xref>] developed a yoga self-training system using transfer learning techniques.</p>
<p>Considering that most data or tasks are related, through transfer learning, we can share the learned model parameters with the new model in some way to speed up and optimize the learning efficiency of the model. It is one of the advantages of transfer learning that we do not need to learn from zero like most networks. In addition, in the case of small data sets, transfer learning can get good results, and we can also use transfer learning to reduce training cost sets.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Ensemble Learning</title>
<p>Ensemble learning (EL) is to construct and combine multiple machine learning machines to complete learning tasks. The process generates a group of &#x201C;individual learning machines&#x201D; and then combines them with a certain strategy [<xref ref-type="bibr" rid="ref-126">126</xref>]. Individual learning machines are common machine learning algorithms, such as decision trees and neural networks [<xref ref-type="bibr" rid="ref-127">127</xref>]. Ensemble learning can be used for classification problem integration, regression problem integration, feature selection integration, outlier detection integration, and so on.</p>
<p>Ensemble learning is used in sensor-based posture recognition systems to overcome the problems of data imbalance, instant recognition, sensor deployment, and selection when collecting data with wearable devices [<xref ref-type="bibr" rid="ref-128">128</xref>]. Liang et al. [<xref ref-type="bibr" rid="ref-129">129</xref>] designed a sitting posture recognition system using an ensemble learning classification model to ensure the generalization ability of the system. Esmaeili et al. [<xref ref-type="bibr" rid="ref-130">130</xref>] designed a posture recognition integrated model by superimposing two classification layers based on the deep convolution method.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Graph Neural Networks</title>
<p>Graph neural networks (GNN) is a framework that uses deep learning to learn the graph structure data directly. Aggregating features of adjacent nodes calculate the features of each node, and the graph dependency is established by passing messages between nodes [<xref ref-type="bibr" rid="ref-131">131</xref>]. In GNN, graph properties (such as points, edges, and global information) are transformed without changing the connectivity of the graph.</p>
<p>GNN has achieved excellent results in posture recognition tasks. Guo [<xref ref-type="bibr" rid="ref-132">132</xref>] formed a multi-person posture estimation algorithm based on a graph neural network by using multilevel feature maps, which greatly improved the positioning accuracy of each part of the human body. Li et al. [<xref ref-type="bibr" rid="ref-133">133</xref>] used the graph neural network to optimize posture graphs, which achieves good efficiency and robustness. Taiana et al. [<xref ref-type="bibr" rid="ref-134">134</xref>] constructed a system based on graph neural networks that can produce accurate relative poses.</p>
<p>In recent years, the variants of GNN variants, such as graph convolutional networks (GCNs), graph attention networks (GATs), and gated graph neural networks (GGNNs), have shown breakthrough performance in many posture recognition tasks [<xref ref-type="bibr" rid="ref-135">135</xref>&#x2013;<xref ref-type="bibr" rid="ref-137">137</xref>].</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Analysis and Discussion</title>
<p>We have detailedly reviewed the techniques and methods of posture recognition, including the process of posture recognition, feature extraction, and classification techniques. Compared with the existing reviews in recent years, this paper presents the following advantages: (i) This paper combs the pose recognition technologies and methods based on traditional machine learning and deep learning-based posture recognition technologies and methods and summarizes and analyzes 2D and 3D datasets, which is more comprehensive in content; (ii) In order to timely share the latest technologies and methods of posture recognition with readers, this review focuses on the latest development of posture recognition technologies and methods. The literature on posture recognition technologies and methods is relatively new, and most of them are related research papers from the past five years, which have been updated in time. We list the comparison of recent reviews on posture recognition as shown in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>The comparison of recent reviews on posture recognition</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>References</th>
<th>Year</th>
<th>Focus</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-138">138</xref>]</td>
<td>2020</td>
<td>Monocular 3D human pose estimation</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-139">139</xref>]</td>
<td>2021</td>
<td>Monocular multi-person pose estimation</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-140">140</xref>]</td>
<td>2021</td>
<td>3D human pose estimation algorithms for markerless motion capture</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-141">141</xref>]</td>
<td>2021</td>
<td>2D multi-person pose estimation methods</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-142">142</xref>]</td>
<td>2021</td>
<td>Deep 3D human pose estimation</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-143">143</xref>]</td>
<td>2021</td>
<td>Human pose estimation and its application to action recognition</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-144">144</xref>]</td>
<td>2022</td>
<td>The application of hardware technology in the posture recognition system</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To help understand more clearly, we created a table of abbreviations and corresponding full names for posture recognition terms as follows (<xref ref-type="table" rid="table-3">Table 3</xref>):</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The list of all abbreviations and full names</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Abbreviations</th>
<th>Full names</th>
</tr>
</thead>
<tbody>
<tr>
<td>AI</td>
<td>Artificial intelligence</td>
</tr>
<tr>
<td>ANN</td>
<td>Artificial neural network</td>
</tr>
<tr>
<td>AP</td>
<td>Average precision</td>
</tr>
<tr>
<td>BN</td>
<td>Batch normalization</td>
</tr>
<tr>
<td>CACT</td>
<td>Cascade-adaboosting-CART</td>
</tr>
<tr>
<td>CFN</td>
<td>Coarse-fine network</td>
</tr>
<tr>
<td>CL</td>
<td>Convolutional layer</td>
</tr>
<tr>
<td>CNN</td>
<td>Convolutional neural network</td>
</tr>
<tr>
<td>CPM</td>
<td>Convolutional pose machine</td>
</tr>
<tr>
<td>CPN</td>
<td>Cascaded pyramid network</td>
</tr>
<tr>
<td>CRF</td>
<td>Conditional random field</td>
</tr>
<tr>
<td>DA</td>
<td>Data augmentation</td>
</tr>
<tr>
<td>DNN</td>
<td>Deep neural network</td>
</tr>
<tr>
<td>DoG</td>
<td>Difference of Gaussian</td>
</tr>
<tr>
<td>DRN</td>
<td>Deep residual network</td>
</tr>
<tr>
<td>DSC</td>
<td>Depthwise separable convolution</td>
</tr>
<tr>
<td>DT</td>
<td>Decision tree</td>
</tr>
<tr>
<td>DTW</td>
<td>Dynamic time warping</td>
</tr>
<tr>
<td>DWT</td>
<td>Discrete wavelet transform</td>
</tr>
<tr>
<td>ECG</td>
<td>Electrocardiogram</td>
</tr>
<tr>
<td>EEG</td>
<td>Electroencephalogram</td>
</tr>
<tr>
<td>EL</td>
<td>Ensemble learning</td>
</tr>
<tr>
<td>EMG</td>
<td>Electromyogram</td>
</tr>
<tr>
<td>FCL</td>
<td>Fully connected layer</td>
</tr>
<tr>
<td>FD</td>
<td>Fourier descriptor</td>
</tr>
<tr>
<td>FPN</td>
<td>Feature pyramid network</td>
</tr>
<tr>
<td>GAN</td>
<td>Generative adversarial network</td>
</tr>
<tr>
<td>GATs</td>
<td>Graph attention networks</td>
</tr>
<tr>
<td>GCNs</td>
<td>Graph convolutional networks</td>
</tr>
<tr>
<td>GGNNs</td>
<td>Gated graph neural networks</td>
</tr>
<tr>
<td>GLCM</td>
<td>Gray-level co-occurrence matrix</td>
</tr>
<tr>
<td>GMM</td>
<td>Gaussian mixture model</td>
</tr>
<tr>
<td>GNN</td>
<td>Graph neural networks</td>
</tr>
<tr>
<td>GSR</td>
<td>Galvanic skin response</td>
</tr>
<tr>
<td>HAR</td>
<td>Human activity recognition</td>
</tr>
<tr>
<td>HMI</td>
<td>Hu moment invariant</td>
</tr>
<tr>
<td>HMM</td>
<td>Hidden Markov model</td>
</tr>
<tr>
<td>HMR</td>
<td>Human mesh recovery</td>
</tr>
<tr>
<td>HOD</td>
<td>Histogram of oriented displacement</td>
</tr>
<tr>
<td>HPO</td>
<td>Hyperparameter optimization</td>
</tr>
<tr>
<td>HOG</td>
<td>Histogram of oriented gradients</td>
</tr>
<tr>
<td>HRNet</td>
<td>High-resolution net</td>
</tr>
<tr>
<td>HSV</td>
<td>Hue saturation value</td>
</tr>
<tr>
<td>IEF</td>
<td>Iterative error feedback</td>
</tr>
<tr>
<td>IMU</td>
<td>Inertial measurement unit</td>
</tr>
<tr>
<td>IoT</td>
<td>Internet of things</td>
</tr>
<tr>
<td>k-NN</td>
<td>k-nearest neighbor</td>
</tr>
<tr>
<td>LDA</td>
<td>Linear discriminant analysis</td>
</tr>
<tr>
<td>LMC</td>
<td>Leap motion controller</td>
</tr>
<tr>
<td>LoG</td>
<td>Laplacian of Gaussian</td>
</tr>
<tr>
<td>LReLU</td>
<td>Leaky rectified linear unit</td>
</tr>
<tr>
<td>LSTM</td>
<td>Long short-term memory</td>
</tr>
<tr>
<td>MPJPE</td>
<td>Mean per joint position error</td>
</tr>
<tr>
<td>MPR</td>
<td>Multi-pose recognition</td>
</tr>
<tr>
<td>MSPN</td>
<td>Multi-stage pose estimation network</td>
</tr>
<tr>
<td>MSST-ResNet</td>
<td>Multi-scale spatio-temporal residual network</td>
</tr>
<tr>
<td>NBC</td>
<td>Naive Bayes classifier</td>
</tr>
<tr>
<td>NWFE</td>
<td>Nonparametric weighted feature extraction</td>
</tr>
<tr>
<td>PCA</td>
<td>Principal component analysis</td>
</tr>
<tr>
<td>PL</td>
<td>Pooling layer</td>
</tr>
<tr>
<td>PReLU</td>
<td>Parametric rectified linear unit</td>
</tr>
<tr>
<td>PRN</td>
<td>Pose residual network</td>
</tr>
<tr>
<td>R-CNN</td>
<td>Region-CNN</td>
</tr>
<tr>
<td>ReLU</td>
<td>Rectified linear unit</td>
</tr>
<tr>
<td>ResNet</td>
<td>Residual neural network</td>
</tr>
<tr>
<td>RF</td>
<td>Random forest</td>
</tr>
<tr>
<td>RFID</td>
<td>Radio frequency identification</td>
</tr>
<tr>
<td>RMPE</td>
<td>Regional multi-person pose estimation</td>
</tr>
<tr>
<td>RNN</td>
<td>Recurrent neural network</td>
</tr>
<tr>
<td>RReLU</td>
<td>Randomized leaky rectified linear unit</td>
</tr>
<tr>
<td>SCGA</td>
<td>Squeezed convolutional gated attention</td>
</tr>
<tr>
<td>SGD</td>
<td>Stochastic gradient descent</td>
</tr>
<tr>
<td>SIFT</td>
<td>Scale-invariant feature transform</td>
</tr>
<tr>
<td>SSC</td>
<td>Spatially separable convolution</td>
</tr>
<tr>
<td>SVM</td>
<td>Support vector machine</td>
</tr>
<tr>
<td>TL</td>
<td>Transfer learning</td>
</tr>
<tr>
<td>VGG</td>
<td>Visual geometry group</td>
</tr>
<tr>
<td>VHMM</td>
<td>Validation hidden Markov model</td>
</tr>
<tr>
<td>WE</td>
<td>Wavelet entropy</td>
</tr>
<tr>
<td>WVS</td>
<td>Wireless visual sensor</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s6_1">
<label>6.1</label>
<title>Main Recognition Techniques</title>
<p>According to data acquisition, posture recognition technology is divided into sensor-based recognition technology, vision-based recognition technology, and RF-based recognition technology.</p>
<p>Sensor-based recognition methods are less costly and simple to operate but are limited to devices and require the real-time wearing of sensors [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-145">145</xref>].</p>
<p>Vision-based recognition method has high accuracy and overcomes the problem of wearing. It is easy to obtain the trajectory, contour, and other information about human movement. However, this method is affected by light, background environment, and other factors and is prone to recognition errors due to occlusion and privacy exposure [<xref ref-type="bibr" rid="ref-146">146</xref>,<xref ref-type="bibr" rid="ref-147">147</xref>].</p>
<p>RF-based identification technology has the characteristics of non-contact and is very sensitive to environmental changes. It is easily affected by the human body&#x2019;s absorption, reflection, and scattering of RF signals [<xref ref-type="bibr" rid="ref-31">31</xref>]. The characteristics of the three recognition methods are shown in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Summary of main recognition techniques for posture</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Technology</th>
<th>Advantages</th>
<th>Disadvantages</th>
</tr>
</thead>
<tbody>
<tr>
<td>Sensor&#x2011;based</td>
<td>Smartphone, accelerometer, gyroscope</td>
<td>Low cost</td>
<td>Constrains of carrying device</td>
</tr>
<tr>
<td>Vision&#x2011;based</td>
<td>Camera</td>
<td>High accuracy</td>
<td>High cost, complex computation, privacy issue</td>
</tr>
<tr>
<td>RF-based</td>
<td>Wi-Fi</td>
<td>Cost-effective, Widely available</td>
<td>Environmental disturbance, unable to provide fine-grained recognition</td>
</tr>
<tr>
<td/>
<td>RFID</td>
<td>Cost-effective, Widely available</td>
<td>Environmental disturbance</td>
</tr>
<tr>
<td/>
<td>Radar</td>
<td>Widely available</td>
<td>Environmental disturbance, unable to provide fine-grained recognition</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>2D Posture Recognition and 3D Posture Recognition</title>
<p>According to the difference in human posture dimensions, the human posture recognition task can be divided into two-dimensional human posture recognition and three-dimensional human posture recognition. The purpose of two-dimensional human posture recognition is to locate and identify the keypoints of the human body. Then these key points are connected in the order of joints, which are projected on the two-dimensional plane of the image to form the human skeleton.</p>
<p>There are currently many 2D recognition algorithms, and the accuracy and processing speed have been greatly improved. However, the keypoints of 2D are greatly affected by wearing, posture and perspective. They are also affected by the environment, such as occlusion, illumination, and fog, which require high requirements for data annotation. In addition, the keypoints of 2D are not easy to estimate the positions between human body parts through vision.</p>
<p>3D posture recognition can give images a more stable and understandable interpretation. In recognition of human 3D posture, the 3D coordinate position and angle of human joints are mainly predicted. We can use the 3D posture estimator to convert objects in the image into 3D objects by adding depth to the prediction, that is, to realize the mapping between 2D keypoints and 3D keypoints. There are two specific methods: One is to directly regress 3D coordinates from 2D images [<xref ref-type="bibr" rid="ref-148">148</xref>,<xref ref-type="bibr" rid="ref-149">149</xref>], and the other is to obtain the data of 2D first and then &#x201C;lift&#x201D; to 3D posture [<xref ref-type="bibr" rid="ref-150">150</xref>,<xref ref-type="bibr" rid="ref-151">151</xref>].</p>
<p>In 3D posture recognition, due to the addition of depth information on the basis of 2D posture recognition, the expression of human posture is more accurate than in 2D, but there will be occlusion, and it also faces challenges such as the inherent deep ambiguity and inadequacy in single-view 2D to 3D mapping, and the lack of large outdoor datasets. Currently, the mainstream datasets are established in the laboratory environment, and the model&#x2019;s generalization ability is weak. In addition, there is a lack of special posture datasets, such as falling and rolling.</p>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Recognition Based on Traditional Machine Learning and Deep Neural Network</title>
<p>Traditional machine learning-based recognition methods mainly describe and infer human posture based on the human body models and extract image posture features through algorithms, which have high requirements on feature representation and spatial position relationship of keypoints. Excluding low level features (such as boundary and color), typical high-level features, such as scale-invariant feature transformation and gradient histogram, have stronger expression ability and can effectively compress the spatial dimension of features, showing advantages in terms of time efficiency.</p>
<p>Posture recognition based on deep learning can be trained and learned through the image data of the network model, and the most effective representation method can be directly obtained. The core of posture recognition based on deep learning is the depth of neural networks. Semantic information is extracted from the image through a convolutional neural network, richer and more accurate and reflects better robustness than artificial features.</p>
<p>Moreover, the expressive ability of the network model will increase exponentially with the increase of the network stack number. However, overcoming factors such as occlusion, inadequate training data, and depth blur is still difficult. The commonly used posture recognition algorithms [<xref ref-type="bibr" rid="ref-143">143</xref>,<xref ref-type="bibr" rid="ref-152">152</xref>] in recent years are shown in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Common algorithms for posture recognition research</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>References</th>
<th>Method</th>
<th>Year</th>
<th>Datasets</th>
<th>Accuracy/performance</th>
<th>Characteristics</th>
</tr>
</thead>
<tbody>
<tr>
<td>Pishchulin et al. [<xref ref-type="bibr" rid="ref-153">153</xref>]</td>
<td>DeepCut</td>
<td>2016</td>
<td>MPII</td>
<td>54.10% (pckh-0.5)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Pishchulin et al. [<xref ref-type="bibr" rid="ref-153">153</xref>]</td>
<td>DeeperCut</td>
<td>2016</td>
<td>MPII</td>
<td>59.40% (pckh-0.5)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Wei et al. [<xref ref-type="bibr" rid="ref-88">88</xref>]</td>
<td>CPM</td>
<td>2016</td>
<td>MPII</td>
<td>87.95% (pckh-0.5)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Newell et al. [<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
<td>Stacked hourglass networks</td>
<td>2016</td>
<td>MPII</td>
<td>90.90% (pckh-0.5)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Carreira et al. [<xref ref-type="bibr" rid="ref-154">154</xref>]</td>
<td>IEF</td>
<td>2016</td>
<td>MPII</td>
<td>81.3% (pckh-0.5)</td>
<td>Single-person</td>
</tr>
<tr>
<td>Fang et al. [<xref ref-type="bibr" rid="ref-155">155</xref>]</td>
<td>RMPE</td>
<td>2017</td>
<td>MS COCO</td>
<td>61.80% (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>He et al. [<xref ref-type="bibr" rid="ref-156">156</xref>]</td>
<td>Mask R-CNN</td>
<td>2017</td>
<td>MS COCO</td>
<td>63.10% (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Newell et al. [<xref ref-type="bibr" rid="ref-157">157</xref>]</td>
<td>Associative embedding</td>
<td>2017</td>
<td>MS COCO</td>
<td>65.50% (AP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Huang et al. [<xref ref-type="bibr" rid="ref-158">158</xref>]</td>
<td>CFN</td>
<td>2017</td>
<td>MS COCO</td>
<td>72.60% (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Newell et al. [<xref ref-type="bibr" rid="ref-157">157</xref>]</td>
<td>Associative embedding</td>
<td>2017</td>
<td>MPII</td>
<td>77.50% (mAP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Fang et al. [<xref ref-type="bibr" rid="ref-155">155</xref>]</td>
<td>RMPE</td>
<td>2017</td>
<td>MPII</td>
<td>82.10% (pckh-0.5)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Chu et al. [<xref ref-type="bibr" rid="ref-159">159</xref>]</td>
<td>CRF</td>
<td>2017</td>
<td>MPII</td>
<td>91.50% (pckh-0.5)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Fang et al. [<xref ref-type="bibr" rid="ref-155">155</xref>]</td>
<td>AlphaPose</td>
<td>2017</td>
<td>MPII</td>
<td>76.7% (mAP-0.5)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Fang et al. [<xref ref-type="bibr" rid="ref-155">155</xref>]</td>
<td>AlphaPose</td>
<td>2017</td>
<td>MS COCO</td>
<td>71.0 (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Kocabas et al. [<xref ref-type="bibr" rid="ref-160">160</xref>]</td>
<td>PRN</td>
<td>2018</td>
<td>MS COCO</td>
<td>69.60% (AP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Chen et al. [<xref ref-type="bibr" rid="ref-91">91</xref>]</td>
<td>CPN</td>
<td>2018</td>
<td>MS COCO</td>
<td>73.00% (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Xiao et al. [<xref ref-type="bibr" rid="ref-92">92</xref>]</td>
<td>Simple baseline</td>
<td>2018</td>
<td>MS COCO</td>
<td>73.70% (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Kanazawa et al. [<xref ref-type="bibr" rid="ref-161">161</xref>]</td>
<td>HMR</td>
<td>2018</td>
<td>Human3.6M</td>
<td>56.80 mm (average MPJPE)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Kocabas et al. [<xref ref-type="bibr" rid="ref-160">160</xref>]</td>
<td>MultiPoseNet</td>
<td>2018</td>
<td>MS COCO</td>
<td>70.5 (AP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Kreiss et al. [<xref ref-type="bibr" rid="ref-162">162</xref>]</td>
<td>PifPaf</td>
<td>2019</td>
<td>MS COCO</td>
<td>66.70% (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Li et al. [<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
<td>MSPN</td>
<td>2019</td>
<td>MS COCO</td>
<td>76.10% (AP)</td>
<td>Top-down</td>
</tr>
<tr>
<td>Sun et al. [<xref ref-type="bibr" rid="ref-102">102</xref>]</td>
<td>HRNet-W48</td>
<td>2019</td>
<td>MS COCO</td>
<td>77.00% (AP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Sun et al. [<xref ref-type="bibr" rid="ref-102">102</xref>]</td>
<td>HRNet-W48</td>
<td>2019</td>
<td>MPII</td>
<td>90.80% (pckh-0.5)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Xu et al. [<xref ref-type="bibr" rid="ref-163">163</xref>]</td>
<td>DenseRaC</td>
<td>2019</td>
<td>Human3.6M</td>
<td>48.00 mm (average MPJPE)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Zhao et al. [<xref ref-type="bibr" rid="ref-137">137</xref>]</td>
<td>SemGCN</td>
<td>2019</td>
<td>Human3.6M</td>
<td>43.80 mm (average MPJPE)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Gujjar et al. [<xref ref-type="bibr" rid="ref-164">164</xref>]</td>
<td>Res-EnDec</td>
<td>2019</td>
<td>JAAD</td>
<td>81.14% (AP)</td>
<td>deep learning</td>
</tr>
<tr>
<td>Huang et al. [<xref ref-type="bibr" rid="ref-165">165</xref>]</td>
<td>DeepFuse</td>
<td>2020</td>
<td>Human3.6M</td>
<td>37.50 mm (average MPJPE)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Zhong et al. [<xref ref-type="bibr" rid="ref-166">166</xref>]</td>
<td>SocialGAN</td>
<td>2020</td>
<td>3D Pedstria Trajectory</td>
<td>71.60% (prediction error)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Cao et al. [<xref ref-type="bibr" rid="ref-167">167</xref>]</td>
<td>OpenPose</td>
<td>2021</td>
<td>MS COCO</td>
<td>60.50% (AP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Liu et al. [<xref ref-type="bibr" rid="ref-168">168</xref>]</td>
<td>UDP-Pose-PSA</td>
<td>2021</td>
<td>MS COCO</td>
<td>79.50% (AP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Cao et al. [<xref ref-type="bibr" rid="ref-167">167</xref>]</td>
<td>OpenPose</td>
<td>2021</td>
<td>MPII</td>
<td>76.50% (AP)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Groos et al. [<xref ref-type="bibr" rid="ref-169">169</xref>]</td>
<td>EfficientPose IV</td>
<td>2021</td>
<td>MPII</td>
<td>91.20% (pckh-0.5)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Shan et al. [<xref ref-type="bibr" rid="ref-170">170</xref>]</td>
<td>Pose3D-RIE</td>
<td>2021</td>
<td>Human3.6M</td>
<td>30.10 mm (average MPJPE)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Reddy et al. [<xref ref-type="bibr" rid="ref-171">171</xref>]</td>
<td>TesseTrack</td>
<td>2021</td>
<td>Human3.6M</td>
<td>18.70 mm (average MPJPE)</td>
<td>Bottom-up</td>
</tr>
<tr>
<td>Yau et al. [<xref ref-type="bibr" rid="ref-172">172</xref>]</td>
<td>Graph-SIM</td>
<td>2021</td>
<td>PePScenes</td>
<td>94.40% (accuracy)</td>
<td>Deep learning</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6_4">
<label>6.4</label>
<title>Datasets</title>
<p>In the field of posture recognition, the successful application of deep learning has significantly improved the accuracy and generalization ability of two-dimensional posture recognition, where the datasets play a crucial role in the system [<xref ref-type="bibr" rid="ref-143">143</xref>]. We list widely used 2D posture benchmark datasets, as shown in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Datasets for 2D human posture recognition (I &#x003D; Image, V &#x003D; Video, S &#x003D; Single-person, M &#x003D; Multi-person)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Year</th>
<th>Data source</th>
<th>Single/<break/>Multi person</th>
<th>#Keypoints</th>
<th>#Train</th>
<th>#Test</th>
</tr>
</thead>
<tbody>
<tr>
<td>LSP [<xref ref-type="bibr" rid="ref-173">173</xref>]</td>
<td>2010</td>
<td>I</td>
<td>S</td>
<td>14</td>
<td>1,000 images</td>
<td>1,000 images</td>
</tr>
<tr>
<td>LSP extended [<xref ref-type="bibr" rid="ref-174">174</xref>]</td>
<td>2011</td>
<td>I</td>
<td>S</td>
<td>14</td>
<td>10,000 images</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>FashionPose [<xref ref-type="bibr" rid="ref-175">175</xref>]</td>
<td>2013</td>
<td>I</td>
<td>S</td>
<td>13</td>
<td>6,530 images</td>
<td>1,000 images</td>
</tr>
<tr>
<td>J-HMDB [<xref ref-type="bibr" rid="ref-176">176</xref>]</td>
<td>2013</td>
<td>V</td>
<td>S</td>
<td>13</td>
<td>31,838 frames</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>FLIC [<xref ref-type="bibr" rid="ref-177">177</xref>]</td>
<td>2013</td>
<td>I</td>
<td>S</td>
<td>10</td>
<td>3,987 images</td>
<td>1,016 images</td>
</tr>
<tr>
<td>Penn Action [<xref ref-type="bibr" rid="ref-178">178</xref>]</td>
<td>2013</td>
<td>V</td>
<td>S</td>
<td>13</td>
<td>1,163 videos</td>
<td>1,163 videos</td>
</tr>
<tr>
<td>MPII [<xref ref-type="bibr" rid="ref-179">179</xref>]</td>
<td>2014</td>
<td>I</td>
<td>S</td>
<td>16</td>
<td>28,821 images (40,522 people)</td>
<td>11,701 images</td>
</tr>
<tr>
<td>MPII (Multi-person) [<xref ref-type="bibr" rid="ref-179">179</xref>]</td>
<td>2014</td>
<td>I</td>
<td>M</td>
<td>16</td>
<td>3,844 images</td>
<td>1,758 images</td>
</tr>
<tr>
<td>MSCOCO Keypoints [<xref ref-type="bibr" rid="ref-180">180</xref>]</td>
<td>2014</td>
<td>I</td>
<td>M</td>
<td>17</td>
<td>64,115 images<break/>(262,465 people)</td>
<td>40,670 images (test-std)<break/>20,288 images (test-dev)</td>
</tr>
<tr>
<td>AI challenger [<xref ref-type="bibr" rid="ref-181">181</xref>]</td>
<td>2017</td>
<td>I</td>
<td>M</td>
<td>14</td>
<td>210,000 images</td>
<td>30,000 images</td>
</tr>
<tr>
<td>PoseTrack [<xref ref-type="bibr" rid="ref-182">182</xref>]</td>
<td>2017</td>
<td>V</td>
<td>M</td>
<td>14</td>
<td>20 videos</td>
<td>20 videos</td>
</tr>
<tr>
<td>PoseTrack [<xref ref-type="bibr" rid="ref-183">183</xref>]</td>
<td>2018</td>
<td>V</td>
<td>M</td>
<td>15</td>
<td>292 videos</td>
<td>208 videos</td>
</tr>
<tr>
<td>CrowdPose [<xref ref-type="bibr" rid="ref-184">184</xref>]</td>
<td>2019</td>
<td>I</td>
<td>M</td>
<td>14</td>
<td>10,000 images</td>
<td>8,000 images</td>
</tr>
<tr>
<td>Human-in-Events (HiEve) [<xref ref-type="bibr" rid="ref-185">185</xref>]</td>
<td>2020</td>
<td>V</td>
<td>M</td>
<td>14</td>
<td>49,820 frames<break/>(1,099,357 people)</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Compared with 2D posture recognition, 3D posture recognition faces more challenges, among which deep learning algorithms rely on huge training data. However, due to the difficulty and high cost of 3D posture labeling, the current mainstream datasets are collected in the laboratory environment and lack large outdoor datasets. This will inevitably affect the generalization performance of the algorithm on outdoor data [<xref ref-type="bibr" rid="ref-138">138</xref>,<xref ref-type="bibr" rid="ref-143">143</xref>]. The widely used 3D posture recognition datasets are shown in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Datasets for 3D human posture recognition (S &#x003D; Single-person, M &#x003D; Multi-person)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead valign="top">
<tr>
<th>Dataset</th>
<th>Year</th>
<th>#Frame<break/>#Video Sequence</th>
<th>Size/Characters</th>
<th>Single/<break/>Multi-person</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td>HumanEva-I&#x0026;II [<xref ref-type="bibr" rid="ref-186">186</xref>]</td>
<td>2010</td>
<td>80,000<break/>56</td>
<td>4 subjects, lab environment</td>
<td>S</td>
</tr>
<tr>
<td>Human3.6M [<xref ref-type="bibr" rid="ref-187">187</xref>]</td>
<td>2014</td>
<td>3.6 millions<break/>1 376</td>
<td>About 3.6 &#x00D7; 106 poses, lab environment</td>
<td>S</td>
</tr>
<tr>
<td>CMU Panoptic [<xref ref-type="bibr" rid="ref-188">188</xref>]</td>
<td>2015</td>
<td>1.5 million<break/>65</td>
<td>Large scale, multiple perspectives, multiple people</td>
<td>M</td>
</tr>
<tr>
<td>Joint Track Auto (JTA) [<xref ref-type="bibr" rid="ref-189">189</xref>]</td>
<td>2016</td>
<td>460,800<break/>512</td>
<td>Contains high-definition videos of pedestrians walking in urban scenes</td>
<td>M</td>
</tr>
<tr>
<td>MPI-INF-3DHP [<xref ref-type="bibr" rid="ref-190">190</xref>]</td>
<td>2017</td>
<td>1.3 millions<break/>64</td>
<td>8 subjects, indoor &#x0026; outdoor</td>
<td>S</td>
</tr>
<tr>
<td>SURREAL [<xref ref-type="bibr" rid="ref-191">191</xref>]</td>
<td>2017</td>
<td>6 million</td>
<td>The texture SMPL model on the background image is rendered to form a large composite dataset</td>
<td>S</td>
</tr>
<tr>
<td>MuCo-3DHP [<xref ref-type="bibr" rid="ref-192">192</xref>]</td>
<td>2018</td>
<td>&#x2013;</td>
<td>Datasets were synthesized from MPI-INF-3DHP by data augmentation</td>
<td>M</td>
</tr>
<tr>
<td>3DPW [<xref ref-type="bibr" rid="ref-193">193</xref>]</td>
<td>2018</td>
<td>&#x003E;50,000<break/>60</td>
<td>Collect 3D human poses in the field with IMUs and a moving camera</td>
<td>M</td>
</tr>
<tr>
<td>MuPoTS-3D [<xref ref-type="bibr" rid="ref-192">192</xref>]</td>
<td>2018</td>
<td>8,000<break/>20</td>
<td>Test set of 3D posture estimation for multiple people in the wild</td>
<td>M</td>
</tr>
<tr>
<td>AMASS [<xref ref-type="bibr" rid="ref-194">194</xref>]</td>
<td>2019</td>
<td>N/A<break/>(&#x003E;40 h)</td>
<td>Fifteen different marker-based MoCap datasets were unified into 3D human meshes</td>
<td>S</td>
</tr>
<tr>
<td>MoVi [<xref ref-type="bibr" rid="ref-195">195</xref>]</td>
<td>2020</td>
<td>N/A<break/>(17 h)</td>
<td>Large single-player video dataset with 3DMoCap annotations<break/>Can provides SMPL parameters obtained through MoSh&#x002B;&#x002B;</td>
<td>S</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6_5">
<label>6.5</label>
<title>Current Research Direction</title>
<p>At present, posture recognition is divided into the following research directions:
<list list-type="simple">
<list-item><label>(1)</label><p>Pose machines. The pose machine is a mature 2D human posture recognition method. In order to make use of the excellent image feature extraction ability of the convolutional neural network, the convolutional neural network is integrated into the framework of the pose machine [<xref ref-type="bibr" rid="ref-88">88</xref>].</p></list-item>
<list-item><label>(2)</label><p>Convolutional network structure. In recent years, significant progress has been made in posture recognition based on convolutional network structure, but there is still room for optimization in recognition performance. Many researchers focused on the optimization of convolutional network structures, and some optimization models were proposed, such as stacked hourglass network [<xref ref-type="bibr" rid="ref-89">89</xref>], iterative error feedback (IEF) [<xref ref-type="bibr" rid="ref-154">154</xref>], Mask R-CNN [<xref ref-type="bibr" rid="ref-156">156</xref>], and the EfficientPose [<xref ref-type="bibr" rid="ref-169">169</xref>].</p></list-item>
<list-item><label>(3)</label><p>Multi-person posture recognition in natural scenes. Due to many factors, such as complex background, occlusive congestion, and posture difference in the natural environment, many posture recognition methods with a fine performance in the experimental environment are ineffective in multi-person posture recognition tasks. However, with the development of the field of posture recognition, multi-person recognition in natural scenes is very worthy of study. Fortunately, the recognition of multiple people in natural scenes has attracted the attention of many scholars [<xref ref-type="bibr" rid="ref-155">155</xref>].</p></list-item>
<list-item><label>(4)</label><p>Attention mechanism. By designing different attention mechanism characteristics for each part of the human body, more accurate human posture recognition results can be obtained. Some attention-related strategies have been proposed, for example, the attention regularization loss based on local feature identity to constrain attention weight [<xref ref-type="bibr" rid="ref-135">135</xref>], the convolutional neural network with multi-context attention mechanism is incorporated into the end-to-end framework of posture recognition [<xref ref-type="bibr" rid="ref-159">159</xref>], and the polarization self-attention block is realized through polarization filtering and enhancement techniques [<xref ref-type="bibr" rid="ref-168">168</xref>].</p></list-item>
<list-item><label>(5)</label><p>Data fusion. The performance of the data fusion algorithm directly affects the accuracy of posture recognition and the reliability of the system [<xref ref-type="bibr" rid="ref-196">196</xref>,<xref ref-type="bibr" rid="ref-197">197</xref>]. Data fusion strategies include multi-sensor-based data fusion [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-198">198</xref>,<xref ref-type="bibr" rid="ref-199">199</xref>], position and posture-based fusion [<xref ref-type="bibr" rid="ref-200">200</xref>], multi-feature fusion [<xref ref-type="bibr" rid="ref-201">201</xref>,<xref ref-type="bibr" rid="ref-202">202</xref>], and so on.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Conclusion and Future Directions</title>
<p>This paper reviews and summarizes the methods and techniques of posture recognition. It mainly includes the following aspects: (i) The structure and related algorithms based on traditional machine learning and deep neural network are presented; (ii) The background and application of three posture recognition techniques are presented, and their characteristics are compared; (iii) Several common posture recognition network structures based on CNN are presented and compared; (iv) Three typical lightweight network design methods are presented; (v) The commonly used datasets for posture recognition are summarized, and the limitations of 2D and 3D datasets are talked about. In summary, the framework of our review is shown in <xref ref-type="fig" rid="fig-17">Fig. 17</xref>.</p>
<fig id="fig-17">
<label>Figure 17</label>
<caption>
<title>The systematic diagram of our study</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_27676-fig-17.tif"/>
</fig>
<sec id="s7_1">
<label>7.1</label>
<title>Limitations and Challenges</title>
<p>Although the techniques and methods of posture recognition have made great progress in recent years, posture research will still face challenges due to the complexity of the task and the different requirements of different fields. Through the research, we believe that the challenges facing posture recognition at this stage mainly include the following aspects, as shown in <xref ref-type="table" rid="table-8">Table 8</xref>.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Possible posture recognition challenges</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Number</th>
<th>Challenges</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>Datasets problems</td>
<td>Lack of special posture datasets and large outdoor 3D datasets.</td>
</tr>
<tr>
<td>2</td>
<td>Poor generalization ability</td>
<td>Poor generalization ability leads to low accuracy of posture recognition.</td>
</tr>
<tr>
<td>3</td>
<td>Human body occlusion problem</td>
<td>It includes the occlusion of the human body itself, the occlusion of other objects on the human body, and the occlusion of other human bodies on the human body.</td>
</tr>
<tr>
<td>4</td>
<td>The contradiction between model accuracy and computational power and large storage space</td>
<td>The increase in the complexity of neural network models leads to an increase in the number of parameters and the demand for computing resources.</td>
</tr>
<tr>
<td>5</td>
<td>Depth ambiguity problem</td>
<td>There may be multiple postures in the 3D space that correspond to the human posture in the 2D image.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>(1) Datasets problems
<list list-type="simple">
<list-item><label>(i)</label><p>Lack of special posture datasets. Currently, the existing public datasets have a large amount of data, but most of the human posture is normal, such as standing, walking, and so on. Lack of special postures, such as falling, crowding, etc.</p></list-item>
<list-item><label>(ii)</label><p>Lack of large outdoor 3D datasets. The production of 3D posture datasets mostly relies on motion capture equipment, which has restrictions on the environment and the range of human activity, so 3D datasets in outdoor scenes are relatively scarce.</p></list-item>
</list></p>
<p>(2) Poor generalization ability</p>
<p>Since many datasets are established in the experimental environment, the generalization ability of the human posture recognition model in natural scenes is poor, and it is difficult to achieve an accurate posture recognition effect in practical applications [<xref ref-type="bibr" rid="ref-203">203</xref>].</p>
<p>(3) Human body occlusion problem</p>
<p>Human body occlusion is one of the most important problems in the process of posture recognition, especially in the natural environment of multi-person posture recognition. Human body occlusion is very common. The phenomenon of human-body occlusion includes the occlusion of the human body itself, the occlusion of other objects on the human body, and the occlusion of other human bodies on the human body [<xref ref-type="bibr" rid="ref-204">204</xref>,<xref ref-type="bibr" rid="ref-205">205</xref>]. The occlusion of the human body has a great influence on the prediction of human body joints.</p>
<p>(4) The contradiction between model accuracy and computational power and large storage space</p>
<p>Deep learning algorithm has become the mainstream method of posture recognition. Many existing posture recognition technologies based on deep learning blindly pursue accuracy, and the design of complex and multi-level networks leads to high requirements on hardware, which is not good for the wide application of neural networks. Therefore, it is particularly important to carry out lightweight design on the network while maintaining recognition accuracy.</p>
<p>(5) Depth ambiguity problem</p>
<p>Depth ambiguity is a problem in 3D posture recognition, which may result in multiple 3D postures corresponding to the same 2D projection. Additional information needs to be added by the algorithm to recover the correct 3D posture [<xref ref-type="bibr" rid="ref-206">206</xref>]. Many approaches attempt to solve this problem by using a variety of prior information, such as geometric prior knowledge, statistical models, and temporal smoothness [<xref ref-type="bibr" rid="ref-207">207</xref>]. However, there are still some unsolved challenges and gaps between research and practical application.</p>
</sec>
<sec id="s7_2">
<label>7.2</label>
<title>Future Research Directions</title>
<p>In the future, the research of posture recognition can proceed from the following two aspects of the above discussion of the challenges. (i) Establish an appropriate posture benchmark database, which can be integrated and improved. (ii) The technology based on CNN and other deep neural networks have the potential for improvement, which can be researched in feature extraction, information fusion, and other aspects. (iii) The robustness and stability of body mesh reconstruction under heavy occlusion need to be further explored [<xref ref-type="bibr" rid="ref-208">208</xref>]. (iv) Lightweight network design can be used to solve the contradiction between model accuracy, computing power, and large storage space. It still has a lot of room for improvement in recognition accuracy.</p>
</sec>
</sec>
</body>
<back>
<sec><title>Funding Statement</title>
<p>The paper is partially supported by <funding-source>British Heart Foundation Accelerator Award</funding-source>, UK (<award-id>AA/18/3/34220</award-id>); <funding-source>Royal Society International Exchanges Cost Share Award</funding-source>, UK (<award-id>RP202G0230</award-id>); <funding-source>Hope Foundation for Cancer Research</funding-source>, UK (<award-id>RM60G0680</award-id>); <funding-source>Medical Research Council Confidence in Concept Award</funding-source>, UK (<award-id>MC_PC_17171</award-id>); <funding-source>Sino-UK Industrial Fund</funding-source>, UK (<award-id>RP202G0289</award-id>); <funding-source>Global Challenges Research Fund (GCRF)</funding-source>, UK (<award-id>P202PF11</award-id>); <funding-source>LIAS Pioneering Partnerships award</funding-source>, UK (<award-id>P202ED10</award-id>); <funding-source>Data Science Enhancement Fund</funding-source>, UK (<award-id>P202RE237</award-id>); <funding-source>Fight for Sight</funding-source>, UK (<award-id>24NN201</award-id>); <funding-source>Sino-UK Education Fund</funding-source>, UK (<award-id>OP202006</award-id>).</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>G&#x00F3;rriz</surname>, <given-names>J. M.</given-names></string-name>, <string-name><surname>Ram&#x00ED;rez</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Ort&#x00ED;z</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Mart&#x00ED;nez-Murcia</surname>, <given-names>F. J.</given-names></string-name>, <string-name><surname>Segovia</surname>, <given-names>F.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Artificial intelligence within the interplay between natural and artificial computation: Advances in data science, trends and applications</article-title>. <source>Neurocomputing</source><italic>,</italic> <volume>410</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>237</fpage>&#x2013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.05.078</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Islam</surname>, <given-names>M. M.</given-names></string-name>, <string-name><surname>Nooruddin</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Karray</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Multimodal human activity recognition for smart healthcare applications</article-title>. <conf-name>2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC)</conf-name>, pp. <fpage>196</fpage>&#x2013;<lpage>203</lpage>. <publisher-loc>Prague, Czech Republic</publisher-loc>,
 <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>C. W.</given-names></string-name>, <string-name><surname>Hong</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Lin</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Bird posture recognition based on target keypoints estimation in dual-task convolutional neural networks</article-title>. <source>Ecological Indicators</source><italic>,</italic> <volume>135</volume><italic>,</italic> <fpage>108506</fpage>. <pub-id pub-id-type="doi">10.1016/j.ecolind.2021.108506</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Pu</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Mu</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Pig-posture recognition based on computer vision: Dataset and exploration</article-title>. <source>Animals</source><italic>,</italic> <volume>11</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>1295</fpage>. <pub-id pub-id-type="doi">10.3390/ani11051295</pub-id>; <pub-id pub-id-type="pmid">33946472</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Tian</surname>, <given-names>J. Y.</given-names></string-name></person-group> (<year>2018</year>). <article-title>The recognition of pig posture based on target features and decision tree support vector machine</article-title>. <source>Science Technology and Engineering</source><italic>,</italic> <volume>18</volume><italic>,</italic> <fpage>297</fpage>&#x2013;<lpage>301</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Norton</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Behaviour recognition of pigs and cattle: Journey from computer vision to deep learning</article-title>. <source>Computers and Electronics in Agriculture</source><italic>,</italic> <volume>187</volume><italic>(</italic><issue>1&#x2013;3</issue><italic>),</italic> <fpage>106255</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2021.106255</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Siddiqui</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Human pose estimation from a single view point</article-title>. <conf-name>Computer Vision and Pattern Recognition Workshops (CVPRW), 2010 IEEE Computer Society Conference</conf-name>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <publisher-loc>San Francisco, CA, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Lin</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Guo</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2019</year>). <article-title>A posture recognition method applied to smart product service</article-title>. <source>Procedia CIRP</source><italic>,</italic> <volume>83</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>425</fpage>&#x2013;<lpage>428</lpage>. <pub-id pub-id-type="doi">10.1016/j.procir.2019.04.145</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Satapathy</surname>, <given-names>S. C.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>S. H.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Y. D.</given-names></string-name></person-group> (<year>2020</year>). <article-title>A survey on artificial intelligence in Chinese sign language recognition</article-title>. <source>Arabian Journal for Science and Engineering</source><italic>,</italic> <volume>45</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>9859</fpage>&#x2013;<lpage>9894</lpage>. <pub-id pub-id-type="doi">10.1007/s13369-020-04758-2</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bao</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Intille</surname>, <given-names>S. S.</given-names></string-name></person-group> (<year>2004</year>). <article-title>Activity recognition from user-annotated acceleration data</article-title>. <conf-name>International Conference on Pervasive Computing</conf-name>, pp. <fpage>1</fpage>&#x2013;<lpage>17</lpage>. <publisher-loc>Linz/Vienna, Austria</publisher-loc>, <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yadav</surname>, <given-names>S. K.</given-names></string-name>, <string-name><surname>Agarwal</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Kumar</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Tiwari</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Pandey</surname>, <given-names>H. M.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>YogNet: A two-stream network for realtime multiperson yoga action recognition and posture correction</article-title>. <source>Knowledge-Based Systems</source><italic>,</italic> <volume>250</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>109097</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2022.109097</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nooruddin</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Islam</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Sharna</surname>, <given-names>F. A.</given-names></string-name>, <string-name><surname>Alhetari</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Kabir</surname>, <given-names>M. N.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Sensor-based fall detection systems: A review</article-title>. <source>Journal of Ambient Intelligence and Humanized Computing</source><italic>,</italic> <volume>13</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>2735</fpage>&#x2013;<lpage>2751</lpage>. <pub-id pub-id-type="doi">10.1007/s12652-021-03248-z</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Multi-sensor detection and control network technology based on parallel computing model in robot target detection and recognition</article-title>. <source>Computer Communications</source><italic>,</italic> <volume>159</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>215</fpage>&#x2013;<lpage>221</lpage>. <pub-id pub-id-type="doi">10.1016/j.comcom.2020.05.006</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Ding</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2009</year>). <article-title>An adaptive PNN-DS approach to classification using multi-sensor information fusion</article-title>. <source>Neural Computing and Applications</source><italic>,</italic> <volume>18</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>455</fpage>&#x2013;<lpage>467</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-008-0220-4</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Shen</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Gao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Deng</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Learning multi-level features for sensor-based human action recognition</article-title>. <source>Pervasive and Mobile Computing</source><italic>,</italic> <volume>40</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>324</fpage>&#x2013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1016/j.pmcj.2017.07.001</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmed</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Antar</surname>, <given-names>A. D.</given-names></string-name>, <string-name><surname>Ahad</surname>, <given-names>M. A. R.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Static postural transition-based technique and efficient feature extraction for sensor-based activity recognition</article-title>. <source>Pattern Recognition Letters</source><italic>,</italic> <volume>147</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>25</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2021.04.001</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qiu</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Zhao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>L.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>Multi-sensor information fusion based on machine learning for real applications in human activity recognition: State-of-the-art and research challenges</article-title>. <source>Information Fusion</source><italic>,</italic> <volume>80</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>241</fpage>&#x2013;<lpage>265</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2021.11.006</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Antwi-Afari</surname>, <given-names>M. F.</given-names></string-name>, <string-name><surname>Qarout</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Herzallah</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Anwer</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Umer</surname>, <given-names>W.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>Deep learning-based networks for automated recognition and classification of awkward working postures in construction using wearable insole sensor data</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>136</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>104181</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2022.104181</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hong</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Hong</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Ma</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>X.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>A wearable-based posture recognition system with AI-assisted approach for healthcare IoT</article-title>. <source>Future Generation Computer Systems</source><italic>,</italic> <volume>127</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>286</fpage>&#x2013;<lpage>296</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2021.08.030</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Bi</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Hybrid lightweight deep-learning model for sensor-fusion basketball shooting-posture recognition</article-title>. <source>Measurement</source><italic>,</italic> <volume>189</volume><italic>(</italic><issue>8</issue><italic>),</italic> <fpage>110595</fpage>. <pub-id pub-id-type="doi">10.1016/j.measurement.2021.110595</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sardar</surname>, <given-names>A. W.</given-names></string-name>, <string-name><surname>Ullah</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Bacha</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Khan</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Ali</surname>, <given-names>F.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>Mobile sensors based platform of human physical activities recognition for COVID-19 spread minimization</article-title>. <source>Computers in Biology and Medicine</source><italic>,</italic> <volume>146</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>105662</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2022.105662</pub-id>; <pub-id pub-id-type="pmid">35654623</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Xiang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Pan</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2020</year>). <article-title>3D PostureNet: A unified framework for skeleton-based posture recognition</article-title>. <source>Pattern Recognition Letters</source><italic>,</italic> <volume>140</volume><italic>(</italic><issue>8</issue><italic>),</italic> <fpage>143</fpage>&#x2013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2020.09.029</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abedi</surname>, <given-names>W. M. S.</given-names></string-name>, <string-name><surname>Ibraheem Nadher</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Sadiq</surname>, <given-names>A. T.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Modified deep learning method for body postures recognition</article-title>. <source>International Journal of Advanced Science and Technology</source><italic>,</italic> <volume>29</volume><italic>,</italic> <fpage>3830</fpage>&#x2013;<lpage>3841</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tome</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Russell</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Agapito</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Lifting from the deep: Convolutional 3D pose estimation from a single image</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>2500</fpage>&#x2013;<lpage>2509</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fang</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Ma</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Vision-based posture-consistent teleoperation of robotic arm using multi-stage deep neural network</article-title>. <source>Robotics and Autonomous Systems</source><italic>,</italic> <volume>131</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>103592</fpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2020.103592</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kumar</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sangwan</surname>, <given-names>K. S.</given-names></string-name></person-group> (<year>2021</year>). <article-title>A computer vision based approach fordriver distraction recognition using deep learning and genetic algorithm based ensemble</article-title>. <conf-name>International Conference on Artificial Intelligence and Soft Computing</conf-name>, pp. <fpage>44</fpage>&#x2013;<lpage>56</lpage>. <publisher-loc>Zakopane, Poland</publisher-loc>, <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mehrizi</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Peng</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Metaxas</surname>, <given-names>D.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2018</year>). <article-title>A computer vision based method for 3D posture estimation of symmetrical lifting</article-title>. <source>Journal of Biomechanics</source><italic>,</italic> <volume>69</volume><italic>(</italic><issue>1&#x2013;2</issue><italic>),</italic> <fpage>40</fpage>&#x2013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1016/j.jbiomech.2018.01.012</pub-id>; <pub-id pub-id-type="pmid">29398001</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yao</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Sheng</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Ruan</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Gu</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2015</year>). <article-title>RF-care: Device-free posture recognition for elderly people using a passive rfid tag array</article-title>. <conf-name>Proceedings of the 12th International Conference on Mobile and Ubiquitous Systems: Computing, Networking and Services</conf-name>, pp. <fpage>120</fpage>&#x2013;<lpage>129</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yao</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Sheng</surname>, <given-names>Q. Z.</given-names></string-name>, <string-name><surname>Xue</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Tao</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Wan</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Compressive representation for device-free activity recognition with passive RFID signal strength</article-title>. <source>IEEE Transactions on Mobile Computing</source><italic>,</italic> <volume>17</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>293</fpage>&#x2013;<lpage>306</lpage>. <pub-id pub-id-type="doi">10.1109/TMC.2017.2706282</pub-id></mixed-citation></ref>
<ref id="ref-30"><label>30.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2019</year>). <article-title>TagSheet: Sleeping posture recognition with an unobtrusive passive tag matrix</article-title>. <conf-name>IEEE Conference on Computer Communications</conf-name>, pp. <fpage>874</fpage>&#x2013;<lpage>882</lpage>. <publisher-loc>Paris, France</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hussain</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Sheng</surname>, <given-names>Q. Z.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>W. E.</given-names></string-name></person-group> (<year>2020</year>). <article-title>A review and categorization of techniques on device-free human activity recognition</article-title>. <source>Journal of Network and Computer Applications</source><italic>,</italic> <volume>167</volume><italic>,</italic> <fpage>102738</fpage>. <pub-id pub-id-type="doi">10.1016/j.jnca.2020.102738</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>32.</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Islam</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Nooruddin</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Karray</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Muhammad</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Human activity recognition using tools of convolutional neural networks: A state of the art review, data sets, challenges and future prospects</article-title>. <comment>arXiv preprint arXiv: 220203274</comment>.</mixed-citation></ref>
<ref id="ref-33"><label>33.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Research on pedestrian detection technology based on the SVM classifier trained by HOG and LTP features</article-title>. <source>Future Generation Computer Systems</source><italic>,</italic> <volume>125</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>604</fpage>&#x2013;<lpage>615</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2021.06.016</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>34.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vashisth</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Saurav</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Histogram of oriented gradients based reduced feature for traffic sign recognition</article-title>. <conf-name>2018 International Conference on Advances in Computing, Communications and Informatics (ICACCI)</conf-name>, pp. <fpage>2206</fpage>&#x2013;<lpage>2212</lpage>. <publisher-loc>Bangalore, India</publisher-loc>.</mixed-citation></ref>
<ref id="ref-35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lowe</surname>, <given-names>D. G.</given-names></string-name></person-group> (<year>2004</year>). <article-title>Distinctive image features from scale-invariant key-points</article-title>. <source>International Journal of Computer Vision</source><italic>,</italic> <volume>60</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>91</fpage>&#x2013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1023/B:VISI.0000029664.99615.94</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>36.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Giveki</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Soltanshahi</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Montazer</surname>, <given-names>G. A.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A new image feature descriptor for content based image retrieval using scale invariant feature transform and local derivative pattern</article-title>. <source>Optik</source><italic>,</italic> <volume>131</volume><italic>(</italic><issue>8</issue><italic>),</italic> <fpage>242</fpage>&#x2013;<lpage>254</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijleo.2016.11.046</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>37.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Handwriting posture prediction based on unsupervised model</article-title>. <source>Pattern Recognition</source><italic>,</italic> <volume>100</volume><italic>,</italic> <fpage>107093</fpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2019.107093</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>38.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Oszust</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Krupski</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Isolated sign language recognition with depth cameras</article-title>. <source>Procedia Computer Science</source><italic>,</italic> <volume>192</volume><italic>,</italic> <fpage>2085</fpage>&#x2013;<lpage>2094</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2021.08.216</pub-id></mixed-citation></ref>
<ref id="ref-39"><label>39.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ramezanpanah</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Mallem</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Davesne</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Human action recognition using laban movement analysis and dynamic time warping</article-title>. <source>Procedia Computer Science</source><italic>,</italic> <volume>176</volume><italic>(</italic><issue>27</issue><italic>),</italic> <fpage>390</fpage>&#x2013;<lpage>399</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2020.08.040</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>40.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yoon</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Lee</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Chun</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Son</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Investigation of the relationship between Ironworker&#x2019;s gait stability and different types of load carrying using wearable sensors</article-title>. <source>Advanced Engineering Informatics</source><italic>,</italic> <volume>51</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>101521</fpage>. <pub-id pub-id-type="doi">10.1016/j.aei.2021.101521</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>41.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ghersi</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Ferrando</surname>, <given-names>M. H.</given-names></string-name>, <string-name><surname>Fliger</surname>, <given-names>C. G.</given-names></string-name>, <string-name><surname>Castro Arenas</surname>, <given-names>C. F.</given-names></string-name>, <string-name><surname>Edwards Molina</surname>, <given-names>D. J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Gait-cycle segmentation method based on lower-trunk acceleration signals and dynamic time warping</article-title>. <source>Medical Engineering &#x0026; Physics</source><italic>,</italic> <volume>82</volume><italic>,</italic> <fpage>70</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1016/j.medengphy.2020.06.001</pub-id>; <pub-id pub-id-type="pmid">32709267</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>42.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hern&#x00E1;ndez-Vela</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Bautista</surname>, <given-names>M. &#x00C1;.</given-names></string-name>, <string-name><surname>Perez-Sala</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Ponce-L&#x00F3;pez</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Escalera</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2014</year>). <article-title>Probability-based dynamic time warping and bag-of-visual-and-depth-words for human gesture recognition in RGB-D</article-title>. <source>Pattern Recognition Letters</source><italic>,</italic> <volume>50</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>112</fpage>&#x2013;<lpage>121</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2013.09.009</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>43.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>&#x017D;emgulys</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Raudonis</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Maskeli&#x016B;nas</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Dama&#x0161;evi&#x010D;ius</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Recognition of basketball referee signals from videos using Histogram of Oriented Gradients (HOG) and Support Vector Machine (SVM)</article-title>. <source>Procedia Computer Science</source><italic>,</italic> <volume>130</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>953</fpage>&#x2013;<lpage>960</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2018.04.095</pub-id></mixed-citation></ref>
<ref id="ref-44"><label>44.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Onishi</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Takiguchi</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Ariki</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2008</year>). <article-title>3D human posture estimation using the HOG features from monocular image</article-title>. <conf-name>2008 19th International Conference on Pattern Recognition</conf-name>, pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>. <publisher-loc>Tampa, FL, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-45"><label>45.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Seemanthini</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Manjunath</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Human detection and tracking using HOG for action recognition</article-title>. <source>Procedia Computer Science</source><italic>,</italic> <volume>132</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1317</fpage>&#x2013;<lpage>1326</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2018.05.048</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>46.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cheng</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Ogunbona</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2009</year>). <article-title>Kernel PCA of HOG features for posture detection</article-title>. <conf-name>2009 24th International Conference Image and Vision Computing</conf-name>, pp. <fpage>415</fpage>&#x2013;<lpage>420</lpage>. <publisher-loc>New Zealand</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-47"><label>47.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>C. C.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>K. C.</given-names></string-name></person-group> (<year>2007</year>). <chapter-title>Hand posture recognition using adaboost with sift for human robot interaction</chapter-title>. In: <source>Recent progress in robotics: Viable robotic service to human</source>, pp. <fpage>317</fpage>&#x2013;<lpage>329</lpage>. <publisher-loc>Berlin, Heidelberg</publisher-loc>, <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-48"><label>48.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Xi</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2020</year>). <article-title>A human body based on sift-neural network algorithm attitude recognition method</article-title>. <source>Journal of Medical Imaging and Health Informatics</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>129</fpage>&#x2013;<lpage>133</lpage>. <pub-id pub-id-type="doi">10.1166/jmihi.2020.2867</pub-id></mixed-citation></ref>
<ref id="ref-49"><label>49.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Liang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Liang</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Head pose estimation with combined 2D SIFT and 3D HOG features</article-title>. <conf-name>2013 Seventh International Conference on Image and Graphics</conf-name>, pp. <fpage>650</fpage>&#x2013;<lpage>655</lpage>. <publisher-loc>Qingdao, China</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-50"><label>50.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ning</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Na</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Deep spatial/temporal-level feature engineering for Tennis-based action recognition</article-title>. <source>Future Generation Computer Systems</source><italic>,</italic> <volume>125</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>188</fpage>&#x2013;<lpage>193</lpage>. <pub-id pub-id-type="doi">10.1016/j.future.2021.06.022</pub-id></mixed-citation></ref>
<ref id="ref-51"><label>51.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Atrevi</surname>, <given-names>D. F.</given-names></string-name>, <string-name><surname>Vivet</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Duculty</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Emile</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A very simple framework for 3D human poses estimation using a single 2D image: Comparison of geometric moments descriptors</article-title>. <source>Pattern Recognition</source><italic>,</italic> <volume>71</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>389</fpage>&#x2013;<lpage>401</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2017.06.024</pub-id></mixed-citation></ref>
<ref id="ref-52"><label>52.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>C. J.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Review on vision-based pose estimation of UAV based on landmark</article-title>. <conf-name>2017 2nd International Conference on Frontiers of Sensors Technologies (ICFST)</conf-name>, pp. <fpage>453</fpage>&#x2013;<lpage>457</lpage>. <publisher-loc>Shenzhen, China</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-53"><label>53.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Comellini</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Le Le Ny</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Zenou</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Espinosa</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Dubanchet</surname>, <given-names>V.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Global descriptors for visual pose estimation of a noncooperative target in space rendezvous</article-title>. <source>IEEE Transactions on Aerospace and Electronic Systems</source><italic>,</italic> <volume>57</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>4197</fpage>&#x2013;<lpage>4212</lpage>. <pub-id pub-id-type="doi">10.1109/TAES.2021.3086888</pub-id></mixed-citation></ref>
<ref id="ref-54"><label>54.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rong</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Kong</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Yin</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2018</year>). <article-title>RGB-D hand pose estimation using fourier descriptor</article-title>. <conf-name>2018 7th International Conference on Digital Home (ICDH)</conf-name>, pp. <fpage>50</fpage>&#x2013;<lpage>56</lpage>. <publisher-loc>Guilin, China</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-55"><label>55.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dedeolu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Treyin</surname>, <given-names>B. U.</given-names></string-name>, <string-name><surname>G&#x00FC;d&#x00FC;kbay</surname>, <given-names>U.</given-names></string-name>, <string-name><surname>Etin</surname>, <given-names>A. E.</given-names></string-name></person-group> (<year>2006</year>). <article-title>Silhouette-based method for object classification and human action recognition in video</article-title>. <conf-name>International Conference on Computer Vision</conf-name>, pp. <fpage>64</fpage>&#x2013;<lpage>77</lpage>. <publisher-loc>Graz, Austria</publisher-loc>.</mixed-citation></ref>
<ref id="ref-56"><label>56.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cherla</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Kulkarni</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Kale</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Ramasubramanian</surname>, <given-names>V.</given-names></string-name></person-group> (<year>2008</year>). <article-title>Towards fast, view-invariant human action recognition</article-title>. <conf-name>2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops</conf-name>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <publisher-loc>Anchorage, AK, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-57"><label>57.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Hu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>W. M.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>D. W.</given-names></string-name>, <string-name><surname>Ma</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2022</year>). <article-title>A deep learning based 2-dimensional hip pressure signals analysis method for sitting posture recognition</article-title>. <source>Biomedical Signal Processing and Control</source><italic>,</italic> <volume>73</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>103432</fpage>. <pub-id pub-id-type="doi">10.1016/j.bspc.2021.103432</pub-id></mixed-citation></ref>
<ref id="ref-58"><label>58.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koprowski</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2015</year>). <article-title>Automatic analysis of the trunk thermal images from healthy subjects and patients with faulty posture</article-title>. <source>Computers in Biology and Medicine</source><italic>,</italic> <volume>62</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>110</fpage>&#x2013;<lpage>118</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2015.04.017</pub-id>; <pub-id pub-id-type="pmid">25929672</pub-id></mixed-citation></ref>
<ref id="ref-59"><label>59.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Federolf</surname>, <given-names>P. A.</given-names></string-name></person-group> (<year>2016</year>). <article-title>A novel approach to study human posture control: Principal movements obtained from a principal component analysis of kinematic marker data</article-title>. <source>Journal of Biomechanics</source><italic>,</italic> <volume>49</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>364</fpage>&#x2013;<lpage>370</lpage>. <pub-id pub-id-type="doi">10.1016/j.jbiomech.2015.12.030</pub-id>; <pub-id pub-id-type="pmid">26768228</pub-id></mixed-citation></ref>
<ref id="ref-60"><label>60.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Iosifidis</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Tefas</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Nikolaidis</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Pitas</surname>, <given-names>I.</given-names></string-name></person-group> (<year>2012</year>). <article-title>Multi-view human movement recognition based on fuzzy distances and linear discriminant analysis</article-title>. <source>Computer Vision and Image Understanding</source><italic>,</italic> <volume>116</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>347</fpage>&#x2013;<lpage>360</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2011.08.008</pub-id></mixed-citation></ref>
<ref id="ref-61"><label>61.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hsia</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Liou</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Aung</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Foo</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>W.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2009</year>). <article-title>Analysis and comparison of sleeping posture classification methods using pressure sensitive bed system</article-title>. <conf-name>2009 Annual International Conference of the IEEE Engineering in Medicine and Biology Society</conf-name>, pp. <fpage>6131</fpage>&#x2013;<lpage>6134</lpage>. <publisher-loc>Minneapolis, MN, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-62"><label>62.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dolphens</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Cagnie</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Coorevits</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Vleeming</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Palmans</surname>, <given-names>T.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2014</year>). <article-title>Posture class prediction of pre-peak height velocity subjects according to gross body segment orientations using linear discriminant analysis</article-title>. <source>European Spine Journal</source><italic>,</italic> <volume>23</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>530</fpage>&#x2013;<lpage>535</lpage>. <pub-id pub-id-type="doi">10.1007/s00586-013-3058-0</pub-id>; <pub-id pub-id-type="pmid">24097292</pub-id></mixed-citation></ref>
<ref id="ref-63"><label>63.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cortes</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Vapnik</surname>, <given-names>V. N.</given-names></string-name></person-group> (<year>1995</year>). <article-title>Support vector networks</article-title>. <source>Machine Learning</source><italic>,</italic> <volume>20</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>273</fpage>&#x2013;<lpage>297</lpage>. <pub-id pub-id-type="doi">10.1007/BF00994018</pub-id></mixed-citation></ref>
<ref id="ref-64"><label>64.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Shi</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A novel local feature descriptor based on energy information for human activity recognition</article-title>. <source>Neurocomputing</source><italic>,</italic> <volume>228</volume><italic>(</italic><issue>11</issue><italic>),</italic> <fpage>19</fpage>&#x2013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2016.07.058</pub-id></mixed-citation></ref>
<ref id="ref-65"><label>65.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alcaraz</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Labb&#x00E9;</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Landete</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Support vector machine with feature selection: A multiobjective approach</article-title>. <source>Expert Systems with Applications</source><italic>,</italic> <volume>204</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>117485</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2022.117485</pub-id></mixed-citation></ref>
<ref id="ref-66"><label>66.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Sang</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Face-mask recognition for fraud prevention using Gaussian mixture model</article-title>. <source>Journal of Visual Communication and Image Representation</source><italic>,</italic> <volume>55</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>795</fpage>&#x2013;<lpage>801</lpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2018.08.016</pub-id></mixed-citation></ref>
<ref id="ref-67"><label>67.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Rajan</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Jeon</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Chang</surname>, <given-names>J. H.</given-names></string-name>, <string-name><surname>Dajani</surname>, <given-names>H. R.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Oscillometric blood pressure estimation by combining nonparametric bootstrap with Gaussian mixture model</article-title>. <source>Computers in Biology and Medicine</source><italic>,</italic> <volume>85</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>112</fpage>&#x2013;<lpage>124</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2015.11.008</pub-id>; <pub-id pub-id-type="pmid">26654485</pub-id></mixed-citation></ref>
<ref id="ref-68"><label>68.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname>, <given-names>C. L.</given-names></string-name>, <string-name><surname>Wu</surname>, <given-names>M. S.</given-names></string-name>, <string-name><surname>Jeng</surname>, <given-names>S. H.</given-names></string-name></person-group> (<year>2000</year>). <article-title>Gesture recognition using the multi-PDM method and hidden Markov model</article-title>. <source>Image and Vision Computing</source><italic>,</italic> <volume>18</volume><italic>(</italic><issue>11</issue><italic>),</italic> <fpage>865</fpage>&#x2013;<lpage>879</lpage>. <pub-id pub-id-type="doi">10.1016/S0262-8856(99)00042-6</pub-id></mixed-citation></ref>
<ref id="ref-69"><label>69.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>S&#x00E1;nchez</surname>, <given-names>V. G.</given-names></string-name>, <string-name><surname>Lysaker</surname>, <given-names>O. M.</given-names></string-name>, <string-name><surname>Skeie</surname>, <given-names>N. O.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Human behaviour modelling for welfare technology using hidden Markov models</article-title>. <source>Pattern Recognition Letters</source><italic>,</italic> <volume>137</volume><italic>,</italic> <fpage>71</fpage>&#x2013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2019.09.022</pub-id></mixed-citation></ref>
<ref id="ref-70"><label>70.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Ding</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2022</year>). <article-title>A fusion of a deep neural network and a hidden Markov model to recognize the multiclass abnormal behavior of elderly people</article-title>. <source>Knowledge-Based Systems</source><italic>,</italic> <volume>252</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>109351</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2022.109351</pub-id></mixed-citation></ref>
<ref id="ref-71"><label>71.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rabiner</surname>, <given-names>L. R.</given-names></string-name></person-group> (<year>1989</year>). <article-title>A tutorial on hidden Markov models and selected applications in speech recognition</article-title>. <source>Proceedings of the IEEE</source><italic>,</italic> <volume>77</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>257</fpage>&#x2013;<lpage>286</lpage>. <pub-id pub-id-type="doi">10.1109/5.18626</pub-id></mixed-citation></ref>
<ref id="ref-72"><label>72.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Haider</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Abbasi</surname>, <given-names>Q. H.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Post-surgical fall detection by exploiting the 5G C-band technology for eHealth paradigm</article-title>. <source>Applied Soft Computing</source><italic>,</italic> <volume>81</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>105537</fpage>. <pub-id pub-id-type="doi">10.1016/j.asoc.2019.105537</pub-id></mixed-citation></ref>
<ref id="ref-73"><label>73.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nasirahmadi</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sturm</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Olsson</surname>, <given-names>A. C.</given-names></string-name>, <string-name><surname>Jeppsson</surname>, <given-names>K. H.</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Automatic scoring of lateral and sternal lying posture in grouped pigs using image processing and support vector machine</article-title>. <source>Computers and Electronics in Agriculture</source><italic>,</italic> <volume>156</volume><italic>(</italic><issue>1&#x2013;3</issue><italic>),</italic> <fpage>475</fpage>&#x2013;<lpage>481</lpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2018.12.009</pub-id></mixed-citation></ref>
<ref id="ref-74"><label>74.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yin</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Person re-identification by integrating metric learning and support vector machine</article-title>. <source>Signal Processing</source><italic>,</italic> <volume>166</volume><italic>,</italic> <fpage>107277</fpage>. <pub-id pub-id-type="doi">10.1016/j.sigpro.2019.107277</pub-id></mixed-citation></ref>
<ref id="ref-75"><label>75.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bonneau</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Benet</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Labrune</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Bailly</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Ricard</surname>, <given-names>E.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>Predicting sow postures from video images: Comparison of convolutional neural networks and segmentation combined with support vector machines under various training and testing setups</article-title>. <source>Biosystems Engineering</source><italic>,</italic> <volume>212</volume><italic>(</italic><issue>3&#x2013;4</issue><italic>),</italic> <fpage>19</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1016/j.biosystemseng.2021.09.014</pub-id></mixed-citation></ref>
<ref id="ref-76"><label>76.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ameli</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Naghdy</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Stirling</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Naghdy</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Aghmesheh</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Objective clinical gait analysis using inertial sensors and six minute walking test</article-title>. <source>Pattern Recognition</source><italic>,</italic> <volume>63</volume><italic>(</italic><issue>9604</issue><italic>),</italic> <fpage>246</fpage>&#x2013;<lpage>257</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2016.08.002</pub-id></mixed-citation></ref>
<ref id="ref-77"><label>77.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zheng</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Target detection algorithm for dance moving images based on sensor and motion capture data</article-title>. <source>Microprocessors and Microsystems</source><italic>,</italic> <volume>81</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>103743</fpage>. <pub-id pub-id-type="doi">10.1016/j.micpro.2020.103743</pub-id></mixed-citation></ref>
<ref id="ref-78"><label>78.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mallick</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Das</surname>, <given-names>P. P.</given-names></string-name>, <string-name><surname>Majumdar</surname>, <given-names>A. K.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Posture and sequence recognition for Bharatanatyam dance performances using machine learning approaches</article-title>. <source>Journal of Visual Communication and Image Representation</source><italic>,</italic> <volume>87</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>103548</fpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2022.103548</pub-id></mixed-citation></ref>
<ref id="ref-79"><label>79.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Guo</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Classification of human movements with and without spinal orthosis based on surface electromyogram signals</article-title>. <source>Medicine in Novel Technology and Devices</source><italic>,</italic> <volume>16</volume><italic>,</italic> <fpage>100165</fpage>. <pub-id pub-id-type="doi">10.1016/j.medntd.2022.100165</pub-id></mixed-citation></ref>
<ref id="ref-80"><label>80.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nunes</surname>, <given-names>U. M.</given-names></string-name>, <string-name><surname>Faria</surname>, <given-names>D. R.</given-names></string-name>, <string-name><surname>Peixoto</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A human activity recognition framework using max-min features and key poses with differential evolution random forests classifier</article-title>. <source>Pattern Recognition Letters</source><italic>,</italic> <volume>99</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>21</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2017.05.004</pub-id></mixed-citation></ref>
<ref id="ref-81"><label>81.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Imbeault-Nepton</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Maitre</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Bouchard</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Gaboury</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Filtering data bins of UWB radars for activity recognition with random forest</article-title>. <source>Procedia Computer Science</source><italic>,</italic> <volume>201</volume><italic>,</italic> <fpage>48</fpage>&#x2013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2022.03.009</pub-id></mixed-citation></ref>
<ref id="ref-82"><label>82.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Subedi</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Pradhananga</surname>, <given-names>N.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Sensor-based computational approach to preventing back injuries in construction workers</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>131</volume><italic>(</italic><issue>196</issue><italic>),</italic> <fpage>103920</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2021.103920</pub-id></mixed-citation></ref>
<ref id="ref-83"><label>83.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dimitrijevic</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Lepetit</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Fua</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2006</year>). <article-title>Human body pose detection using Bayesian spatio-temporal templates</article-title>. <source>Computer Vision and Image Understanding</source><italic>,</italic> <volume>104</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>127</fpage>&#x2013;<lpage>139</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2006.07.007</pub-id></mixed-citation></ref>
<ref id="ref-84"><label>84.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pajak</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Krutz</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Patalas-Maliszewska</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Rehm</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Pajak</surname>, <given-names>I.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>An approach to sport activities recognition based on an inertial sensor and deep learning</article-title>. <source>Sensors and Actuators A: Physical</source><italic>,</italic> <volume>345</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>113773</fpage>. <pub-id pub-id-type="doi">10.1016/j.sna.2022.113773</pub-id></mixed-citation></ref>
<ref id="ref-85"><label>85.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Singh</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Garg</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Hybrid machine learning algorithm for human activity recognition using decision tree and particle swarm optimization</article-title>. <source>International Journal of Engineering Science and Computing</source><italic>,</italic> <volume>6</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>8379</fpage>&#x2013;<lpage>8389</lpage>.</mixed-citation></ref>
<ref id="ref-86"><label>86.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Roan</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Smith</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Lockhart</surname>, <given-names>T. E.</given-names></string-name></person-group> (<year>2009</year>). <article-title>Gait analysis to classify external load conditions using linear discriminant analysis</article-title>. <source>Human Movement Science</source><italic>,</italic> <volume>28</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>226</fpage>&#x2013;<lpage>235</lpage>. <pub-id pub-id-type="doi">10.1016/j.humov.2008.10.008</pub-id>; <pub-id pub-id-type="pmid">19162355</pub-id></mixed-citation></ref>
<ref id="ref-87"><label>87.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Balaji</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Brindha</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Elumalai</surname>, <given-names>V. K.</given-names></string-name>, <string-name><surname>Umesh</surname>, <given-names>K.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Data-driven gait analysis for diagnosis and severity rating of Parkinson&#x2019;s disease</article-title>. <source>Medical Engineering &#x0026; Physics</source><italic>,</italic> <volume>91</volume><italic>(</italic><issue>21</issue><italic>),</italic> <fpage>54</fpage>&#x2013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.1016/j.medengphy.2021.03.005</pub-id>; <pub-id pub-id-type="pmid">34074466</pub-id></mixed-citation></ref>
<ref id="ref-88"><label>88.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wei</surname>, <given-names>S. -E.</given-names></string-name>, <string-name><surname>Ramakrishna</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Kanade</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Sheikh</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Convolutional pose machines</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>4724</fpage>&#x2013;<lpage>4732</lpage>. <publisher-loc>Las Vegas, NV, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-89"><label>89.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Newell</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Jia</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Stacked hourglass networks for human pose estimation</article-title>. <conf-name>European Conference on Computer Vision</conf-name>, pp. <fpage>483</fpage>&#x2013;<lpage>499</lpage>. <publisher-loc>Amsterdam, The Netherlands</publisher-loc>.</mixed-citation></ref>
<ref id="ref-90"><label>90.</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Yin</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Peng</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Du</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Rethinking on multi-stage networks for human pose estimation</article-title>. <comment>arXiv preprint arXiv:190100148</comment>.</mixed-citation></ref>
<ref id="ref-91"><label>91.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Peng</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>G.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2018</year>). <article-title>Cascaded pyramid network for multi-person pose estimation</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>7103</fpage>&#x2013;<lpage>7112</lpage>. <publisher-loc>Salt Lake City, UT, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-92"><label>92.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xiao</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Wu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Simple baselines for human pose estimation and tracking</article-title>. <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>, pp. <fpage>466</fpage>&#x2013;<lpage>481</lpage>. <publisher-loc>Munich, Germany</publisher-loc>.</mixed-citation></ref>
<ref id="ref-93"><label>93.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yan</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Coenen</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Driving posture recognition by convolutional neural networks</article-title>. <source>IET Computer Vision</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>103</fpage>&#x2013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1049/iet-cvi.2015.0175</pub-id></mixed-citation></ref>
<ref id="ref-94"><label>94.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>Q.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Application of human posture recognition based on the convolutional neural network in physical training guidance</article-title>. <source>Computational Intelligence and Neuroscience</source><italic>,</italic> <volume>2022</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1155/2022/5277157</pub-id>; <pub-id pub-id-type="pmid">35800679</pub-id></mixed-citation></ref>
<ref id="ref-95"><label>95.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yun</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Tao</surname>, <given-names>B.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>Grasping pose detection for loose stacked object based on convolutional neural network with multiple self-powered sensors information</article-title>. <source>IEEE Sensors Journal</source><italic>,</italic> <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1109/JSEN.2022.3190560</pub-id></mixed-citation></ref>
<ref id="ref-96"><label>96.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rani</surname>, <given-names>M. C. J.</given-names></string-name>, <string-name><surname>Devarakonda</surname>, <given-names>D. N.</given-names></string-name></person-group> (<year>2022</year>). <article-title>An effectual classical dance pose estimation and classification system employing convolution neural network&#x2013;long shortterm memory (CNN-LSTM) network for video sequences</article-title>. <source>Microprocessors and Microsystems</source><italic>,</italic> <volume>95</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>104651</fpage>. <pub-id pub-id-type="doi">10.1016/j.micpro.2022.104651</pub-id></mixed-citation></ref>
<ref id="ref-97"><label>97.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Zheng</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Gan</surname>, <given-names>H.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Automatic recognition of lactating sow postures by refined two-stream RGB-D faster R-CNN</article-title>. <source>Biosystems Engineering</source><italic>,</italic> <volume>189</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>116</fpage>&#x2013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1016/j.biosystemseng.2019.11.013</pub-id></mixed-citation></ref>
<ref id="ref-98"><label>98.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kharghanian</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Peiravi</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Moradi</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Iosifidis</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Pain detection using batch normalized discriminant restricted Boltzmann machine layers</article-title>. <source>Journal of Visual Communication and Image Representation</source><italic>,</italic> <volume>76</volume><italic>(</italic><issue>8</issue><italic>),</italic> <fpage>103062</fpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2021.103062</pub-id></mixed-citation></ref>
<ref id="ref-99"><label>99.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Indira</surname>, <given-names>D. N. V. S. L. S.</given-names></string-name>, <string-name><surname>Markapudi</surname>, <given-names>B. R.</given-names></string-name>, <string-name><surname>Chaduvula</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Jyothi Chaduvula</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Visual and buying sequence features-based product image recommendation using optimization based deep residual network</article-title>. <source>Gene Expression Patterns</source><italic>,</italic> <volume>45</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>119261</fpage>. <pub-id pub-id-type="doi">10.1016/j.gep.2022.119261</pub-id>; <pub-id pub-id-type="pmid">35817289</pub-id></mixed-citation></ref>
<ref id="ref-100"><label>100.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2019</year>). <article-title>MSST-ResNet: Deep multi-scale spatiotemporal features for robust visual object tracking</article-title>. <source>Knowledge-Based Systems</source><italic>,</italic> <volume>164</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>235</fpage>&#x2013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2018.10.044</pub-id></mixed-citation></ref>
<ref id="ref-101"><label>101.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Son</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Choi</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Seong</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Detection of construction workers under varying poses and changing background in image sequences via very deep residual networks</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>99</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>27</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2018.11.033</pub-id></mixed-citation></ref>
<ref id="ref-102"><label>102.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Xiao</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Deep high-resolution representation learning for human pose estimation</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>5693</fpage>&#x2013;<lpage>5703</lpage>. <publisher-loc>Long Beach, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-103"><label>103.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Cai</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Ju</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Deep cascaded convolutional models for cattle pose estimation</article-title>. <source>Computers and Electronics in Agriculture</source><italic>,</italic> <volume>164</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>104885</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2019.104885</pub-id></mixed-citation></ref>
<ref id="ref-104"><label>104.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Ren</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Gao</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>UULPN: An ultra-lightweight network for human pose estimation based on unbiased data processing</article-title>. <source>Neurocomputing</source><italic>,</italic> <volume>480</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>220</fpage>&#x2013;<lpage>233</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.12.083</pub-id></mixed-citation></ref>
<ref id="ref-105"><label>105.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Ling</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2021</year>). <article-title>ATT squeeze U-Net: A lightweight network for forest fire detection and recognition</article-title>. <source>IEEE Access</source><italic>,</italic> <volume>9</volume><italic>,</italic> <fpage>10858</fpage>&#x2013;<lpage>10870</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3050628</pub-id></mixed-citation></ref>
<ref id="ref-106"><label>106.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Xie</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Yao</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Q.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Research on a surface defect detection algorithm based on MobileNet-SSD</article-title>. <source>Applied Sciences</source><italic>,</italic> <volume>8</volume><italic>(</italic><issue>9</issue><italic>),</italic> <fpage>1678</fpage>. <pub-id pub-id-type="doi">10.3390/app8091678</pub-id></mixed-citation></ref>
<ref id="ref-107"><label>107.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Lin</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2018</year>). <article-title>ShuffleNet: An extremely efficient convolutional neural network for mobile devices</article-title>. <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>, pp. <fpage>6848</fpage>&#x2013;<lpage>6856</lpage>. <publisher-loc>Salt Lake City, UT, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-108"><label>108.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chollet</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Xception: Deep learning with depthwise separable convolutions</article-title>. <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>, pp. <fpage>1251</fpage>&#x2013;<lpage>1258</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-109"><label>109.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ma</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zheng</surname>, <given-names>H. T.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2018</year>). <article-title>ShuffleNet V2: Practical guidelines for efficient CNN architecture design</article-title>. <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>, pp. <fpage>116</fpage>&#x2013;<lpage>131</lpage>. <publisher-loc>Munich, Germany</publisher-loc>.</mixed-citation></ref>
<ref id="ref-110"><label>110.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>C. F.</given-names></string-name></person-group> (<year>2018</year>). <article-title>A basic introduction to separable convolutions</article-title>. <source>Towards Data Science</source>.</mixed-citation></ref>
<ref id="ref-111"><label>111.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tseng</surname>, <given-names>F. H.</given-names></string-name>, <string-name><surname>Yeh</surname>, <given-names>K. H.</given-names></string-name>, <string-name><surname>Kao</surname>, <given-names>F. Y.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>C. Y.</given-names></string-name></person-group> (<year>2022</year>). <article-title>MiniNet: Dense squeeze with depthwise separable convolutions for image classification in resource-constrained autonomous systems</article-title>. <source>ISA Transactions</source><italic>,</italic> <volume>20</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>273</fpage>. <pub-id pub-id-type="doi">10.1016/j.isatra.2022.07.030</pub-id>; <pub-id pub-id-type="pmid">36038366</pub-id></mixed-citation></ref>
<ref id="ref-112"><label>112.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>T. Y.</given-names></string-name>, <string-name><surname>Dollar</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Hariharan</surname>, <given-names>B.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Feature pyramid networks for object detection</article-title>. <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>, pp. <fpage>2117</fpage>&#x2013;<lpage>2125</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-113"><label>113.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Ioffe</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Szegedy</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2015</year>). <chapter-title>Batch normalization: Accelerating deep network training by reducing internal covariate shift</chapter-title>. <source>International Conference on Machine Learning</source>, pp. <fpage>448</fpage>&#x2013;<lpage>456</lpage>. <publisher-loc>Lille, France</publisher-loc>.</mixed-citation></ref>
<ref id="ref-114"><label>114.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>An</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Qian</surname>, <given-names>W.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Batch-normalized deep neural networks for achieving fast intelligent fault diagnosis of machines</article-title>. <source>Neurocomputing</source><italic>,</italic> <volume>329</volume><italic>,</italic> <fpage>53</fpage>&#x2013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2018.10.049</pub-id></mixed-citation></ref>
<ref id="ref-115"><label>115.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Ren</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>. <publisher-loc>Las Vegas, NV, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-116"><label>116.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Srivastava</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Hinton</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Krizhevsky</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sutskever</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Salakhutdinov</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2014</year>). <article-title>Dropout: A simple way to prevent neural networks from overfitting</article-title>. <source>Journal of Machine Learning Research</source><italic>,</italic> <volume>15</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1929</fpage>&#x2013;<lpage>1958</lpage>.</mixed-citation></ref>
<ref id="ref-117"><label>117.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hinton</surname>, <given-names>G. E.</given-names></string-name>, <string-name><surname>Srivastava</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Krizhevsky</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sutskever</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Salakhutdinov</surname>, <given-names>R. R.</given-names></string-name></person-group> (<year>2012</year>). <article-title>Improving neural networks by preventing co-adaptation of feature detectors</article-title>. <source>Computer Science</source><italic>,</italic> <volume>3</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>212</fpage>&#x2013;<lpage>223</lpage>.</mixed-citation></ref>
<ref id="ref-118"><label>118.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nair</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Hinton</surname>, <given-names>G. E.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Rectified linear units improve restricted boltzmann machines</article-title>. <conf-name>ICML</conf-name>, <publisher-loc>Haifa, Israel</publisher-loc>.</mixed-citation></ref>
<ref id="ref-119"><label>119.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Maas</surname>, <given-names>A. L.</given-names></string-name>, <string-name><surname>Hannun</surname>, <given-names>A. Y.</given-names></string-name>, <string-name><surname>Ng</surname>, <given-names>A. Y.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Rectifier nonlinearities improve neural network acoustic models</article-title>. <conf-name>Proceedings of ICML</conf-name>, vol. <comment>30</comment>, no. 1, pp. <fpage>3</fpage>. <publisher-loc>Atlanta, Georgia, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-120"><label>120.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Ren</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2015</year>). <article-title>Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>1026</fpage>&#x2013;<lpage>1034</lpage>. <publisher-loc>Santiago, Chile</publisher-loc>.</mixed-citation></ref>
<ref id="ref-121"><label>121.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ding</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Qian</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Activation functions and their characteristics in deep neural networks</article-title>. <conf-name>2018 Chinese Control and Decision Conference (CCDC)</conf-name>, pp. <fpage>1836</fpage>&#x2013;<lpage>1841</lpage>. <publisher-loc>Shenyang, China</publisher-loc>. <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-122"><label>122.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Weiss</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Khoshgoftaar</surname>, <given-names>T. M.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2016</year>). <article-title>A survey of transfer learning</article-title>. <source>Journal of Big Data</source><italic>,</italic> <volume>3</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1186/s40537-016-0043-6</pub-id></mixed-citation></ref>
<ref id="ref-123"><label>123.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hu</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Tang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Tang</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2020</year>). <article-title>A real-time patient-specific sleeping posture recognition system using pressure sensitive conductive sheet and transfer learning</article-title>. <source>IEEE Sensors Journal</source><italic>,</italic> <volume>21</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>6869</fpage>&#x2013;<lpage>6879</lpage>. <pub-id pub-id-type="doi">10.1109/JSEN.2020.3043416</pub-id></mixed-citation></ref>
<ref id="ref-124"><label>124.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ogundokun</surname>, <given-names>R. O.</given-names></string-name>, <string-name><surname>Maskeli&#x016B;nas</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Dama&#x0161;evi&#x010D;ius</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Human posture detection using image augmentation and hyperparameter-optimized transfer learning algorithms</article-title>. <source>Applied Sciences</source><italic>,</italic> <volume>12</volume><italic>(</italic><issue>19</issue><italic>),</italic> <fpage>10156</fpage>. <pub-id pub-id-type="doi">10.3390/app121910156</pub-id></mixed-citation></ref>
<ref id="ref-125"><label>125.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Long</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Jo</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Nam</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Development of a yoga posture coaching system using an interactive display based on transfer learning</article-title>. <source>The Journal of Supercomputing</source><italic>,</italic> <volume>78</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>5269</fpage>&#x2013;<lpage>5284</lpage>. <pub-id pub-id-type="doi">10.1007/s11227-021-04076-w</pub-id>; <pub-id pub-id-type="pmid">34566258</pub-id></mixed-citation></ref>
<ref id="ref-126"><label>126.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Lv</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Image recognition of wind turbine blade damage based on a deep learning model with transfer learning and an ensemble learning classifier</article-title>. <source>Renewable Energy</source><italic>,</italic> <volume>163</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>386</fpage>&#x2013;<lpage>397</lpage>. <pub-id pub-id-type="doi">10.1016/j.renene.2020.08.125</pub-id></mixed-citation></ref>
<ref id="ref-127"><label>127.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sagi</surname>, <given-names>O.</given-names></string-name>, <string-name><surname>Rokach</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Ensemble learning: A survey</article-title>. <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery</source><italic>,</italic> <volume>8</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>e1249</fpage>. <pub-id pub-id-type="doi">10.1002/widm.1249</pub-id></mixed-citation></ref>
<ref id="ref-128"><label>128.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Patil</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Patil</surname>, <given-names>K.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2016</year>). <article-title>Wearable sensor based human posture recognition</article-title>. <conf-name>2016 IEEE International Conference on Big Data (Big Data)</conf-name>, pp. <fpage>3432</fpage>&#x2013;<lpage>3438</lpage>. <publisher-loc>Washington DC, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-129"><label>129.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liang</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Cao</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Smart cushion: A practical system for fine-grained sitting posture recognition</article-title>. <conf-name>2017 IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops)</conf-name>, pp. <fpage>419</fpage>&#x2013;<lpage>424</lpage>. <publisher-loc>Kona, HI, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-130"><label>130.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Esmaeili</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>AkhavanPour</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Bosaghzadeh</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2020</year>). <article-title>An ensemble model for human posture recognition</article-title>. <conf-name>2020 International Conference on Machine Vision and Image Processing (MVIP)</conf-name>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>. <publisher-loc>Iran</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-131"><label>131.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Defferrard</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Bresson</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Vandergheynst</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2016</year>). <chapter-title>Convolutional neural networks on graphs with fast localized spectral filtering</chapter-title>. In: <source>Advances in neural information processing systems</source>.</mixed-citation></ref>
<ref id="ref-132"><label>132.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Research on multiplayer posture estimation technology of sports competition video based on graph neural network algorithm</article-title>. <source>Computational Intelligence and Neuroscience</source><italic>,</italic> <volume>2022</volume><italic>(</italic><issue>11</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1155/2022/4727375</pub-id>; <pub-id pub-id-type="pmid">35401733</pub-id></mixed-citation></ref>
<ref id="ref-133"><label>133.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Ling</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2021</year>). <article-title>PoGO-Net: Pose graph optimization with graph neural networks</article-title>. <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>, pp. <fpage>5895</fpage>&#x2013;<lpage>5905</lpage>. <publisher-loc>Montreal, QC, Canada</publisher-loc>.</mixed-citation></ref>
<ref id="ref-134"><label>134.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Taiana</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Toso</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>James</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Del Bue</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2022</year>). <article-title>PoserNet: Refining relative camera poses exploiting object detections</article-title>. <conf-name>European Conference on Computer Vision</conf-name>, pp. <fpage>247</fpage>&#x2013;<lpage>263</lpage>. <publisher-loc>Tel Aviv, Israel</publisher-loc>, <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-135"><label>135.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Person re-identification using heterogeneous local graph attention networks</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>12136</fpage>&#x2013;<lpage>12145</lpage>. <publisher-loc>Nashville, TN, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-136"><label>136.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zeng</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Lin</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Semantic relationships guided representation learning for facial action unit recognition</article-title>. <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, vol. <comment>1</comment>, pp. <fpage>8594</fpage>&#x2013;<lpage>8601</lpage>. <publisher-loc>Honolulu, Hawaii, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-137"><label>137.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Peng</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Tian</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Kapadia</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Metaxas</surname>, <given-names>D. N.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Semantic graph convolutional networks for 3D human pose regression</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>3425</fpage>&#x2013;<lpage>3435</lpage>. <publisher-loc>Long Beach, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-138"><label>138.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Fang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Dong</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Shuai</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>W.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>A survey on monocular 3D human pose estimation</article-title>. <source>Virtual Reality &#x0026; Intelligent Hardware</source><italic>,</italic> <volume>2</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>471</fpage>&#x2013;<lpage>500</lpage>. <pub-id pub-id-type="doi">10.1016/j.vrih.2020.04.005</pub-id></mixed-citation></ref>
<ref id="ref-139"><label>139.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Souza dos Reis</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Seewald</surname>, <given-names>L. A.</given-names></string-name>, <string-name><surname>Antunes</surname>, <given-names>R. S.</given-names></string-name>, <string-name><surname>Rodrigues</surname>, <given-names>V. F.</given-names></string-name>, <string-name><surname>da Rosa Righi</surname>, <given-names>R.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>Monocular multi-person pose estimation: A survey</article-title>. <source>Pattern Recognition</source><italic>,</italic> <volume>118</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>108046</fpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2021.108046</pub-id></mixed-citation></ref>
<ref id="ref-140"><label>140.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Desmarais</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Mottet</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Slangen</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Montesinos</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2021</year>). <article-title>A review of 3D human pose estimation algorithms for markerless motion capture</article-title>. <source>Computer Vision and Image Understanding</source><italic>,</italic> <volume>212</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>103275</fpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2021.103275</pub-id></mixed-citation></ref>
<ref id="ref-141"><label>141.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Ge</surname>, <given-names>S. S.</given-names></string-name></person-group> (<year>2021</year>). <article-title>A comprehensive survey on 2D multi-person pose estimation methods</article-title>. <source>Engineering Applications of Artificial Intelligence</source><italic>,</italic> <volume>102</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>104260</fpage>. <pub-id pub-id-type="doi">10.1016/j.engappai.2021.104260</pub-id></mixed-citation></ref>
<ref id="ref-142"><label>142.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Tan</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Zhen</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Zheng</surname>, <given-names>F.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>Deep 3D human pose estimation: A review</article-title>. <source>Computer Vision and Image Understanding</source><italic>,</italic> <volume>210</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>103225</fpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2021.103225</pub-id></mixed-citation></ref>
<ref id="ref-143"><label>143.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Song</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Yuan</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Human pose estimation and its application to action recognition: A survey</article-title>. <source>Journal of Visual Communication and Image Representation</source><italic>,</italic> <volume>76</volume><italic>,</italic> <fpage>103055</fpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2021.103055</pub-id></mixed-citation></ref>
<ref id="ref-144"><label>144.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Hussain</surname>, <given-names>A. J.</given-names></string-name>, <string-name><surname>Sadiq</surname>, <given-names>A. T.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Human body posture recognition approaches</article-title>. <source>ARO-The Scientific Journal of Koya University</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>75</fpage>&#x2013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.14500/aro.10930</pub-id></mixed-citation></ref>
<ref id="ref-145"><label>145.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ran</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Xiao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Gao</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>Z.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>A portable sitting posture monitoring system based on a pressure sensor array and machine learning</article-title>. <source>Sensors and Actuators A: Physical</source><italic>,</italic> <volume>331</volume><italic>,</italic> <fpage>112900</fpage>. <pub-id pub-id-type="doi">10.1016/j.sna.2021.112900</pub-id></mixed-citation></ref>
<ref id="ref-146"><label>146.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Vision-based framework for automatic interpretation of construction workers&#x2019; hand gestures</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>130</volume><italic>,</italic> <fpage>103872</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2021.103872</pub-id></mixed-citation></ref>
<ref id="ref-147"><label>147.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Saremi</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Mirjalili</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Lewis</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Vision-based hand posture estimation using a new hand model made of simple components</article-title>. <source>Optik</source><italic>,</italic> <volume>167</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>15</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijleo.2018.02.069</pub-id></mixed-citation></ref>
<ref id="ref-148"><label>148.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Chan</surname>, <given-names>A. B.</given-names></string-name></person-group> (<year>2014</year>). <article-title>3D human pose estimation from monocular images with deep convolutional neural network</article-title>. <conf-name>Asian Conference on Computer Vision</conf-name>, pp. <fpage>332</fpage>&#x2013;<lpage>347</lpage>. <publisher-loc>Singapore</publisher-loc>, <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-149"><label>149.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pavlakos</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Derpanis</surname>, <given-names>K. G.</given-names></string-name>, <string-name><surname>Daniilidis</surname>, <given-names>K.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Coarse-to-fine volumetric prediction for single-image 3D human pose</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>7025</fpage>&#x2013;<lpage>7034</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-150"><label>150.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ge</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Ren</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Xue</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>3D hand shape and pose estimation from a single RGB image</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>10833</fpage>&#x2013;<lpage>10842</lpage>. <publisher-loc>Long Beach, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-151"><label>151.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhou</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Xue</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Towards 3D human pose estimation in the wild: A weakly-supervised approach</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>398</fpage>&#x2013;<lpage>407</lpage>. <publisher-loc>Venice, Italy</publisher-loc>.</mixed-citation></ref>
<ref id="ref-152"><label>152.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Wu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Cheung</surname>, <given-names>Y. M.</given-names></string-name>, <string-name><surname>Guo</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Gao</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>A survey of human action recognition and posture prediction</article-title>. <source>Tsinghua Science and Technology</source><italic>,</italic> <volume>27</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>973</fpage>&#x2013;<lpage>1001</lpage>. <pub-id pub-id-type="doi">10.26599/TST.2021.9010068</pub-id></mixed-citation></ref>
<ref id="ref-153"><label>153.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pishchulin</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Insafutdinov</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Tang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Andres</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Andriluka</surname>, <given-names>M.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2016</year>). <article-title>Deepcut: Joint subset partition and labeling for multi person pose estimation</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>4929</fpage>&#x2013;<lpage>4937</lpage>. <publisher-loc>Las Vegas, NV, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-154"><label>154.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Carreira</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Agrawal</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Fragkiadaki</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Malik</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Human pose estimation with iterative error feedback</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>4733</fpage>&#x2013;<lpage>4742</lpage>. <publisher-loc>Las Vegas, NV, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-155"><label>155.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fang</surname>, <given-names>H. S.</given-names></string-name>, <string-name><surname>Xie</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Tai</surname>, <given-names>Y. W.</given-names></string-name>, <string-name><surname>Lu</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2017</year>). <article-title>RMPE: Regional multi-person pose estimation</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>2334</fpage>&#x2013;<lpage>2343</lpage>. <publisher-loc>Venice, Italy</publisher-loc>.</mixed-citation></ref>
<ref id="ref-156"><label>156.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Gkioxari</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Doll&#x00E1;r</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Mask R-CNN</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>2961</fpage>&#x2013;<lpage>2969</lpage>. <publisher-loc>Venice, Italy</publisher-loc>.</mixed-citation></ref>
<ref id="ref-157"><label>157.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Newell</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Deng</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2017</year>). <chapter-title>Associative embedding: End-to-end learning for joint detection and grouping</chapter-title>. In: <source>Advances in neural information processing systems</source>, <publisher-loc>Long Beach, California, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-158"><label>158.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Gong</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Tao</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A coarse-fine network for keypoint localization</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>3028</fpage>&#x2013;<lpage>3037</lpage>. <publisher-loc>Venice, Italy</publisher-loc>.</mixed-citation></ref>
<ref id="ref-159"><label>159.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Ouyang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Ma</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Yuille</surname>, <given-names>A. L.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Multi-context attention for human pose estimation</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>1831</fpage>&#x2013;<lpage>1840</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-160"><label>160.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kocabas</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Karagoz</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Akbas</surname>, <given-names>E.</given-names></string-name></person-group> (<year>2018</year>). <article-title>MultiPoseNet: Fast multi-person pose estimation using pose residual network</article-title>. <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>, pp. <fpage>417</fpage>&#x2013;<lpage>433</lpage>. <publisher-loc>Munich, Germany</publisher-loc>.</mixed-citation></ref>
<ref id="ref-161"><label>161.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kanazawa</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Black</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Jacobs</surname>, <given-names>D. W.</given-names></string-name>, <string-name><surname>Malik</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2018</year>). <article-title>End-to-end recovery of human shape and pose</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>7122</fpage>&#x2013;<lpage>7131</lpage>. <publisher-loc>Salt Lake City, UT, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-162"><label>162.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kreiss</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Bertoni</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Alahi</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2019</year>). <article-title>PifPaf: Composite fields for human pose estimation</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>11977</fpage>&#x2013;<lpage>11986</lpage>. <publisher-loc>Long Beach, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-163"><label>163.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>S. C.</given-names></string-name>, <string-name><surname>Tung</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2019</year>). <article-title>DenseRaC: Joint 3D pose and shape estimation by dense render-and-compare</article-title>. <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>, pp. <fpage>7760</fpage>&#x2013;<lpage>7770</lpage>. <publisher-loc>Seoul, Korea (South)</publisher-loc>.</mixed-citation></ref>
<ref id="ref-164"><label>164.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gujjar</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Vaughan</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Classifying pedestrian actions in advance using predicted video of urban driving scenes</article-title>. <conf-name>2019 International Conference on Robotics and Automation (ICRA)</conf-name>, pp. <fpage>2097</fpage>&#x2013;<lpage>2103</lpage>. <publisher-loc>Montreal, QC, Canada</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-165"><label>165.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Zeng</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Lai</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>Q.</given-names></string-name></person-group> (<year>2020</year>). <article-title>DeepFuse: An IMU-aware network for real-time 3D human pose estimation from multi-view image</article-title>. <conf-name>Workshop on Applications of Computer Vision</conf-name>, pp. <fpage>429</fpage>&#x2013;<lpage>438</lpage>. <publisher-loc>Snowmass, CO, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-166"><label>166.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhong</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Cao</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Pedestrian motion trajectory prediction with stereo-based 3D deep pose estimation and trajectory learning</article-title>. <source>IEEE Access</source><italic>,</italic> <volume>8</volume><italic>,</italic> <fpage>23480</fpage>&#x2013;<lpage>23486</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.2969994</pub-id></mixed-citation></ref>
<ref id="ref-167"><label>167.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cao</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Simon</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>S. E.</given-names></string-name>, <string-name><surname>Sheikh</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Realtime multi-person 2D pose estimation using part affinity fields</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>7291</fpage>&#x2013;<lpage>7299</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-168"><label>168.</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Fan</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Polarized self-attention: Towards high-quality pixel-wise regression</article-title>. <comment>arXiv preprint arXiv:210700782</comment>.</mixed-citation></ref>
<ref id="ref-169"><label>169.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Groos</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Ramampiaro</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Ihlen</surname>, <given-names>E. A.</given-names></string-name></person-group> (<year>2021</year>). <article-title>EfficientPose: Scalable single-person pose estimation</article-title>. <source>Applied Intelligence</source><italic>,</italic> <volume>51</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>2518</fpage>&#x2013;<lpage>2533</lpage>. <pub-id pub-id-type="doi">10.1007/s10489-020-01918-7</pub-id></mixed-citation></ref>
<ref id="ref-170"><label>170.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shan</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Lu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Gao</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Improving robustness and accuracy via relative information encoding in 3D human pose estimation</article-title>. <conf-name>Proceedings of the 29th ACM International Conference on Multimedia</conf-name>, <publisher-loc>Virtual Event, China</publisher-loc>.</mixed-citation></ref>
<ref id="ref-171"><label>171.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Reddy</surname>, <given-names>N. D.</given-names></string-name>, <string-name><surname>Guigues</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Pishchulin</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Eledath</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Narasimhan</surname>, <given-names>S. G.</given-names></string-name></person-group> (<year>2021</year>). <article-title>TesseTrack: End-to-end learnable multi-person articulated 3D pose tracking</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>15190</fpage>&#x2013;<lpage>15200</lpage>. <publisher-loc>Nashville, TN, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-172"><label>172.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yau</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Malekmohammadi</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Rasouli</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Lakner</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Rohani</surname>, <given-names>M.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>Graph-sim: A graph-based spatiotemporal interaction modelling for pedestrian action prediction</article-title>. <conf-name>2021 IEEE International Conference on Robotics and Automation (ICRA)</conf-name>, pp. <fpage>8580</fpage>&#x2013;<lpage>8586</lpage>. <publisher-loc>Xi'an, China</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-173"><label>173.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Johnson</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Everingham</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Clustered pose and nonlinear appearance models for human pose estimation</article-title>. <conf-name>BMVC</conf-name>, vol. <comment>4</comment>, pp. <fpage>5</fpage>. <publisher-loc>Aberystwyth, Wales, UK</publisher-loc>, <publisher-name>Citeseer</publisher-name>.</mixed-citation></ref>
<ref id="ref-174"><label>174.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Johnson</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Everingham</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Learning effective human pose estimation from inaccurate annotation</article-title>. <conf-name>CVPR 2011</conf-name>, pp. <fpage>1465</fpage>&#x2013;<lpage>1472</lpage>. <publisher-loc>Colorado Springs, CO, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-175"><label>175.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dantone</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Gall</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Leistner</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>van Gool</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Human pose estimation using body parts dependent joint regressors</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>3041</fpage>&#x2013;<lpage>3048</lpage>. <publisher-loc>Portland, OR, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-176"><label>176.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jhuang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Gall</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Zuffi</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Schmid</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Black</surname>, <given-names>M. J.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Towards understanding action recognition</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>3192</fpage>&#x2013;<lpage>3199</lpage>. <publisher-loc>Sydney, NSW, Australia</publisher-loc>.</mixed-citation></ref>
<ref id="ref-177"><label>177.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sapp</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Taskar</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Modec: Multimodal decomposable models for human pose estimation</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>3674</fpage>&#x2013;<lpage>3681</lpage>. <publisher-loc>Portland, OR, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-178"><label>178.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Derpanis</surname>, <given-names>K. G.</given-names></string-name></person-group> (<year>2013</year>). <article-title>From actemes to action: A strongly-supervised representation for detailed action understanding</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>2248</fpage>&#x2013;<lpage>2255</lpage>. <publisher-loc>Sydney, NSW, Australia</publisher-loc>.</mixed-citation></ref>
<ref id="ref-179"><label>179.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Andriluka</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Pishchulin</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Gehler</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Schiele</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2014</year>). <article-title>2D human pose estimation: New benchmark and state of the art analysis</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>3686</fpage>&#x2013;<lpage>3693</lpage>. <publisher-loc>Columbus, OH, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-180"><label>180.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>T. Y.</given-names></string-name>, <string-name><surname>Maire</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Belongie</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Hays</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Perona</surname>, <given-names>P.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2014</year>). <article-title>Microsoft coco: Common objects in context</article-title>. <conf-name>European Conference on Computer Vision</conf-name>, pp. <fpage>740</fpage>&#x2013;<lpage>755</lpage>. <publisher-loc>Zurich, Switzerland</publisher-loc>, <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-181"><label>181.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Zheng</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Zhao</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Yan</surname>, <given-names>B.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Large-scale datasets for going deeper in image understanding</article-title>. <conf-name>2019 IEEE International Conference on Multimedia and Expo (ICME)</conf-name>, pp. <fpage>1480</fpage>&#x2013;<lpage>1485</lpage>. <publisher-loc>Shanghai, China</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-182"><label>182.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Iqbal</surname>, <given-names>U.</given-names></string-name>, <string-name><surname>Milan</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Gall</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Posetrack: Joint multi-person pose estimation and tracking</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>2011</fpage>&#x2013;<lpage>2020</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-183"><label>183.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Andriluka</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Iqbal</surname>, <given-names>U.</given-names></string-name>, <string-name><surname>Insafutdinov</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Pishchulin</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Milan</surname>, <given-names>A.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2018</year>). <article-title>Posetrack: A benchmark for human pose estimation and tracking</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>5167</fpage>&#x2013;<lpage>5176</lpage>. <publisher-loc>Salt Lake City, UT, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-184"><label>184.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Mao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Fang</surname>, <given-names>H. S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Crowdpose: Efficient crowded scenes pose estimation and a new benchmark</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>10863</fpage>&#x2013;<lpage>10872</lpage>. <publisher-loc>Long Beach, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-185"><label>185.</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Qian</surname>, <given-names>R.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Human in events: A large-scale benchmark for human-centric video analysis in complex events</article-title>. <comment>arXiv preprint arXiv: 200504490</comment>.</mixed-citation></ref>
<ref id="ref-186"><label>186.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sigal</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Balan</surname>, <given-names>A. O.</given-names></string-name>, <string-name><surname>Black</surname>, <given-names>M. J.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion</article-title>. <source>International Journal of Computer Vision</source><italic>,</italic> <volume>87</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>4</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-009-0273-6</pub-id></mixed-citation></ref>
<ref id="ref-187"><label>187.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ionescu</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Papava</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Olaru</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Sminchisescu</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments</article-title>. <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source><italic>,</italic> <volume>36</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>1325</fpage>&#x2013;<lpage>1339</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2013.248</pub-id>; <pub-id pub-id-type="pmid">26353306</pub-id></mixed-citation></ref>
<ref id="ref-188"><label>188.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Joo</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Tan</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Gui</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Nabbe</surname>, <given-names>B.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2015</year>). <article-title>Panoptic studio: A massively multiview system for social motion capture</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>3334</fpage>&#x2013;<lpage>3342</lpage>. <publisher-loc>Santiago, Chile</publisher-loc>.</mixed-citation></ref>
<ref id="ref-189"><label>189.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Banks</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Flood</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2016</year>). <article-title>JointTrack auto: An open-source programme for automatic measurement of 3D implant kinematics from single-or bi-plane radiographic images</article-title>. <source>Orthopaedic Proceedings</source><italic>,</italic> vol. <volume>98</volume>, no. <supplement>SUPP_1</supplement>, pp. <fpage>38</fpage>. <comment>The British Editorial Society of Bone &#x0026; Joint Surgery</comment>.</mixed-citation></ref>
<ref id="ref-190"><label>190.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mehta</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Rhodin</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Casas</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Fua</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Sotnychenko</surname>, <given-names>O.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Monocular 3D human pose estimation in the wild using improved CNN supervision</article-title>. <conf-name>2017 International Conference on 3D Vision (3DV)</conf-name>, pp. <fpage>506</fpage>&#x2013;<lpage>516</lpage>. <publisher-loc>Qingdao, China</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-191"><label>191.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Varol</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Romero</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Martin</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Mahmood</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Black</surname>, <given-names>M. J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Learning from synthetic humans</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>109</fpage>&#x2013;<lpage>117</lpage>. <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-192"><label>192.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mehta</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Sotnychenko</surname>, <given-names>O.</given-names></string-name>, <string-name><surname>Mueller</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Sridhar</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2018</year>). <article-title>Single-shot multi-person 3D pose estimation from monocular RGB</article-title>. <conf-name>2018 International Conference on 3D Vision (3DV)</conf-name>, pp. <fpage>120</fpage>&#x2013;<lpage>130</lpage>. <publisher-loc>Verona, Italy</publisher-loc>, <publisher-name>IEEE</publisher-name>.</mixed-citation></ref>
<ref id="ref-193"><label>193.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Von Marcard</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Henschel</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Black</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Rosenhahn</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Pons-Moll</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Recovering accurate 3D human pose in the wild using IMUs and a moving camera</article-title>. <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>, pp. <fpage>601</fpage>&#x2013;<lpage>617</lpage>. <publisher-loc>Munich, Germany</publisher-loc>.</mixed-citation></ref>
<ref id="ref-194"><label>194.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mahmood</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Ghorbani</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Troje</surname>, <given-names>N. F.</given-names></string-name>, <string-name><surname>Pons-Moll</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Black</surname>, <given-names>M. J.</given-names></string-name></person-group> (<year>2019</year>). <article-title>AMASS: Archive of motion capture as surface shapes</article-title>. <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>, pp. <fpage>5442</fpage>&#x2013;<lpage>5451</lpage>. <publisher-loc>Seoul, Korea (South)</publisher-loc>.</mixed-citation></ref>
<ref id="ref-195"><label>195.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ghorbani</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Mahdaviani</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Thaler</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Kording</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Cook</surname>, <given-names>D. J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>MoVi: A large multi-purpose human motion and video dataset</article-title>. <source>PLoS One</source><italic>,</italic> <volume>16</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>e0253157</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0253157</pub-id>; <pub-id pub-id-type="pmid">34138926</pub-id></mixed-citation></ref>
<ref id="ref-196"><label>196.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>Y. D.</given-names></string-name>, <string-name><surname>Dong</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>S. H.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Yao</surname>, <given-names>X.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Advances in multimodal data fusion in neuroimaging: Overview, challenges, and novel orientation</article-title>. <source>Information Fusion</source><italic>,</italic> <volume>64</volume><italic>(</italic><issue>Suppl 3</issue><italic>),</italic> <fpage>149</fpage>&#x2013;<lpage>187</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2020.07.006</pub-id>; <pub-id pub-id-type="pmid">32834795</pub-id></mixed-citation></ref>
<ref id="ref-197"><label>197.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Celebi</surname>, <given-names>M. E.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Y. D.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Lu</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>Advances in data preprocessing for biomedical data fusion: An overview of the methods, challenges, and prospects</article-title>. <source>Information Fusion</source><italic>,</italic> <volume>76</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>376</fpage>&#x2013;<lpage>421</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2021.07.001</pub-id></mixed-citation></ref>
<ref id="ref-198"><label>198.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gravina</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Alinia</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Ghasemzadeh</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Fortino</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Multi-sensor fusion in body sensor networks: State-of-the-art and research challenges</article-title>. <source>Information Fusion</source><italic>,</italic> <volume>35</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>68</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2016.09.005</pub-id></mixed-citation></ref>
<ref id="ref-199"><label>199.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Guan</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Research on intelligent position posture detection and control based on multi-sensor fusion method</article-title>. <source>Journal of Physics: Conference Series</source><italic>,</italic> <volume>1</volume><italic>,</italic> <fpage>012078</fpage>. <comment>IOP Publishing</comment>.</mixed-citation></ref>
<ref id="ref-200"><label>200.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Luo</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Zheng</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Ke</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2019</year>). <article-title>A proactive workers&#x2019; safety risk evaluation framework based on position and posture data fusion</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>98</volume><italic>,</italic> <fpage>275</fpage>&#x2013;<lpage>288</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2018.11.026</pub-id></mixed-citation></ref>
<ref id="ref-201"><label>201.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wu</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2021</year>). <article-title>A multi-feature motion posture recognition model based on genetic algorithm</article-title>. <source>Traitement du Signal</source><italic>,</italic> <volume>38</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>599</fpage>&#x2013;<lpage>605</lpage>. <pub-id pub-id-type="doi">10.18280/ts.380307</pub-id></mixed-citation></ref>
<ref id="ref-202"><label>202.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zeng</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Che</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Convolutional neural network based multi-feature fusion for non-rigid 3D model retrieval</article-title>. <source>Journal of Information Processing Systems</source><italic>,</italic> <volume>14</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>176</fpage>&#x2013;<lpage>190</lpage>.</mixed-citation></ref>
<ref id="ref-203"><label>203.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yan</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Ergonomic posture recognition using 3D view-invariant features from single ordinary camera</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>94</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2018.05.033</pub-id></mixed-citation></ref>
<ref id="ref-204"><label>204.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chakraborty</surname>, <given-names>B. K.</given-names></string-name>, <string-name><surname>Sarma</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Bhuyan</surname>, <given-names>M. K.</given-names></string-name>, <string-name><surname>MacDorman</surname>, <given-names>K. F.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Review of constraints on vision-based gesture recognition for human-computer interaction</article-title>. <source>IET Computer Vision</source><italic>,</italic> <volume>12</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>3</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1049/iet-cvi.2017.0052</pub-id></mixed-citation></ref>
<ref id="ref-205"><label>205.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Angelini</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Fu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Long</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Shao</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Naqvi</surname>, <given-names>S. M.</given-names></string-name></person-group> (<year>2019</year>). <article-title>2D pose-based real-time human action recognition with occlusion-handling</article-title>. <source>IEEE Transactions on Multimedia</source><italic>,</italic> <volume>22</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>1433</fpage>&#x2013;<lpage>1446</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2019.2944745</pub-id></mixed-citation></ref>
<ref id="ref-206"><label>206.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Dong</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Fan</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2022</year>). <article-title>A survey on depth ambiguity of 3D human pose estimation</article-title>. <source>Applied Sciences</source><italic>,</italic> <volume>12</volume><italic>(</italic><issue>20</issue><italic>),</italic> <fpage>10591</fpage>. <pub-id pub-id-type="doi">10.3390/app122010591</pub-id></mixed-citation></ref>
<ref id="ref-207"><label>207.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Bao</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Mei</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Recent advances of monocular 2D and 3D human pose estimation: A deep learning perspective</article-title>. <source>ACM Computing Surveys</source><italic>,</italic> <volume>55</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>41</lpage>.</mixed-citation></ref>
<ref id="ref-208"><label>208.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zaka-Ud-Din</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Khan</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2022</year>). <article-title>A review of 3D human body pose estimation and mesh recovery</article-title>. <source>Digital Signal Processing</source><italic>,</italic> <volume>128</volume><italic>,</italic> <fpage>103628</fpage>.</mixed-citation></ref>
</ref-list>
</back>
</article>