<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">16871</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2021.016871</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Convolutional Bi-LSTM Based Human Gait Recognition Using Video Sequences</article-title>
<alt-title alt-title-type="left-running-head">Convolutional Bi-LSTM Based Human Gait Recognition Using Video Sequences</alt-title>
<alt-title alt-title-type="right-running-head">Convolutional Bi-LSTM Based Human Gait Recognition Using Video Sequences</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western">
<surname>Amin</surname>
<given-names>Javaria</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western">
<surname>Anjum</surname>
<given-names>Muhammad Almas</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western">
<surname>Sharif</surname>
<given-names>Muhammad</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western">
<surname>Kadry</surname>
<given-names>Seifedine</given-names>
</name>
<xref ref-type="aff" rid="aff-4">4</xref></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western">
<surname>Nam</surname>
<given-names>Yunyoung</given-names>
</name>
<xref ref-type="aff" rid="aff-5">5</xref><email>ynam@sch.ac.kr</email></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western">
<surname>Wang</surname>
<given-names>ShuiHua</given-names>
</name>
<xref ref-type="aff" rid="aff-6">6</xref></contrib>
<aff id="aff-1"><label>1</label><institution>University of Wah</institution>, <addr-line>Wah Cantt, 47040</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-2"><label>2</label><institution>National University of Technology (NUTECH)</institution>, <addr-line>Islamabad, 44000</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-3"><label>3</label><institution>COMSATS University Islamabad</institution>, <addr-line>Wah Campus, Wah Cantt</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-4"><label>4</label><institution>Faculty of Applied Computing and Technology, Noroff University College</institution>, <addr-line>Kristiansand</addr-line>, <country>Norway</country></aff>
<aff id="aff-5"><label>5</label><institution>Department of Computer Science and Engineering, Soonchunhyang University</institution>, <addr-line>Asan, 31538, Korea</addr-line></aff>
<aff id="aff-6"><label>6</label><institution>Department of Mathematics, University of Leicester</institution>, <addr-line>Leicester</addr-line>, <country>UK</country></aff>
</contrib-group>
<author-notes><corresp id="cor1">&#x002A;Corresponding Author: Yunyoung Nam. Email: <email>ynam@sch.ac.kr</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2021-02-24">
<day>24</day>
<month>02</month>
<year>2021</year>
</pub-date>
<volume>68</volume>
<issue>2</issue>
<fpage>2693</fpage>
<lpage>2709</lpage>
<history>
<date date-type="received">
<day>12</day>
<month>01</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>02</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2021 Amin et al.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Amin et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_16871.pdf"></self-uri>
<abstract>
<p>Recognition of human gait is a difficult assignment, particularly for unobtrusive surveillance in a video and human identification from a large distance. Therefore, a method is proposed for the classification and recognition of different types of human gait. The proposed approach is consisting of two phases. In phase I, the new model is proposed named convolutional bidirectional long short-term memory (Conv-BiLSTM) to classify the video frames of human gait. In this model, features are derived through convolutional neural network (CNN) named ResNet-18 and supplied as an input to the LSTM model that provided more distinguishable temporal information. In phase II, the YOLOv2-squeezeNet model is designed, where deep features are extricated using the fireconcat-02 layer and fed/passed to the tinyYOLOv2 model for recognized/localized the human gaits with predicted scores. The proposed method achieved up to 90% correct prediction scores on  CASIA-A,  CASIA-B, and the CASIA-C benchmark datasets. The proposed method achieved better/improved prediction scores as compared to the recent existing works.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Bi-LSTM</kwd>
<kwd>YOLOv2</kwd>
<kwd>open neural network</kwd>
<kwd>resNet-18</kwd>
<kwd>gait</kwd>
<kwd>squeezeNet</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Gait biometric represents a person walking styles and more powerful as compared to other biometrics [<xref ref-type="bibr" rid="ref-1">1</xref>] i.e., iris, palmprint, face, and fingerprint [<xref ref-type="bibr" rid="ref-2">2</xref>], etc. Therefore, it can be utilized for person identification from a long-distance [<xref ref-type="bibr" rid="ref-3">3</xref>]. Human gait with different styles is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Gait recognition methodologies have attained more attention in the last two decades in real-time applications such as forensic identification, video surveillance, and crime investigation [<xref ref-type="bibr" rid="ref-4">4</xref>]. In literature, some research works proposed improved feature vectors to discriminate the gait patterns based on the motion [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>]. The recognition of human body parts in motion is achieving more attention from researchers [<xref ref-type="bibr" rid="ref-8">8</xref>]. However, it is a more challenging and difficult task to accurately track each part of the human body [<xref ref-type="bibr" rid="ref-9">9</xref>]. The appearance-based gait recognition methodologies commonly utilized human silhouettes as input.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Human gait with different actions at 90<inline-formula id="ieqn-1"><!--<alternatives><inline-graphic xlink:href="ieqn-1.png"/>--><!--<tex-math id="tex-ieqn-1"><![CDATA[$^{\circ}$]]></tex-math>--><mml:math id="mml-ieqn-1"><mml:msup><mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math><!--</alternatives>--></inline-formula> (a) normal (b) human with a bag (c) human wearing a coat (d) woman (e) man</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-1.png"/>
</fig>
<p>These approaches might obtain maximum recognition scores when there is less variation in consecutive frames. When the variation increased in the consecutive frames, the performance of these algorithms decreased in real-time applications [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>The gait features are drastically changed in case of different variations i.e., illumination, view, clothing, and carrying [<xref ref-type="bibr" rid="ref-11">11</xref>]. Model-based features are utilized to track the human body parts and movement [<xref ref-type="bibr" rid="ref-12">12</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>]. The main contribution of the presented approach is based on feature vectors that are extracted from LSTM and ResNet-18 model. The extracted feature vectors contain more prominent discriminative information to classify the different types of human gaits based on fully connected and softmax layers. Furthermore, in phase II classified images are recognized using a proposed modified YOLOv2-ONNX model, which consists of 20 layers that are configured by applying the open neural network (ONNX) model and SqueezeNet architecture as the base-network of the tinyYOLOv2 model. The best recognition results are achieved by extracting deep features using the fireconcat-02 layer to the squeezeNet architecture and further fed as an input to the YOLOv2 model. The proposed method accurately recognizes the different kinds of human gaits.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Several machine learning approaches are used in the literature for human gait recognition (HGR) [<xref ref-type="bibr" rid="ref-15">15</xref>]. For HGR, features play a vital role to extract the discriminant information. Modified Local Optimal Oriented Pattern (MLOOP) features are extracted for HGR, and selected best features from MLOOP features vector [<xref ref-type="bibr" rid="ref-16">16</xref>]. The histogram oriented gradient (HOG) with Harlick features are combined for HGR and tested on the CASIA (A&#x2013;B) datasets [<xref ref-type="bibr" rid="ref-17">17</xref>]. The Gabor wavelet features are extracted from the input images in different orientations [<xref ref-type="bibr" rid="ref-18">18</xref>] for HGR. The method performance is computed on CASIA (A and B) datasets [<xref ref-type="bibr" rid="ref-19">19</xref>]. The multi-scale LBP and Gabor features are extracted and selected the best features by spectra discriminant analysis-based regression method [<xref ref-type="bibr" rid="ref-20">20</xref>&#x2013;<xref ref-type="bibr" rid="ref-22">22</xref>]. Principle component analysis (PCA) along with gait energy image (GEI) feature vectors are utilized for human identification [<xref ref-type="bibr" rid="ref-23">23</xref>]. However, it is difficult to recognize the variations in frames such as clothing, angle, and view [<xref ref-type="bibr" rid="ref-24">24</xref>]. To improve the recognition results, the fusion of structural gait profile and the energy shifted image is performed [<xref ref-type="bibr" rid="ref-25">25</xref>]. The deep features are extracted [<xref ref-type="bibr" rid="ref-26">26</xref>] using pre-trained AlexNet and VGG-19 and fused using skewness &#x0026; entropy. The informative features are selected by the FEcS method for HGR. The method is evaluated on CASIA A, B, and C datasets [<xref ref-type="bibr" rid="ref-27">27</xref>]. The gait flow image &#x0026; Gaussian image features are extracted to create a features vector and fed to the extended neural network classifier for HGR [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>]. The stacked progressive work autoencoders (SPAE) model is employed for gait recognition at different angles and views, in which some temporal information is missing [<xref ref-type="bibr" rid="ref-30">30</xref>]. GaitSet is applied for the extraction of invariant features for action recognition. The component-based frequency features are extracted for the identification of human actions [<xref ref-type="bibr" rid="ref-31">31</xref>]. The temporal features among the frame might obtain improved results as compared with the GEI [<xref ref-type="bibr" rid="ref-32">32</xref>]. However, classifying the cross-clothing and cross-carrying conditions is still a difficult activity due to changes in human shape and appearance [<xref ref-type="bibr" rid="ref-33">33</xref>]. Feng et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] extracted the heat map from the joint of the human body in an RGB input image instead of utilizing a binary silhouette. The extracted heat maps are supplied further to the LSTM model for temporal feature extraction. Recently, the skeleton and joints of the body are also utilized for the recognition of person identification [<xref ref-type="bibr" rid="ref-35">35</xref>]. It is observed that gait recognition with higher accuracy is still a challenging task [<xref ref-type="bibr" rid="ref-36">36</xref>].</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methodology</title>
<p>The proposed model contains two phases; robust feature extraction and classification is a challenging task for human gait recognition. Therefore, in phase I, the Conv-BiLSTM model is developed, in which deep features are extracted from the localized images using Resnet-18 and supplied to the LSTM network to classify the different types of human gaits. In phase II, input images are passed to the proposed YOLOv2-Squeeze model, which extracts deep from the fireconcat-02 layer of the squeeze-Net model and is supplied as an input to the tinyYOLOv2 model for localization/recognition of the different types of human gaits. The proposed model steps are displayed in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>General proposed approach steps</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-2.png"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Proposed Conv-BiLSTM Model for Classification of the Localized Images</title>
<p>The video frames are classified using the proposed Conv-BiLSTM model, in which deep features are extracted from the input frames by the CNN model such as Resnet18. Next, the sequence structures are restored and output is reshaped into sequence vectors using the unfolding sequence layer. After that, resultant vector sequences are created using BiLSTM and output layers. Finally, assembled both networks into a single network.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Convolutional Neural Network</title>
<p>The convolutional layers extract the feature vectors from the localized images. These feature vectors are used as the input of the activations function on the last pooling layer of the Resnet18 model as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. In the training phase, the model creates padding due to a large sequence of frames which has a negative impact on the accuracy of the gait classification.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Feature vectors extraction using Resnet-18 model</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-3.png"/>
</fig>
<p>To overcome this problem, the classification results are improved by removing the sequences with more than 600-time steps with class labels. The bar length of the histogram represents the selected sequences in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Visualization of the training data sequences</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-4.png"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Bidirectional Long Short-Term Memory (BiLSTM) Model</title>
<p>The modified BiLSTM model is used for the classification of human gaits, in which LSTM layers are used for more efficient temporal feature learning. The selection of hyperparameters for model training is done after the extensive experiment as given in <xref ref-type="table" rid="table-1">Tab. 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Experiment for parameter selection for model training</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Hidden units</th>
<th>Minibatch size</th>
<th>Accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td>1000</td>
<td>4</td>
<td>0.89</td>
</tr>
<tr>
<td>1500</td>
<td>8</td>
<td>0.94</td>
</tr>
<tr>
<td><bold><italic>2000</italic></bold></td>
<td><bold><italic>16</italic></bold></td>
<td><bold><italic>0.98</italic></bold></td>
</tr>
<tr>
<td>2500</td>
<td>32</td>
<td>0.94</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-1">Tab. 1</xref>, shows the experiment of the parameter&#x2019;s selection, where 2000 hidden Units, 16 batch size is used for the further experiment because increase/decrease the HU obtained accuracy is decreased. The Hyperparameters of the BiLSTM model are stated in <xref ref-type="table" rid="table-2">Tab. 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Selected hyperparameters of BiLSTM</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Hidden units</th>
<th>2000</th>
</tr>
</thead>
<tbody>
<tr>
<td>Dropout</td>
<td>0.5</td>
</tr>
<tr>
<td>MiniBatchSize</td>
<td>16</td>
</tr>
<tr>
<td>Gradient threshold</td>
<td>2</td>
</tr>
<tr>
<td>InitialLearnRate</td>
<td>1e-4</td>
</tr>
<tr>
<td>Maximum epochs</td>
<td>30</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The model specification is as: Sequence input (1024 dimensions), LSTM layers (2000 hidden units (HU)), 50% dropout, fully connected layers, softmax, and a classification layer. The activation functions of the proposed BiLSTM model are mentioned in <xref ref-type="table" rid="table-3">Tab. 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>BiLSTM layers with corresponding activations</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Layers</th>
<th>Activation functions</th>
<th>Learnable</th>
</tr>
</thead>
<tbody>
<tr>
<td>Input sequence</td>
<td>1024</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>BiLSTM</td>
<td>4000</td>
<td><inline-formula id="ieqn-2"><!--<alternatives><inline-graphic xlink:href="ieqn-2.png"/>--><!--<tex-math id="tex-ieqn-2"><![CDATA[$\text{W}\text{e}\text{i}\text{g}\text{h}\text{ts}= 16000\times 1024$]]></tex-math>--><mml:math id="mml-ieqn-2"><mml:mstyle class="text"><mml:mtext>W</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>e</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>i</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>g</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>h</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>ts</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mn>16000</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1024</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td/>
<td/>
<td>Recurrent <inline-formula id="ieqn-3"><!--<alternatives><inline-graphic xlink:href="ieqn-3.png"/>--><!--<tex-math id="tex-ieqn-3"><![CDATA[$\text{b}\text{i}\text{a}\text{s}= 16000 \times 2000$]]></tex-math>--><mml:math id="mml-ieqn-3"><mml:mstyle class="text"><mml:mtext>b</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>i</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>a</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>s</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mn>16000</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>2000</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td/>
<td/>
<td><inline-formula id="ieqn-4"><!--<alternatives><inline-graphic xlink:href="ieqn-4.png"/>--><!--<tex-math id="tex-ieqn-4"><![CDATA[$\text{B}\text{i}\text{a}\text{s}= 16000\times 1$]]></tex-math>--><mml:math id="mml-ieqn-4"><mml:mstyle class="text"><mml:mtext>B</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>i</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>a</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>s</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mn>16000</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>Drop</td>
<td>4000</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>FC</td>
<td>3</td>
<td><inline-formula id="ieqn-5"><!--<alternatives><inline-graphic xlink:href="ieqn-5.png"/>--><!--<tex-math id="tex-ieqn-5"><![CDATA[$\text{W}\text{e}\text{i}\text{g}\text{h}\text{ts}= 3\times 4000$]]></tex-math>--><mml:math id="mml-ieqn-5"><mml:mstyle class="text"><mml:mtext>W</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>e</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>i</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>g</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>h</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>ts</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>4000</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td/>
<td/>
<td><inline-formula id="ieqn-6"><!--<alternatives><inline-graphic xlink:href="ieqn-6.png"/>--><!--<tex-math id="tex-ieqn-6"><![CDATA[$\text{B}\text{i}\text{a}\text{s}= 3\times 1$]]></tex-math>--><mml:math id="mml-ieqn-6"><mml:mstyle class="text"><mml:mtext>B</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>i</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>a</mml:mtext></mml:mstyle><mml:mstyle class="text"><mml:mtext>s</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>Softmax</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Classification</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The LSTM [<xref ref-type="bibr" rid="ref-37">37</xref>] cell has four gates, i.e., input, forget, output gate, and cell candidate. In the LSTM block, three weights are learnable, i.e., input f, recurrent weights <italic>R<sub>W</sub></italic>, and bias b. The matrices of the learnable weights are expressed mathematically as:</p>
<p><disp-formula id="eqn-1">
<label>(1)</label>
<!--<alternatives><graphic mimetype="image" mime-subtype="png" xlink:href="eqn-1.png"/>-->
<!--<tex-math id="tex-eqn-1"><![CDATA[$$\begin{equation}
f_{W}= \left[\begin{array}{l}
{f_{W}}_{\mathit{input}} \\[6pt]
{f_{W}}_{\mathit{forget}} \\[6pt]
{f_{W}}_{cell \, \mathit{candidate}} \\[6pt]
{f_{W}}_{ \mathit{Output}}\end{array}\right],\quad \mathit{Recurrent}_{W}= \left[\begin{array}{l}
{\mathit{Recurrent}_{W}}_{\mathit{input}} \\[6pt] {\mathit{Recurrent}_{W}}_{\mathit{forget}} \\[6pt]
{\mathit{Recurrent}_{W}}_{cell \, \mathit{candidate}} \\[6pt]
{\mathit{Recurrent}_{W}}_{\mathit{Output}}\end{array}\right],\quad
b= \left[\begin{array}{l}
b_{\mathit{input}} \\[6pt]
b_{\mathit{forget}} \\[6pt]
b_{cell\, \mathit{candidate}} \\[6pt]
b_{\mathit{Output}}\end{array}\right]
 \label{eqn-1}
\end{equation}$$]]></tex-math>-->
<mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable equalrows="false" columnlines="" equalcolumns="false"><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="0.3em"/><mml:mstyle mathvariant="italic"><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>O</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr> </mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em"/><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable equalrows="false" columnlines="" equalcolumns="false"><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="0.3em"/><mml:mstyle mathvariant="italic"><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>O</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr> </mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em"/><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable equalrows="false" columnlines="" equalcolumns="false"><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="0.3em"/><mml:mstyle mathvariant="italic"><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>O</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mtd></mml:mtr> </mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math><!--</alternatives>--></disp-formula></p>
<p>The cell state <inline-formula id="ieqn-7"><!--<alternatives><inline-graphic xlink:href="ieqn-7.png"/>--><!--<tex-math id="tex-ieqn-7"><![CDATA[$\mathrm{c}_{\mathrm{t}}$]]></tex-math>--><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>c</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:math><!--</alternatives>--></inline-formula> at the time step (t) is written as:</p>
<p><disp-formula id="eqn-2">
<label>(2)</label>
<!--<alternatives><graphic mimetype="image" mime-subtype="png" xlink:href="eqn-2.png"/>-->
<!--<tex-math id="tex-eqn-2"><![CDATA[$$\begin{equation}
\text{cell state}_{\mathrm{t}}=F_{t}\odot cell~ \mathit{state}_{t-1}+I_{t}\odot cell~ \mathit{candidate}_{t}
 \label{eqn-2}
\end{equation}$$]]></tex-math>-->
<mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mstyle><mml:mtext>cell&#x00A0;state</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width=".3em" /><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo lspace='0pt' rspace='0pt'>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width=".3em" /><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math><!--</alternatives>--></disp-formula></p>
<p>where <inline-formula id="ieqn-8"><!--<alternatives><inline-graphic xlink:href="ieqn-8.png"/>--><!--<tex-math id="tex-ieqn-8"><![CDATA[$\odot$]]></tex-math>--><mml:math id="mml-ieqn-8"><mml:mo>&#x2299;</mml:mo></mml:math><!--</alternatives>--></inline-formula> denotes Hadamard product. The hidden state <inline-formula id="ieqn-9"><!--<alternatives><inline-graphic xlink:href="ieqn-9.png"/>--><!--<tex-math id="tex-ieqn-9"><![CDATA[$\mathrm{h}_{\mathrm{t}}$]]></tex-math>--><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>h</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:math><!--</alternatives>--></inline-formula> is represented as:</p>
<p><disp-formula id="eqn-3">
<label>(3)</label>
<!--<alternatives><graphic mimetype="image" mime-subtype="png" xlink:href="eqn-3.png"/>-->
<!--<tex-math id="tex-eqn-3"><![CDATA[$$\begin{equation}
\text{hidden state}_{\mathrm{t}}=\mathit{Output}_{t}\odot tanh (cell~ \mathit{state}_{t})
 \label{eqn-3}
\end{equation}$$]]></tex-math>-->
<mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:mstyle><mml:mtext>hidden&#x00A0;state</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>O</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width=".3em" /><mml:msub><mml:mrow><mml:mstyle mathvariant="italic"><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math><!--</alternatives>--></disp-formula></p>
<p>In the LSTM model, based on time steps, feature vectors are computed through LSTM layers and supplied to the next block. The <italic>n</italic>th block output is used for the class label prediction, in which HU follows the fully connected, softmax, and the output layers.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Concatenation of CNN and LSTM Models</title>
<p>In the proposed model, LSTM layers are concatenated with CNN layers, in which frames are transformed into a sequence of vectors to classify the human gaits. <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, shows the steps of the assembled network.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Proposed Conv-BiLSTM model</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-5.png"/>
</fig>
<p>In <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, input sequences are passed to the convolutional layers, where features are extracted by convolutional operators. The convolutional layers follow the sequence folding layer. The sequence unfolding layer is followed by the flatten layer in which the structure of the sequences is restored and output is reshaped into a vector. The gait classification is performed using the output of BiLSTM followed by fully connected and softmax layers.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Localization of Human Gait UsingYOLOv2-SqueezeNet Model</title>
<p>YOLOv2 is fast and effective as compared with recurrent neural network (RCNN) and SSD detectors. Therefore, in this research, YOLOv2-SqueezeNet model is suggested for different types of human gait localization such as female, male, fast walk, slow walk, walk with the bag, normal, and wearing as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>YOLOv2-SqueezeNet model for localization</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-6.png"/>
</fig>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref>, shows proposed YOLOv2-SqueezeNet model, where features are extracted from fireconcat-02 layer of the SqueezeNet model and passed as an input to the pre-train YOLOv2 detector. The proposed model more accurately localized the required regions with class labels.</p>
<p>The model consists of the 20 layers in which 01 input image, 04 convolutional, 04 ReLU, 01 depth concatenation, 01 max-pooling, of the squeezeNet model, and 02 YOLOv2 convolutional, 02 YOLOv2 batch-normalization, 02 YOLOv2ReLU, 01 YOLOv2 transforms, and 01 YOLOv2-output of the YOLOv2 model. The activation functions of the YOLOv2-SqueezeNet model are shown in <xref ref-type="table" rid="table-4">Tab. 4</xref>. <xref ref-type="table" rid="table-5">Tab. 5</xref>, presents the training hyperparameters.</p>
<table-wrap id="table-4"> 
<label>Table 4</label>
<caption>
<title>Layer wise activations of YOLOv2-SqueezeNet model</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Layers</th>
<th>Activations</th>
</tr>
</thead>
<tbody>
<tr>
<td>Data (input image)</td>
<td><inline-formula id="ieqn-10"><!--<alternatives><inline-graphic xlink:href="ieqn-10.png"/>--><!--<tex-math id="tex-ieqn-10"><![CDATA[$ 130 \times 130 \times 3 $]]></tex-math>--><mml:math id="mml-ieqn-10"><mml:mn>130</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>130</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>Conv</td>
<td><inline-formula id="ieqn-11"><!--<alternatives><inline-graphic xlink:href="ieqn-11.png"/>--><!--<tex-math id="tex-ieqn-11"><![CDATA[$ 64 \times 64 \times 64 $]]></tex-math>--><mml:math id="mml-ieqn-11"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>ReLU</td>
<td><inline-formula id="ieqn-12"><!--<alternatives><inline-graphic xlink:href="ieqn-12.png"/>--><!--<tex-math id="tex-ieqn-12"><![CDATA[$ 64 \times 64 \times 64 $]]></tex-math>--><mml:math id="mml-ieqn-12"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>Max pool</td>
<td><inline-formula id="ieqn-13"><!--<alternatives><inline-graphic xlink:href="ieqn-13.png"/>--><!--<tex-math id="tex-ieqn-13"><![CDATA[$ 31 \times 31 \times 64 $]]></tex-math>--><mml:math id="mml-ieqn-13"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>Fire2-squeeze (convolution)</td>
<td><inline-formula id="ieqn-14"><!--<alternatives><inline-graphic xlink:href="ieqn-14.png"/>--><!--<tex-math id="tex-ieqn-14"><![CDATA[$ 31 \times 31 \times 16 $]]></tex-math>--><mml:math id="mml-ieqn-14"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>16</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>Fire2-squeeze (ReLU)</td>
<td><inline-formula id="ieqn-15"><!--<alternatives><inline-graphic xlink:href="ieqn-15.png"/>--><!--<tex-math id="tex-ieqn-15"><![CDATA[$ 31 \times 31 \times 16 $]]></tex-math>--><mml:math id="mml-ieqn-15"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>16</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>02-Fire2-expend (convolution)</td>
<td><inline-formula id="ieqn-16"><!--<alternatives><inline-graphic xlink:href="ieqn-16.png"/>--><!--<tex-math id="tex-ieqn-16"><![CDATA[$ 31 \times 31 \times 64 $]]></tex-math>--><mml:math id="mml-ieqn-16"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>02-Fire2-expend (ReLU)</td>
<td><inline-formula id="ieqn-17"><!--<alternatives><inline-graphic xlink:href="ieqn-17.png"/>--><!--<tex-math id="tex-ieqn-17"><![CDATA[$ 31 \times 31 \times 64 $]]></tex-math>--><mml:math id="mml-ieqn-17"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>64</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>Fire2-concatenation</td>
<td><inline-formula id="ieqn-18"><!--<alternatives><inline-graphic xlink:href="ieqn-18.png"/>--><!--<tex-math id="tex-ieqn-18"><![CDATA[$ 31 \times 31 \times 128 $]]></tex-math>--><mml:math id="mml-ieqn-18"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>02-YOLOv2 (convolution)</td>
<td><inline-formula id="ieqn-19"><!--<alternatives><inline-graphic xlink:href="ieqn-19.png"/>--><!--<tex-math id="tex-ieqn-19"><![CDATA[$ 31 \times 31\times 128 $]]></tex-math>--><mml:math id="mml-ieqn-19"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>02-YOLOv2 (ReLU)</td>
<td><inline-formula id="ieqn-20"><!--<alternatives><inline-graphic xlink:href="ieqn-20.png"/>--><!--<tex-math id="tex-ieqn-20"><![CDATA[$ 31\times 31\times 128 $]]></tex-math>--><mml:math id="mml-ieqn-20"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>02-YOLOv2 (batch-normalization)</td>
<td><inline-formula id="ieqn-21"><!--<alternatives><inline-graphic xlink:href="ieqn-21.png"/>--><!--<tex-math id="tex-ieqn-21"><![CDATA[$ 31 \times 31 \times 128 $]]></tex-math>--><mml:math id="mml-ieqn-21"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>YOLOv2-classification</td>
<td><inline-formula id="ieqn-22"><!--<alternatives><inline-graphic xlink:href="ieqn-22.png"/>--><!--<tex-math id="tex-ieqn-22"><![CDATA[$ 31\times 31\times 2 $]]></tex-math>--><mml:math id="mml-ieqn-22"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>YOLOv2-transform</td>
<td><inline-formula id="ieqn-23"><!--<alternatives><inline-graphic xlink:href="ieqn-23.png"/>--><!--<tex-math id="tex-ieqn-23"><![CDATA[$ 31\times 31\times 2 $]]></tex-math>--><mml:math id="mml-ieqn-23"><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>31</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn></mml:math><!--</alternatives>--></inline-formula></td>
</tr>
<tr>
<td>YOLOv2-output</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
 
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Hyperparameters of YOLOv2-SqueezeNet model</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Size of the mini-batch</th>
<th>14</th>
</tr>
</thead>
<tbody>
<tr>
<td>Epochs</td>
<td>1000</td>
</tr>
<tr>
<td>Frequency validation</td>
<td>50</td>
</tr>
<tr>
<td>Rate of the learning</td>
<td>0.001</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-5">Tab. 5</xref> presents the hyperparameters that are selected to configure the proposed model for human gait classification, in which mini-batch size is selected 14, 1000 epochs are used for model training because greater than equal to the 1000 epochs model results are consistent.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Setup</title>
<p>Gait recognition is a great challenge due to complex recognition patterns that have been utilized in different fields such as machine learning, robotics, studying, biomedical, visual surveillance, and forensic. Therefore, intelligent recognition and the digital security group designed CASIA (A, B &#x0026; C) datasets in the national pattern recognition laboratory [<xref ref-type="bibr" rid="ref-38">38</xref>&#x2013;<xref ref-type="bibr" rid="ref-44">44</xref>].</p>
<p>The presented study is implemented on Matlab 2020RA Toolbox using a Core-i7 desktop Computer with a 740 K Nvidia Graphic Card. 0.5 hold out validation is used for model training. The description of the number of training and testing images are mentioned in <xref ref-type="table" rid="table-6">Tab. 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Description of training and testing number images in the corresponding datasets</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Datasets</th>
<th>Number of images in training and testing</th>
</tr>
</thead>
<tbody>
<tr>
<td>CASIA-A</td>
<td>Total number of: 19139</td>
</tr>
<tr>
<td/>
<td>Training images: 9569</td>
</tr>
<tr>
<td/>
<td>Testing images: 9569</td>
</tr>
<tr>
<td>CASIA-B</td>
<td>Total number of: 51036</td>
</tr>
<tr>
<td/>
<td>Training images: 22518</td>
</tr>
<tr>
<td/>
<td>Testing images: 22518</td>
</tr>
<tr>
<td>CASIA-C</td>
<td>Total number of: 91800</td>
</tr>
<tr>
<td/>
<td>Training images: 45900</td>
</tr>
<tr>
<td/>
<td>Testing images: 45900</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s4_1">
<label>4.1</label>
<title>Results and Discussion</title>
<p>In the developed framework, implement two experiments for the analysis of the proposed approach performance. The first experiment is performed to compute the performance of the YOLOv2-ONNX model and the second experiment is performed for classification results.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experiment #01</title>
<p>In this experiment, extracted feature vectors using the Conv-BiLSTM model are passed to the softmax layer for the classification of different types of human gaits such as female/male, bag, wearing, normal, and fast walk, slow walk, normal walk classes of the CASIA-A, CASIA-B and CASIA-C datasets respectively. <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, represents the proposed approach performance.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Training/testing results with respective loss rate (a) CASIA-A (b) CASIA-B (c) CASIA-C (blue line shows training, red shows loss rate, and dotted black line represent validation accuracy)</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-7.png"/>
</fig>
<p>In <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, the proposed model achieved 1.00 validation accuracy (VA) on CASIA-A and CASIA-C datasets, while it achieved 0.96 VA on the CASIA-B dataset.</p>
<p>The classification outcomes are stated in the <xref ref-type="table" rid="table-7">Tabs. 7</xref>&#x2013;<xref ref-type="table" rid="table-9">9</xref>.</p>
<p><xref ref-type="table" rid="table-7">Tab. 7</xref>, shows experimental results on CASIA-A dataset proposed method achieves 1.00 CPR on two classes of female/male.</p>
<table-wrap id="table-7"> 
<label>Table 7</label>
<caption>
<title>Proposed method results for human gaits recognition on different datasets using CASIA-A dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Classes</th>
<th>SE</th>
<th>SP</th>
<th>PPV</th>
<th>FPR</th>
<th>FNR</th>
<th>NPV</th>
<th>CPR</th>
</tr>
</thead>
<tbody>
<tr>
<td>CASIA-A</td>
<td>Female</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td/>
<td>Male</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-8">Tab. 8</xref>, CASIA-B dataset is considered for performance evaluation, where three classes such as Bag, wearing, and normal are involved. The method achieved 0.92 CPR in bag class, 1.00 CPR on wearing, and 0.88 CPR in the normal class.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Proposed method results for human gaits recognition on different datasets using the CASIA-B dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Classes</th>
<th>SE</th>
<th>SP</th>
<th>PPV</th>
<th>FPR</th>
<th>FNR</th>
<th>NPV</th>
<th>CPR</th>
</tr>
</thead>
<tbody>
<tr>
<td>CASIA-B</td>
<td>Bag</td>
<td>0.94</td>
<td>0.86</td>
<td>0.94</td>
<td>0.13</td>
<td>0.05</td>
<td>0.86</td>
<td>0.92</td>
</tr>
<tr>
<td/>
<td>Wearing</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
</tr>
<tr>
<td/>
<td>Normal</td>
<td>0.91</td>
<td>0.81</td>
<td>0.92</td>
<td>0.18</td>
<td>0.08</td>
<td>0.78</td>
<td>0.88</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The evaluation results in <xref ref-type="table" rid="table-9">Tab. 9</xref> shows that, the proposed method achieved 1.00 CPR on the classes of CASIA-C dataset. The outcomes in <xref ref-type="table" rid="table-7">Tabs. 7</xref>&#x2013;<xref ref-type="table" rid="table-9">9</xref>, depicts that the proposed model obtained a 1.00 correct recognition rate (CPR). The recognition outcomes on the CASIA-B dataset are 0.92 CPR on humans with the bag, 1.00 CPR on wearing class, and 0.88 CPR on a normal class. The predicted labels of human gait recognition are shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The proposed approach comparison is mentioned in <xref ref-type="table" rid="table-10">Tab. 10</xref>.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Proposed method results for human gaits recognition on different datasets using the CASIA-C dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Classes</th>
<th>SE</th>
<th>SP</th>
<th>PPV</th>
<th>FPR</th>
<th>FNR</th>
<th>NPV</th>
<th>CPR</th>
</tr>
</thead>
<tbody>
<tr>
<td>CASIA-C</td>
<td>Slow walk</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td/>
<td>Fast walk</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td/>
<td>Walk with bag</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Predicted labels on benchmark datasets</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-8.png"/>
</fig>
 
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Proposed approach results compared with recent approaches</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Refs.</th>
<th>YEAR</th>
<th>CASIA-A</th>
<th>CASIA-B</th>
<th></th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>2020</td>
<td>0.95 CPR</td>
<td>0.92 CPR</td>
<td/>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>2019</td>
<td>&#x2013;</td>
<td>0.95 CPR</td>
<td/>
</tr>
<tr>
<td>Proposed method</td>
<td>2020</td>
<td>1.00 CPR</td>
<td>0.96 CPR in which 0.88 (human with bag) and 0.92 (normal)</td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
<p>Six recent states of the art approaches are considered for performance evaluation based on some benchmark datasets. In the comparison scenario, the experimental setup is also discussed for existing work with proposed work. Wang et al. [<xref ref-type="bibr" rid="ref-45">45</xref>] used an ensemble learning method for human gait classification on CASIA-A &#x0026; CASIA-B datasets and achieved results are 0.95 and 0.92 CPR respectively. Wang et al. [<xref ref-type="bibr" rid="ref-46">46</xref>] utilized the LSTM model to learn the sequential patterns of the input images and achieved 0.95 CPR on the CASIA-B dataset. The results in <xref ref-type="table" rid="table-11">Tab. 11</xref>, are compared with the latest methodologies which show the proposed approach performance is superior. The proposed model results are better because of strongest feature vectors are obtained using the Conv-BiLSTM model for the classification of different types of human gaits with maximum CPR and also provided good results on a limited range of the input videos.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Localization results</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Datasets</th>
<th>Classes</th>
<th>mAP</th>
</tr>
</thead>
<tbody>
<tr>
<td>CASIA-A</td>
<td>Female</td>
<td>1.00</td>
</tr>
<tr>
<td/>
<td>Male</td>
<td>0.82</td>
</tr>
<tr>
<td>CASIA-B</td>
<td>Bag</td>
<td>1.00</td>
</tr>
<tr>
<td/>
<td>Wearing</td>
<td>0.91</td>
</tr>
<tr>
<td/>
<td>Normal</td>
<td>1.00</td>
</tr>
<tr>
<td>CASIA-C</td>
<td>Fast walk</td>
<td>1.00</td>
</tr>
<tr>
<td/>
<td>Slow walk</td>
<td>0.70</td>
</tr>
<tr>
<td/>
<td>Normal walk</td>
<td>0.95</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Localization results in term of mAP and IoU (a) CASIA-C (b) CASIA-B (c) CASIA-A (d) IoU</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-9.png"/>
</fig>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Gait localization (a, d and g) original gait images (b, e and h) gait labels (c, f and i) prediction scores</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-10.png"/>
</fig>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Gait localization (a, d) original gait images (b, e) gait labels (c, f) prediction scores</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-11.png"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experiment #02</title>
<p>The proposed YOLOv2-ONNX model is validated on CASIA-A, CASIA-B, and CASIA-C in terms of mean average precision (mAP) as mentioned in <xref ref-type="table" rid="table-11">Tab. 11</xref>. The localization outcome according to the respective class labels is graphically depicted in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. <xref ref-type="table" rid="table-11">Tab. 11</xref>, shows the proposed approach obtained mAP of 1.00, 0.91, and 1.00 on different classes such as Bag, wearing, and normal of the CASIA-B dataset respectively.</p>
<p>On different classes of the CASIA-C dataset i.e., fast walk, slow walk, and normal walk achieved mAP is 1.00, 070, and 0.95 respectively, where on the CASIA-A dataset attained mAP is 1.00 and 0.822 on female and male classes respectively. The proposed method more precisely localizes the different types of human gaits as illustrated in <xref ref-type="fig" rid="fig-10">Figs. 10</xref>&#x2013;<xref ref-type="fig" rid="fig-12">12</xref>.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Gait localization (a, d) original gait images (b, e) gait labels (c, f) prediction scores</title>
</caption><graphic mimetype="image" mime-subtype="png" xlink:href="fig-12.png"/>
</fig>
<p><xref ref-type="fig" rid="fig-10">Fig. 10</xref> shows, maximum achieved predicted scores of 0.979 on bag class, 0.986 on normal class, and 0.928 on wearing class.</p>
<p><xref ref-type="fig" rid="fig-10">Figs. 10</xref>&#x2013;<xref ref-type="fig" rid="fig-12">12</xref> reveals that the suggested approach, the obtained higher predicted scores are 0.948 on the fast walk, 0.972 on female class, 0.955 on slow walk class, and 0.978 on male class.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>Due to differences in the multiple viewpoints of human gaits, the HGR is a difficult activity. Therefore, in this study tinyYOLOv2-SqueezeNet model is developed that more accurately localized the different types of human gaits. The proposed method achieved mAP of 1.00, 0.91, and 1.00 on Bag, wearing, and normal classes of CASIA-B dataset respectively. Whereas 1.00, 0.70, and 0.95 mAP on the fast walk, slow walk, and normal walk of CASIA-C dataset respectively. Similarly, 1.00 and 0.82 mAP on female and male classes of the CASIA-A dataset respectively. Furthermore, this research investigates a features extraction model based on Conv-BiLSTM that more accurately classifies human gaits. The experimentation is performed on CASIA-A, B, and C datasets. The model achieves 1.00 CPR to classify human with coat wearing. 0.92 CPR on a human with bag class and 0.87 CPR in a normal class. The overall CPR including three classes (wearing, bag, and normal) achieved 0.91. The 1.00 CPR achieved on CASIA-A as well as CASIA-C datasets on all classes such as female, male, human with a slow walk, human with a fast walk, human with the bag. The computed results proved that a combination of CNN and BiLSTM provides the highest recognition rate as compared with individual CNN or the LSTM models. The proposed method performance is dependent on a selected number of features; however, some useful features may be ignored. Moreover, video sequences in a low-quality resolution that affect recognition accuracy.</p>
</sec>
</body>
<back>
<fn-group><fn fn-type="other"><p><bold>Funding Statement:</bold> This research was supported by the Korea Institute for Advancement of Technology (KIAT) Grant funded by the Korea Government (MOTIE) (P0012724, The Competency, Development Program for Industry Specialist) and the Soonchunhyang University Research Fund.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn></fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Sannidhan</surname></string-name> and <string-name><given-names>G.</given-names> <surname>AnanthPrabhu</surname></string-name></person-group>, &#x201C;<article-title>A comprehensive review on various state-of-the-art techniques for composite sketch matching</article-title>,&#x201D; <source>Imperial Journal of Interdisciplinary Research</source>, vol. <volume>2</volume>, no. <issue>2</issue>, pp. <fpage>1131</fpage>&#x2013;<lpage>1138</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. V.</given-names> <surname>Boulgouris</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Hatzinakos</surname></string-name> and <string-name><given-names>K. N.</given-names> <surname>Plataniotis</surname></string-name></person-group>, &#x201C;<article-title>Gait recognition: A challenging signal processing technology for biometric identification</article-title>,&#x201D; <source>IEEE Signal Processing Magazine</source>, vol. <volume>22</volume>, no. <issue>6</issue>, pp. <fpage>78</fpage>&#x2013;<lpage>90</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Koide</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Miura</surname></string-name></person-group>, &#x201C;<article-title>Identification of a specific person using color, height, and gait features for a person following robot</article-title>,&#x201D; <source>Robotics and Autonomous Systems</source>, vol. <volume>84</volume>, no. <issue>2</issue>, pp. <fpage>76</fpage>&#x2013;<lpage>87</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. P.</given-names> <surname>Singh</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Jain</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Arora</surname></string-name> and <string-name><given-names>U. P.</given-names> <surname>Singh</surname></string-name></person-group>, &#x201C;<article-title>Vision-based gait recognition: A survey</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>6</volume>, pp. <fpage>70497</fpage>&#x2013;<lpage>70527</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Bharadwaj</surname></string-name>, <string-name><given-names>T. I.</given-names> <surname>Dhamecha</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Vatsa</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Singh</surname></string-name></person-group>, &#x201C;<article-title>Computationally efficient face spoofing detection with motion magnification</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition Workshops</conf-name>, <publisher-loc>Portland, Oregon</publisher-loc>, pp. <fpage>105</fpage>&#x2013;<lpage>110</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Jia</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wu</surname></string-name></person-group>, &#x201C;<article-title>A study on gait-based gender classification</article-title>,&#x201D; <source>IEEE Transactions on Image Processing</source>, vol. <volume>18</volume>, no. <issue>8</issue>, pp. <fpage>1905</fpage>&#x2013;<lpage>1910</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Veeraraghavan</surname></string-name>, <string-name><given-names>A. K.</given-names> <surname>Roy-Chowdhury</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Chellappa</surname></string-name></person-group>, &#x201C;<article-title>Matching shape sequences in video with applications in human movement analysis</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>27</volume>, no. <issue>12</issue>, pp. <fpage>1896</fpage>&#x2013;<lpage>1909</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Parameswaran</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Chellappa</surname></string-name></person-group>, &#x201C;<article-title>View invariance for human action recognition</article-title>,&#x201D; <source>International Journal of Computer Vision</source>, vol. <volume>66</volume>, no. <issue>1</issue>, pp. <fpage>83</fpage>&#x2013;<lpage>101</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Shotton</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Fitzgibbon</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Cook</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Sharp</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Finocchio</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<chapter-title>Real-time human pose recognition in parts from single depth images</chapter-title>,&#x201D; in <source>CVPR</source>, <publisher-loc>Providence, RI, USA</publisher-loc>, pp. <fpage>1297</fpage>&#x2013;<lpage>1304</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. F.</given-names> <surname>Abate</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Nappi</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Riccio</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Sabatino</surname></string-name></person-group>, &#x201C;<article-title>2D and 3D face recognition: A survey</article-title>,&#x201D; <source>Pattern Recognition Letters</source>, vol. <volume>28</volume>, no. <issue>14</issue>, pp. <fpage>1885</fpage>&#x2013;<lpage>1906</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. S.</given-names> <surname>Matovski</surname></string-name>, <string-name><given-names>M. S.</given-names> <surname>Nixon</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Mahmoodi</surname></string-name> and <string-name><given-names>J. N.</given-names> <surname>Carter</surname></string-name></person-group>, &#x201C;<article-title>The effect of time on gait recognition performance</article-title>,&#x201D; <source>IEEE Transactions on Information Forensics and Security</source>, vol. <volume>7</volume>, no. <issue>2</issue>, pp. <fpage>543</fpage>&#x2013;<lpage>552</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I. C.</given-names> <surname>Chang</surname></string-name> and <string-name><given-names>S. Y.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>3D human motion tracking based on a progressive particle filter</article-title>,&#x201D; <source>Pattern Recognition</source>, vol. <volume>43</volume>, no. <issue>10</issue>, pp. <fpage>3621</fpage>&#x2013;<lpage>3635</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D. M.</given-names> <surname>Gavrila</surname></string-name> and <string-name><given-names>L. S.</given-names> <surname>Davis</surname></string-name></person-group>, &#x201C;<article-title>3-D model-based tracking of humans in action: a multi-view approach</article-title>,&#x201D; in <conf-name>Proc. CVPR IEEE Computer Society Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>San Francisco, CA, USA</publisher-loc>, pp. <fpage>73</fpage>&#x2013;<lpage>80</lpage>, <year>1996</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Lepetit</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Fua</surname></string-name></person-group>, &#x201C;<article-title>Monocular model-based 3d tracking of rigid objects: A survey</article-title>,&#x201D; <source>Foundations and Trends in Computer Graphics and Vision</source>, vol. <volume>1</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>89</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Wan</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>V. V.</given-names> <surname>Phoha</surname></string-name></person-group>, &#x201C;<article-title>A survey on gait recognition</article-title>,&#x201D; <source>ACM Computing Surveys (CSUR)</source>, vol. <volume>51</volume>, no. <issue>5</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>35</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Anusha</surname></string-name> and <string-name><given-names>C. D.</given-names> <surname>Jaidhar</surname></string-name></person-group>, &#x201C;<article-title>Clothing invariant human gait recognition using modified local optimal oriented pattern binary descriptor</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>79</volume>, no. <issue>3</issue>, pp. <fpage>2873</fpage>&#x2013;<lpage>2896</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Anusha</surname></string-name> and <string-name><given-names>C. D.</given-names> <surname>Jaidhar</surname></string-name></person-group>, &#x201C;<article-title>Human gait recognition based on histogram of oriented gradients and Haralick texture descriptor</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>79</volume>, no. <issue>11&#x2013;12</issue>, pp. <fpage>8213</fpage>&#x2013;<lpage>8234</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. H.</given-names> <surname>Shah</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Sharif</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Yasmin</surname></string-name> and <string-name><given-names>S. L.</given-names> <surname>Fernandes</surname></string-name></person-group>, &#x201C;<article-title>Facial expressions classification and false label reduction using LDA and threefold SVM</article-title>,&#x201D; <source>Pattern Recognition Letters</source>, vol. <volume>139</volume>, pp. <fpage>166</fpage>&#x2013;<lpage>173</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Gait recognition based on the feature extraction of Gabor filter and linear discriminant analysis and improved local coupled extreme learning machine</article-title>,&#x201D; <source>Hindawi</source>, vol. <volume>20</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. O.</given-names> <surname>Lishani</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Boubchir</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Khalifa</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Bouridane</surname></string-name></person-group>, &#x201C;<article-title>Human gait recognition using GEI-based local multi-scale feature descriptors</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>78</volume>, no. <issue>5</issue>, pp. <fpage>5715</fpage>&#x2013;<lpage>5730</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. K.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>G. M.</given-names> <surname>Sultaniya</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Chattopadhyay</surname></string-name></person-group>, &#x201C;<article-title>An efficient descriptor for gait recognition using spatio-temporal cues</article-title>,&#x201D; <source>Emerging Technology in Modelling and Graphics</source>, vol. <volume>5</volume>, pp. <fpage>85</fpage>&#x2013;<lpage>97</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Vinothkanna</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Sasikumar</surname></string-name></person-group>, &#x201C;<chapter-title>A novel multimodal biometrics system with fingerprint and gait recognition traits using contourlet derivative weighted rank fusion</chapter-title>,&#x201D; in <source>Int. Conf. on Computational Vision and Bio Inspired Computing</source>, <publisher-loc>Cham</publisher-loc>, <publisher-name>Springer</publisher-name>, pp. <fpage>950</fpage>&#x2013;<lpage>963</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Pavithra</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Math</surname></string-name></person-group>, &#x201C;<article-title>A review on human gait detection</article-title>,&#x201D; <source>Global Journal of Computer Science and Technology</source>, vol. <volume>2</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Tran</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yin</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Atoum</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Liu</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Gait recognition via disentangled representation learning</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Long Beach, CA</publisher-loc>, pp. <fpage>4710</fpage>&#x2013;<lpage>4719</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Gait feature extraction and gait classification using two-branch CNN</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>79</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Ranjan</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Arya</surname></string-name>, <string-name><given-names>S. L.</given-names> <surname>Fernandes</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Sravya</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Jain</surname></string-name></person-group>, &#x201C;<article-title>A fuzzy neural network approach for automatic K-complex detection in sleep EEG signal</article-title>,&#x201D; <source>Pattern Recognition Letters</source>, vol. <volume>115</volume>, no. <issue>6</issue>, pp. <fpage>74</fpage>&#x2013;<lpage>83</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Arshad</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>M. I.</given-names> <surname>Sharif</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Yasmin</surname></string-name>, <string-name><given-names>J. M. R.</given-names> <surname>Tavares</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>A multilevel paradigm for deep convolutional neural network features selection with an application to human gait recognition</article-title>,&#x201D; <source>Expert Systems</source>, vol. <volume>2</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>25</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Arora</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Srivastava</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Singhal</surname></string-name></person-group>, &#x201C;<chapter-title>Analysis of gait flow image and gait Gaussian image using extension neural network for gait recognition</chapter-title>,&#x201D; in <source>Deep Learning and Neural Networks: Concepts, Methodologies, Tools, and Applications</source>, Management Association, I. (Ed.) <publisher-loc>Hershey PA, USA</publisher-loc>: <publisher-name>IGI Global</publisher-name>, pp. <fpage>429</fpage>&#x2013;<lpage>449</lpage>, <year>2020</year>. <comment>Chapter 25</comment>, [Online]. <uri>http://www.igi.Global.com</uri>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. M.</given-names> <surname>Hasan</surname></string-name>, <string-name><given-names>H. A.</given-names> <surname>Mustafa</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Security</surname></string-name></person-group>, &#x201C;<article-title>Multi-level feature fusion for robust pose-based gait recognition using RNN</article-title>,&#x201D; <source>International Journal of Computer Science and Information Security</source>, vol. <volume>18</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>An</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>A model-based gait recognition method with body pose and human prior knowledge</article-title>,&#x201D; <source>Pattern Recognition</source>, vol. <volume>98</volume>, no. <issue>2</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>11</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>O.</given-names> <surname>Dehzangi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Taherisadr</surname></string-name>, <string-name><given-names>R.</given-names> <surname>ChangalVala</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Asnani</surname></string-name></person-group>, &#x201C;<chapter-title>Motion-based gait identification using spectro-temporal transform and convolutional neural networks</chapter-title>,&#x201D; in <source>Advances in Body Area Networks</source>, <publisher-loc>Switzerland</publisher-loc>: <publisher-name>Springer, Cham</publisher-name>, pp. <fpage>407</fpage>&#x2013;<lpage>421</lpage>, <year>2019</year>. [Online]. <uri>https://doi.org/10.1007/978-3-030-02819-0_31</uri>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yu</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Cross-view gait recognition by discriminative feature learning</article-title>,&#x201D; <source>IEEE Transactions on Image Processing</source>, vol. <volume>29</volume>, pp. <fpage>1001</fpage>&#x2013;<lpage>1015</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Learning joint gait representation via quintuplet loss minimization</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Long Beach, CA</publisher-loc>, pp. <fpage>4700</fpage>&#x2013;<lpage>4709</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>S. F.</given-names> <surname>Chang</surname></string-name></person-group>, &#x201C;<chapter-title>3D shape retrieval using a single depth image from low-cost sensors</chapter-title>,&#x201D; in <source>2016 IEEE Winter Conf. on Applications of Computer Vision</source>, <publisher-loc>Lake Placid, NY, USA</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. M.</given-names> <surname>Hasan</surname></string-name> and <string-name><given-names>H. A.</given-names> <surname>Mustafa</surname></string-name></person-group>, &#x201C;<article-title>Multi-level feature fusion for robust pose-based gait recognition using RNN</article-title>,&#x201D; <source>International Journal of Computer Science and Information Security</source>, vol. <volume>18</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>She</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Nahavandi</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Kouzani</surname></string-name></person-group>, &#x201C;<chapter-title>A review of vision-based gait recognition methods for human identification</chapter-title>,&#x201D; in <source>2010 Int. Conf. on Digital Image Computing: Techniques and Applications</source>, <publisher-loc>Sydney, NSW, Australia</publisher-loc>, pp. <fpage>320</fpage>&#x2013;<lpage>327</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Hochreiter</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Schmidhuber</surname></string-name></person-group>, &#x201C;<article-title>Long short-term memory</article-title>,&#x201D; <source>Neural Computation</source>, vol. <volume>9</volume>, no. <issue>8</issue>, pp. <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>, <year>1997</year>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>He</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<chapter-title>Robust view transformation model for gait recognition</chapter-title>,&#x201D; in <source>2011 18th IEEE Int. Conf. on Image Processing</source>. <publisher-loc>Brussels, Belgium</publisher-loc>, pp. <fpage>2073</fpage>&#x2013;<lpage>2076</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Tan</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<chapter-title>A framework for evaluating the effect of view angle, clothing and carrying condition on gait recognition</chapter-title>,&#x201D; in <source>18th Int. Conf. on Pattern Recognition</source>, <publisher-loc>Hong Kong, China</publisher-loc>, pp. <fpage>441</fpage>&#x2013;<lpage>444</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>He</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Robust recovery of corrupted low-rankmatrix by implicit regularizers</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>36</volume>, no. <issue>4</issue>, pp. <fpage>770</fpage>&#x2013;<lpage>783</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Tao</surname></string-name></person-group>, &#x201C;<article-title>A cascade fusion scheme for gait and cumulative foot pressure image recognition</article-title>,&#x201D; <source>Pattern Recognition</source>, vol. <volume>45</volume>, no. <issue>10</issue>, pp. <fpage>3603</fpage>&#x2013;<lpage>3610</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Huang</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<chapter-title>Evaluation framework on translation-invariant representation for cumulative foot pressure image</chapter-title>,&#x201D; in <source>2011 18th IEEE Int. Conf. on Image Processing</source>, <publisher-loc>Brussels, Belgium</publisher-loc>, pp. <fpage>201</fpage>&#x2013;<lpage>204</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Ning</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Hu</surname></string-name></person-group>, &#x201C;<article-title>Silhouette analysis-based gait recognition for human identification</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>25</volume>, no. <issue>12</issue>, pp. <fpage>1505</fpage>&#x2013;<lpage>1518</lpage>, <year>2003</year>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yu</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<chapter-title>Efficient night gait recognition based on template matching</chapter-title>,&#x201D; in <source>18th Int. Conf. on Pattern Recognition</source>, <publisher-loc>Hong Kong, China</publisher-loc>, pp. <fpage>1000</fpage>&#x2013;<lpage>1003</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>W. Q.</given-names> <surname>Yan</surname></string-name></person-group>, &#x201C;<article-title>Cross-view gait recognition through ensemble learning</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>32</volume>, no. <issue>11</issue>, pp. <fpage>7275</fpage>&#x2013;<lpage>7287</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>W. Q.</given-names> <surname>Yan</surname></string-name></person-group>, &#x201C;<article-title>Human gait recognition based on frame-by-frame gait energy images and convolutional long short-term memory</article-title>,&#x201D; <source>International Journal of Neural Systems</source>, vol. <volume>30</volume>, no. <issue>1</issue>, pp. <fpage>1950027</fpage>&#x2013;<lpage>1950028</lpage>, <year>2020</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>