<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CSSE</journal-id>
<journal-id journal-id-type="nlm-ta">CSSE</journal-id>
<journal-id journal-id-type="publisher-id">CSSE</journal-id>
<journal-title-group>
<journal-title>Computer Systems Science &#x0026; Engineering</journal-title>
</journal-title-group>
<issn pub-type="ppub">0267-6192</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">39479</article-id>
<article-id pub-id-type="doi">10.32604/csse.2023.039479</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Visual Motion Segmentation in Crowd Videos Based on Spatial-Angular Stacked Sparse Autoencoders</article-title>
<alt-title alt-title-type="left-running-head">Visual Motion Segmentation in Crowd Videos Based on Spatial-Angular Stacked Sparse Autoencoders</alt-title>
<alt-title alt-title-type="right-running-head">Visual Motion Segmentation in Crowd Videos Based on Spatial-Angular Stacked Sparse Autoencoders</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Hafeezallah</surname><given-names>Adel</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Al-Dhamari</surname><given-names>Ahlam</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="aff" rid="aff-3">3</xref><email>kmahlam@utm.my</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Abu-Bakar</surname><given-names>Syed Abd Rahman</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Electrical Engineering, Taibah University</institution>, <addr-line>Madinah</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Electronic and Computer Engineering, Faculty of Electrical Engineering, Universiti Teknologi Malaysia</institution>, <addr-line>Johor Bahru, 81310</addr-line>, <country>Malaysia</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Engineering, Hodeidah University</institution>, <addr-line>Hodeidah</addr-line>, <country>Yemen</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Ahlam Al-Dhamari. Email: <email>kmahlam@utm.my</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic"><year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>26</day><month>5</month><year>2023</year></pub-date>
<volume>47</volume>
<issue>1</issue>
<fpage>593</fpage>
<lpage>611</lpage>
<history>
<date date-type="received"><day>31</day><month>1</month><year>2023</year></date>
<date date-type="accepted"><day>20</day><month>3</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Hafeezallah et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Hafeezallah et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CSSE_39479.pdf"></self-uri>
<abstract>
<p>Visual motion segmentation (VMS) is an important and key part of many intelligent crowd systems. It can be used to figure out the flow behavior through a crowd and to spot unusual life-threatening incidents like crowd stampedes and crashes, which pose a serious risk to public safety and have resulted in numerous fatalities over the past few decades. Trajectory clustering has become one of the most popular methods in VMS. However, complex data, such as a large number of samples and parameters, makes it difficult for trajectory clustering to work well with accurate motion segmentation results. This study introduces a spatial-angular stacked sparse autoencoder model (SA-SSAE) with <italic>l2</italic>-regularization and softmax, a powerful deep learning method for visual motion segmentation to cluster similar motion patterns that belong to the same cluster. The proposed model can extract meaningful high-level features using only spatial-angular features obtained from refined tracklets (a.k.a &#x2018;trajectories&#x2019;). We adopt <italic>l2</italic>-regularization and sparsity regularization, which can learn sparse representations of features, to guarantee the sparsity of the autoencoders. We employ the softmax layer to map the data points into accurate cluster representations. One of the best advantages of the SA-SSAE framework is it can manage VMS even when individuals move around randomly. This framework helps cluster the motion patterns effectively with higher accuracy. We put forward a new dataset with its manual ground truth, including 21 crowd videos. Experiments conducted on two crowd benchmarks demonstrate that the proposed model can more accurately group trajectories than the traditional clustering approaches used in previous studies. The proposed SA-SSAE framework achieved a 0.11 improvement in accuracy and a 0.13 improvement in the F-measure compared with the best current method using the CUHK dataset.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Visual motion segmentation</kwd>
<kwd>crowd behavior analysis</kwd>
<kwd>trajectory analysis</kwd>
<kwd>crowd dynamics</kwd>
<kwd>autoencoders</kwd>
<kwd>motion patterns</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Deputyship of Research &#x0026; Innovation, Ministry of Education in Saudi Arabia</funding-source>
<award-id>758</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Crowd behavior analysis (CBA) is one of the most significant and critical topics for ensuring that large events in public areas run smoothly, peacefully, and without casualties. Furthermore, CBA is a multidisciplinary topic that concentrates on a variety of domains, including biology, sociology, and computer vision. Among the important applications of CBA is visual motion segmentation (VMS), which provides a significant amount of information about crowd dynamics in both natural and human communities. Generally, one person&#x2019;s information can only convey a limited amount of local information about the scene. On the other hand, individuals are recognized as union members when crowd motion arises, and they share the same characteristics that are extremely significant for research across many fields. VMS aims to break down a visual image into cohesive subsets corresponding to rigidly and independently moving targets. The pedestrians within each set show collective behaviors and similar motion paths. VMS is an essential preprocessing for a variety of computer vision applications, and it has become a burgeoning study topic in the last few decades. When moving objects are semantically and independently categorized, we obtained motion segments that we can utilize in a variety of video-surveillance tasks, including motion analysis, video indexing, traffic monitoring, activity recognition, crowd counting, crowd tracking, abnormal event detection, disasters recognition, and semantic scene segmentation [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>Even though VMS has progressed considerably [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>] in the past few years, further improvements are still required to accomplish satisfactory performance. The major challenge in VMS is when the target is too small scaled. Because of the high occlusion in crowd images, state-of-the-art (SOTA) approaches use feature points, which then combined into one group using similar motions to avoid directly detecting pedestrians. Depending on how the objects move in the scene, both structured and unstructured crowd images can be formed. A structured crowd scene has consistent spatiotemporal motion patterns formed by objects that move in concert over the whole scene. Put another way, every spatial location in any structured crowd scene has an identical motion pattern, and the motion&#x2019;s direction does not change most of the time (<xref ref-type="fig" rid="fig-1">Fig. 1a</xref>). An unstructured crowd scene, on the other hand, consists of non-uniform spatiotemporal motion patterns formed by irregularly moving objects whose movement direction continually changes and cannot be anticipated (<xref ref-type="fig" rid="fig-1">Fig. 1b</xref>).</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>The motion patterns are exhibited as tracklets; every color represents a particular pattern. (a) Structured crowd (b) unstructured crowd</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-1.tif"/></fig>
<p>With the evolution of surveillance devices, massive amounts of human trajectory data have been captured, making it vitally difficult and crucial to extract valuable data. A human trajectory is a set of sequenced spatio-temporal data from a single person. Human trajectories provide insight into a variety of real-world applications. An effective way to analyze human trajectories is through trajectory clustering. Trajectory clustering approaches are classified into three groups based on the availability of labeled data: supervised, unsupervised, and semi-supervised. The learning of supervised models occurs before trajectory clustering. Labeled data is typically employed to train a function that creates clusters by mapping data to labels. Then, this function is utilized to predict the clusters of unlabeled data. The objective of unsupervised models is to cluster data without the aid of humans or labeled data. By analyzing unlabeled datasets, an inference function can be built, which can then be used to group data. The former two models are combined through semi-supervised modeling, which is modified using unlabeled data after learning it on labeled data [<xref ref-type="bibr" rid="ref-11">11</xref>]. However, traditional clustering methods would struggle to achieve satisfactory performance if the original data were not evenly distributed owing to high intravariance, as shown on the left side of <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>The distribution of the original data is on the left side. Because of the large intravariance, it is hard to separate the data properly. By performing a non-linear transformation function, the points are compacted in a new space with regard to the appropriate cluster centers, as shown on the right side</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-2.tif"/></fig>
<p>To address the earlier issue, we aim to map the spatial-angular feature space to a new feature space that is more suited for clustering tasks. The sparse-autoencoder network is a strong contender to solve the above issue. It uses iterative learning to learn both the encoder (EN) and the decoder (DE) to provide a non-linear transformation function. EN is actually the non-linear transformation function, and DE requires reconstructing precise data from the feature representation produced by the EN. Repeating this procedure ensures that the transformation function is reliable and can accurately represent the data. This paper proposes an effective model to cluster human trajectories in different crowd places based on spatial-angular stacked sparse autoencoders (SA-SSAE). When using high-dimensional large-scale datasets, conventional clustering techniques suffer from major performance concerns. For instance, dimensionality reduction techniques must be used prior to using the clustering algorithm to extract features from raw data [<xref ref-type="bibr" rid="ref-12">12</xref>]. Deep learning has always been at the heart of tackling these concerns [<xref ref-type="bibr" rid="ref-13">13</xref>], so we used stacked sparse autoencoders (SSAE) to extract features from the spatio-temporal trajectory data and to group the trajectories together based on their shared characteristics.</p>
<p>The contributions of our work are as follows:
<list list-type="bullet">
<list-item><p>An efficient framework called SA-SSAE for VMS has been proposed. The SA-SSAE framework is indispensable for high-level crowd behavior analysis. The proposed framework&#x2019;s capacity to handle pedestrian flows that are randomly dispersed is one of its main attractions. To our knowledge, this is the first study that leverages stacked sparse autoencoders for visual motion segmentation in crowd videos using trajectory data employing the Chinese University of Hong Kong (CUHK) crowd benchmark.</p></list-item>
<list-item><p>The generalized Kanade-Lucas-Tomasi key point tracker (gKLT) is applied to the input video sequences to extract the trajectories of motion patterns. Individual trajectory formation is a primary concern in the VMS. For studying and analyzing crowded environments, the accurate extraction of individual trajectories over time is crucial. In this study, the generalized KLT tracker is applied to video sequences to construct individual trajectories.</p></list-item>
<list-item><p>Dataset with ground-truth annotations is presented to validate our framework&#x2019;s performance. The Hajj, a significant pilgrimage to Mecca in Saudi Arabia, is one of Islam&#x2019;s five pillars. Every year, up to four million pilgrims carry out the Hajj rituals. As a result, it ranks as one of the most significant pedestrian issues worldwide. Therefore, we collected our Hajj dataset from real-world crowd scenes in Mecca. The dataset includes 21 crowd videos with different scenes and scenarios.</p></list-item>
<list-item><p>Our framework is evaluated using two real-world datasets with various crowd densities. On both datasets, we found that our framework can construct high-quality clusters and outperform existing methods quantitatively.</p></list-item>
</list></p>
<p>The remainder of the paper is laid out as follows: Section 2 discusses related work. Section 3 presents the proposed spatial-angular stacked sparse autoencoder model for VMS. Comparative results are discussed in Section 4. Concluding remarks are presented in Section 5.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Work</title>
<p>The various methods for VMS are reviewed in this section. The VMS approaches can be categorized into three main subsets depending on the density of movement flows. Approaches addressing a maximum of five humans are described under the domain of low-level density (LLD). Approaches that address between five and fifteen humans are characterized as mid-level density (MLD). Likewise, approaches that target more than fifteen humans fall under the subset of high-level density (HLD) [<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<p><bold>LLD Approaches:</bold> Nguyen et al. proposed a consensual approach for visual motion segmentation in dynamic views [<xref ref-type="bibr" rid="ref-15">15</xref>]. The model combined unsupervised techniques to address the label correspondence issue. Seyedhosseini et al. put forward a discriminative learning scheme, called CHM, for semantic segmentation. It benefits from contextual data at various resolutions in a hierarchy. The ability of CHM to optimize a posterior probability at various resolutions is its major feature. It effectively and greedily implements this optimization. CHM trains a number of classifiers at various resolutions and uses the gained findings to learn a classifier at the original resolution [<xref ref-type="bibr" rid="ref-16">16</xref>]. An approach for motion segmentation by employing optical flow orientations was proposed by Narayana et al. [<xref ref-type="bibr" rid="ref-17">17</xref>]. The over-segmentation of an image into depth-dependent units was addressed by the utilization of optical flow orientations. Their approach could automatically obtain the quantity of foreground motions. Meunier et al. proposed a CNN-based fully unsupervised approach for VMS from optical flow. Meunier et al. hypothesized that the optical flow input could be expressed as a set of parametric motion models, commonly quadratic or affine. The basic principle of their method is to utilize the Expectation-Maximization to rationally build a loss function and a training procedure for their ground-truth-free neural network for VMS [<xref ref-type="bibr" rid="ref-18">18</xref>]. Choudhury et al. proposed a method for VMS by combining the advantages of appearance-based and motion-based segmentation. They put forward supervising a network of image segmentation with the function of predicting areas that are likely to have simple motion patterns and so are likely to match with targets [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p><bold>MLD Approaches:</bold> Mukherjee et al. presented a linear-time video segmentation approach [<xref ref-type="bibr" rid="ref-20">20</xref>] that uses a Gaussian mixture model (GMM) to cluster each video sequence. It also utilizes recursive filtering to re-obtain the parameters of the GMM. In addition to updating the variance iteratively and creating or removing clusters as needed, their hybrid approach can uniquely propagate Gaussian clusters through each subsequent frame. However, a distance threshold parameter is required as its primary input. A cluster similarity criterion, which may be based on a user-defined distance metric, controls how new clusters are included and removed. A unified conditional random field (CRF model) was proposed by [<xref ref-type="bibr" rid="ref-21">21</xref>] for multiple-targets joint-tracking and segmentation in complex video sequences. The model utilizes low-level image features to associate each super pixel with a particular objective or to designate it as a background.</p>
<p><bold>HLD Approaches:</bold> A statistical approach established using the Lagrangian-particle-dynamics [<xref ref-type="bibr" rid="ref-22">22</xref>] for flow segmentation and detection of flow instability was presented in [<xref ref-type="bibr" rid="ref-9">9</xref>]. The flow field created by the crowd motion is handled as an aperiodic dynamical system. To determine the Lagrangian coherent structures existent in the underlying flow, a Finite Time Lyapunov exponent field is built using the greatest eigenvalue of the tensor. In a normalized cuts framework, the Lagrangian coherent structures are used to identify the boundaries of the flow segments by dividing the flow into regions with markedly distinct dynamics. Establishing correspondences between flow segments over time allows for the detection of any alteration in the number of flow segments, called instability. From the point of vision, [<xref ref-type="bibr" rid="ref-6">6</xref>] carefully looked at groups&#x2019; basic and universal aspects that are present in several crowd environments. These aspects are essential to comprehending congested environments and are driven by sociopsychological investigations. Moreover, a group discovery approach was put forward by learning the collective transition priors.</p>
<p>The VMS problem may alternatively be considered a trajectory clustering (TC) challenge. Determining a proper metric to calculate the similarity of trajectories with different attributes and a proper method to cluster the trajectories on the basis of their commonalities are the two primary issues in TC [<xref ref-type="bibr" rid="ref-23">23</xref>]. Most VMS approaches explored to date either have limited applicability to certain categories of crowd scenes or suffer from major performance concerns. Based on spatial-angular stacked sparse autoencoders, this study provides a model for clustering human trajectories in various crowd environments. The proposed framework is effective and robust for various crowd scene scenarios.</p>
</sec>
<sec id="s3"><label>3</label><title>Proposed Spatial-Angular Stacked Sparse Autoencoder Model for VMS</title>
<p>The proposed intelligent VMS framework is described in detail in this section. The SSAE network is built based on spatial-angular motion information and is used to find clusters of similar instances in an unlabeled dataset. The framework is illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. Five main steps are involved in constructing the proposed framework: <italic>generating trajectories</italic>, <italic>refining the generated trajectories</italic>, <italic>obtaining spatial-angular motion information</italic>, <italic>applying spatial-angular stacked sparse autoencoders</italic>, and <italic>softmax layer</italic>. SA-SSAE is proposed to automatically segment motion patterns, expressed as trajectories. The SA-SSAE framework employs the generalized KLT tracker (gKLT tracker [<xref ref-type="bibr" rid="ref-24">24</xref>]), which is proposed as an improvement to the standard KLT tracker [<xref ref-type="bibr" rid="ref-25">25</xref>]), owing to its computing efficiency and tracking accuracy. Every generated trajectory is made up of a set of 2D spatial coordinates <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where <italic>L</italic> &#x2208; [1, <italic>k</italic>] is the trajectory length. The following sub-section describes each step where trajectories are refined following Shao et al.&#x2019;s approach. The motion point features that are made from the new trajectories are then used to make both spatial and angular features. This valuable motion information is then fed into the stack sparse autoencoder network to output motion segments.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Flowchart of the SA-SSAE framework</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-3.tif"/></fig>
<sec id="s3_1"><label>3.1</label><title>Generating Trajectories</title>
<p>A major challenge in the VMS is the formation of individual trajectories. The precise extraction of individual trajectories over time is critical for investigating and examining crowded scenarios. Following state-of-the-art papers [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>], we utilize the gKLT due to its efficacy in determining trajectories, particularly for small objects in crowd images. The gKLT was first proposed by [<xref ref-type="bibr" rid="ref-25">25</xref>]. First, gKLT obtains the feature points in the moving foreground with enough texture data to detect. The feature points are then tracked, and their velocities are calculated frame by frame based on their displacements. A collection of tracked feature points is thus acquired. The gKLT tracker commences with a group of sparse features in the present frame <italic>f<sub>i</sub></italic> and seeks to locate their positions in the subsequent frame <italic>f<sub>i&#x2009;&#x002B;&#x2009;1</sub></italic> by matching a patch of an image surrounding a feature to its identical image patch in the subsequent frame. The assumption of brightness constancy indicates that the intensities of the patches in the subsequent image will not vary significantly. Utilizing a patch enables differentiation between surrounding points of equivalent intensity. The <italic>&#x03C9;</italic>(<italic>x</italic>) window function, which is commonly a Gaussian function, is employed to highlight nearby pixels more than far-off ones. This explains why points nearer to the feature point are likelier to exhibit comparable motion than those further away. Calculating the error function is the next step in the matching:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:munder><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>through the window function&#x2019;s support <italic>&#x03C9;</italic>(<italic>x<sub>0</sub></italic>), which is positioned over the feature point <italic>x<sub>0</sub></italic>. The estimated displacement <italic>d</italic> for the feature point at <italic>x<sub>0</sub></italic> in picture <italic>f<sub>i</sub></italic> is obtained by decreasing the weighted nonlinear least squares formula. The simulation results (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>) show that the gKLT algorithm effectively tracks feature points that indicate crowd movement.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>gKLT algorithm simulation results</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-4.tif"/></fig>
</sec>
<sec id="s3_2"><label>3.2</label><title>Refining the Generated Trajectories</title>
<p>Through the former processing, a collection of tracked feature points was obtained. Nevertheless, the points gKLT generated do not accurately describe the pedestrians because, in some cases, there could be numerous points inside the same part of a moving person. Moreover, the points can be located in the background or formed owing to changes in illumination, producing noisy, short, and static tracklets. In our framework, following Shao et al. [<xref ref-type="bibr" rid="ref-6">6</xref>], such tracklets are filtered out, which increase VMS&#x2019;s overall performance.</p>
</sec>
<sec id="s3_3"><label>3.3</label><title>Obtaining Spatial-Angular Motion Information</title>
<p>The spatial locations of motion features are obtained from the newly refined trajectories. The spatial location features of a trajectory are critical because trajectories that are spaced far apart, even if they are identical in shape, often do not belong to the same cluster. The spatial location features for each trajectory <italic>t<sub>i</sub></italic> can be calculated using the following equation:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>s</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>v<sub>i</sub></italic>&#x2009;&#x003D;&#x2009;[<italic>x<sub>i</sub></italic>, <italic>y<sub>i</sub></italic>], and <italic>k</italic> is the trajectory length. The <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>s</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the spatial location feature vector. Angular features describe the direction of the crowd&#x2019;s motion [<xref ref-type="bibr" rid="ref-28">28</xref>]. To compute the average feature, the average displacement <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mover><mml:mi>A</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> of a trajectory is obtained first using <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> below:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mover><mml:mi>A</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The average displacement vector <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mover><mml:mi>A</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula> contains two components <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mover><mml:mi>v</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula>. Now, using <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, the average angular can be calculated based on the following equation [<xref ref-type="bibr" rid="ref-28">28</xref>]:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mover><mml:mi>A</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mfrac><mml:mn>180</mml:mn><mml:mi>&#x03C0;</mml:mi></mml:mfrac><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mover><mml:mi>v</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mo>[</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mover><mml:mi>A</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mfrac><mml:mn>180</mml:mn><mml:mi>&#x03C0;</mml:mi></mml:mfrac><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2260;</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mover><mml:mi>v</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2264;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mover><mml:mi>v</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> indicates the horizontal direction&#x2019;s unit vector. The value of <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> differs based on the values of the vectors <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mover><mml:mi>v</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula>. Furthermore, a <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> value of zero indicates the absence of motion.</p>
</sec>
<sec id="s3_4"><label>3.4</label><title>Spatial-Angular Stacked Sparse Autoencoders</title>
<p>In our sparse-autoencoder with smoothed <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:math></inline-formula>2-regularization technique, <italic>m</italic> numbers of input spatial-angular training features <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are given such that <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> are the labels. These training samples are fed into the proposed stacked sparse-autoencoders network. The encoder and decoder are two main components of the sparse-autoencoder training process. The encoder maps the input data into the hidden representation while the decoder reconstructs data from the hidden representation. The hidden encoder vector calculated from <italic>X<sub>m</sub></italic> is denoted by the letter <italic>h<sub>m</sub></italic>. <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the output layer decoder vector. Therefore, the following is the encoding procedure:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>
<italic>f<sub>e</sub></italic> shows the encoding function, <italic>W<sub>e</sub></italic> is the encoding weight parameter, and <italic>b<sub>e</sub></italic> represents the corresponding bias. The following is a description of the decoder process:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<italic>f<sub>d</sub></italic> represents the decoding function. <italic>W<sub>d</sub></italic> and <italic>b<sub>d</sub></italic> are the decoding weight and bias, respectively.</p>
<p>The sparse-autoencoder model minimizes the reconstruction error to learn a meaningful hidden representation. Thus, to diminish the reconstruction error and resolve the parameters <italic>W<sub>e</sub></italic>, <italic>W<sub>d</sub></italic>, <italic>b<sub>e</sub></italic>, and <italic>b<sub>d</sub></italic>, the sparse-autoencoder&#x2019;s parameter settings are optimized as follows:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>&#x03D5;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">argmin</mml:mtext></mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the loss function, where <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></inline-formula> <xref ref-type="fig" rid="fig-5">Fig. 5</xref> presents the SSAE architecture. SSAE in our framework is built by stacking <italic>two</italic> sparse-autoencoders into <italic>m</italic> hidden layers using an unsupervised learning method called &#x201C;<italic>layerwise learning</italic>&#x201D;, which is subsequently fine-tuned utilizing a supervised method. By including a regularizer in the cost-function, it is feasible to promote an autoencoder&#x2019;s sparsity. This regularizer is based on a neuron&#x2019;s average output activation value, described as follows:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>k</italic> represents the overall quantity of the training patterns. <italic>x<sub>i</sub></italic> is the <italic>i</italic><sup>th</sup> training pattern, <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the <italic>n</italic><sup>th</sup> row of the weight matrix <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the <italic>n</italic><sup>th</sup> element of the bias vector <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>. The sparse-autoencoder uses back-propagation (bp) to lessen the cost-function, and it can be computed as follows:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">sparse</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>W</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mi>&#x03C5;</mml:mi><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where the first expression of <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref> presents the average sum-of-squares error, <italic>&#x03BB;</italic> is a variable to control the relative weight of the regularization. Incorporating the <italic>l2</italic>-weight regularizer makes the solution &#x201C;smoother&#x201D; and improves its generalization ability. <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>&#x03C5;</mml:mi></mml:math></inline-formula> is the sparsity penalty term&#x2019;s weight. The <italic>&#x03BB;</italic> and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>&#x03C5;</mml:mi></mml:math></inline-formula> can be specified while training the sparse-autoencoder. <italic>S</italic>(&#x22C5;) denotes the sparsity regularizer that regulates the sparsity of the hidden layer output. <italic>S</italic> has a low value for every neuron &#x201C;specializing&#x201D; in the hidden layer, it produces a high output for a few training samples. On account of this, a lower sparsity fraction promotes a high level of sparsity. <italic>a<sub>j</sub></italic> presents the average output of the j<sup>th</sup>-hidden-unit and can be obtained using <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref> below:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:mfrac><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:msubsup><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></disp-formula>where <italic>u</italic> is the <italic>j</italic><sup>th</sup>-hidden-unit-output of an <italic>i</italic><sup>th</sup>-training pattern. The parameters <italic>l2</italic>-weight regularize, and sparsity regularize are utilized to prevent overfitting. The stack sparse-autoencoders are linked to a softmax layer that predicts the probabilistic assignments of clusters. The softmax layer&#x2019;s mathematical model is as follows [<xref ref-type="bibr" rid="ref-29">29</xref>]:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>r</mml:mi><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> represents the model parameters and the expression <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> normalizes the distribution to guarantee that the sum is equal to one.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Stacked sparse autoencoder architecture</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-5.tif"/></fig>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Experiments</title>
<p>In the dense crowd scenes, individuals are a gathering of many groups with similar motion characteristics. Groups are essential components of a crowd. A novel crowd segmentation framework based on spatial-angular stacked sparse autoencoders was presented in this paper to obtain fundamental crowd interactions for subsequent crowd behavior analysis. The proposed framework can be applied to various dense crowd scenes. In this section, we extensively evaluate the performance of the proposed framework for VMS on two crowd video datasets: the Hajj and the CUHK crowd datasets.</p>
<sec id="s4_1"><label>4.1</label><title>Motion Segmentation Benchmarks</title>
<p><bold>Hajj Benchmark:</bold> The Hajj benchmark consists of crowd videos shot in various indoor and outdoor locations in Mecca. Islam&#x2019;s Hajj, a major pilgrimage to Mecca in Saudi Arabia, is one of the five pillars of Islam. The Hajj rituals are performed annually by up to four million pilgrims. Consequently, it is one of the most significant global pedestrian issues. The benchmark has 21 real-world crowd videos with various scenarios and situations. It has several challenges, including occlusions, lighting, and various object scales. The Hajj benchmark videos contain six scenes, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>A few examples of Hajj benchmark frames. The original frames of crowded indoor and outdoor scenes are on the right, and the motion patterns are on the left</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-6.tif"/></fig>
<p><bold>CUHK Benchmark:</bold> The CUHK is a set of crowd videos that includes 55 videos captured by Shao et al. [<xref ref-type="bibr" rid="ref-6">6</xref>], as well as former crowd benchmarks that have already been released from the following SOTA studies [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. The CUHK collection has 300 annotated crowd clips with different crowd densities, and the benchmark has the groups for each clip&#x2019;s motion pattern segmentation. These 300 video clips are used to validate the proposed model following the SOTA approaches in [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]. Based on the crowd dynamics, these clips are grouped into structured and unstructured groups. Additionally, each scene is grouped into indoor and outdoor subcategories based on the location of the recording and the nature of its content. <xref ref-type="fig" rid="fig-7">Fig. 7</xref> shows examples of crowd images from the CUHK benchmark.</p>
<fig id="fig-7"><label>Figure 7</label><caption><title>CUHK benchmark sample frames</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-7.tif"/></fig>
</sec>
<sec id="s4_2"><label>4.2</label><title>Evaluation Methods and Results</title>
<p>VMS is evaluated as a clustering issue, and the performance evaluation can be achieved by utilizing widely well-known cluster assessment measures: Purity, Rand Index (RI), Normalized Mutual Information (NMI), accuracy, and F-measure [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. A larger value denotes better clustering performance; all performance measurements lie within the range [0, 1]. The values of the training parameters for the sparse autoencoders are listed in <xref ref-type="table" rid="table-1">Table 1</xref>. Purelin function and logistic sigmoid function, respectively, serve as the transfer functions for the encoder and decoder. A linear transfer function known as the Purelin function is defined as follows:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>z</mml:mi></mml:math></disp-formula></p>
<p>The Logistic sigmoid function is given as:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>z</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Parameters of SA-SSAE model</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">No</th>
<th align="left">Parameter</th>
<th align="left">Value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">Max epochs</td>
<td align="left">400</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left"><italic>l2</italic> weight regularization</td>
<td align="left">0.1</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Sparsity regularization</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Sparsity proportion</td>
<td align="left">0.01</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Scale data</td>
<td align="left">False</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Moreover, scaled conjugate gradient descent (SCGD) was employed for the autoencoder training process to optimize the weights and bias by minimizing <italic>J<sub>sparse</sub></italic>(<italic>W</italic>, <italic>b</italic>) in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>. The encoder attempts to encode the considerable input features into a smaller hidden representation with only the related information. The decoder turns the process back to reproduce the same set of features as the input. After training, weights will be assigned to every neuron in the hidden layer, enabling them to provide an efficient stimulus against input visual information. The feature set produced from the first sparse AE&#x2019;s weights is utilized to train the second sparse AE. The first sparse AE&#x2019;s feature t-Distributed Stochastic Neighbor Embedding (t-SNE) plot is shown in <xref ref-type="fig" rid="fig-8">Fig. 8a</xref>. The t-SNE demonstrates very clearly that the feature set exhibits good discriminatory behavior. The size of the hidden layers for the first sparse AE is 30, and the size of the hidden layers for the second sparse AE is 10, leading to a smaller representation of the features. A significant difference that may be seen in the t-SNE plots in <xref ref-type="fig" rid="fig-8">Figs. 8a</xref> and <xref ref-type="fig" rid="fig-8">8b</xref> is that the features are compressed about the appropriate cluster centers in the new feature space. <xref ref-type="fig" rid="fig-9">Fig. 9</xref> presents some visual motion segmentation results of real-life crowd scenes from the CUHK and Hajj crowd benchmarks. Different colors stand for different motion segments. It is clear that the proposed SA-SSAE framework works very well compared to the ground truth segmentation groups.</p>
<fig id="fig-8"><label>Figure 8</label><caption><title>The t-SNE plot of features produced from (a) the first sparse AE, (b) the second sparse AE, using the &#x201C;exact&#x201D; method, which maximizes the Kullback-Leibler divergence (KLD) of distributions between the original data space and the embedded data space</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-8.tif"/></fig><fig id="fig-9"><label>Figure 9</label><caption><title>Visual motion segmentation results based on the proposed framework</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-9.tif"/></fig>
<sec id="s4_2_1"><label>4.2.1</label><title>Results on Hajj Crowd Benchmark</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> compares the results and shows the average performance for 21 crowd videos of the Hajj benchmark. It indicates that the proposed model surpasses the recent SADC approach for VMS by a large margin in terms of purity, NMI, and RI. Upon closer examination, we discovered that the NMI-performance-metric always produces a value of zero when one of the two grouping assignments (ground truth or clustering outcome) has only one cluster and the other clustering has multiple clusters. This is not permitted when calculating the NMI. This is because of an underlying mathematical problem with how mutual information is computed. However, such a debate is outside the purview of this study.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Results based on Hajj benchmark</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Method</th>
<th align="left">Purity</th>
<th align="left">NMI</th>
<th align="left">RI</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">SADC [<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td align="left">0.85</td>
<td align="left">0.21</td>
<td align="left">0.74</td>
</tr>
<tr>
<td align="left">SA-SSAE</td>
<td align="left"><bold>0.93</bold></td>
<td align="left"><bold>0.70</bold></td>
<td align="left"><bold>0.91</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2_2"><label>4.2.2</label><title>Results on CUHK Crowd Benchmark</title>
<p>To demonstrate the efficacy of the proposed model, its performance results are compared to the following SOTA methods: MCC [<xref ref-type="bibr" rid="ref-24">24</xref>], CT [<xref ref-type="bibr" rid="ref-6">6</xref>], and SADC [<xref ref-type="bibr" rid="ref-28">28</xref>]. As shown in <xref ref-type="table" rid="table-3">Table 3</xref>, the proposed SA-SSAE model outperforms all other approaches. As mentioned previously, the NMI will equal zero if one of the two clustering assignments has only one cluster and the other clustering has more than one. With 25&#x0025; of the overall videos in the CUHK benchmark having one group ground-truth, this explains why other approaches give smaller values for NMI.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>SA-SSAE CHUK</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Method</th>
<th align="left">Purity</th>
<th align="left">NMI</th>
<th align="left">RI</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">MCC [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td align="left">0.69</td>
<td align="left">0.43</td>
<td align="left">0.70</td>
</tr>
<tr>
<td align="left">CT [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td align="left">0.76</td>
<td align="left">0.41</td>
<td align="left">0.73</td>
</tr>
<tr>
<td align="left">SADC [<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td align="left">0.93</td>
<td align="left">0.78</td>
<td align="left">0.89</td>
</tr>
<tr>
<td align="left">SA-SSAE [ours]</td>
<td align="left"><bold>0.96</bold></td>
<td align="left"><bold>0.84</bold></td>
<td align="left"><bold>0.95</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-4">Table 4</xref> presents a comparative assessment in terms of the Acc and F-measure using nine relevant approaches HC [<xref ref-type="bibr" rid="ref-34">34</xref>], CF [<xref ref-type="bibr" rid="ref-35">35</xref>], CT [<xref ref-type="bibr" rid="ref-10">10</xref>], CDC [<xref ref-type="bibr" rid="ref-26">26</xref>], MCC [<xref ref-type="bibr" rid="ref-24">24</xref>], AMR [<xref ref-type="bibr" rid="ref-36">36</xref>], HSIM [<xref ref-type="bibr" rid="ref-14">14</xref>], MPF-<italic>l1</italic> [<xref ref-type="bibr" rid="ref-33">33</xref>], and MPF-<italic>l2</italic> [<xref ref-type="bibr" rid="ref-33">33</xref>]. Additionally, <xref ref-type="fig" rid="fig-10">Fig. 10</xref> illustrates the relative improvement of the SA-SSAE model over the other approaches. The proposed model gets the highest Acc and F-measure values, which shows that it is better at getting motion pattern segments. It is well known that performance decreases when pedestrian distribution varies. All SOTA approaches [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>] disregard the changes in the distribution of individuals, which only remains true if the motion flow is coherent across time. Such variations generate isolated areas, which in turn, lower overall performance. Furthermore, the SOTA approaches [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>] cannot group structurally comparable pixels into significative segments. Discovering and segregating isolated pedestrian segments presents a highly complicated issue. Even though the HSIM approach is emphasized for its ability to handle randomly distributed pedestrian flows, our SA-SSAE model achieves better results regarding the F-measure. The relative improvement of the SA-SSAE framework compared to the HSIM approach is 0.36.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>Comparative assessment in terms of the Acc and F-measure employing CUHK benchmark</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Method</th>
<th align="left">Acc</th>
<th align="left">F-measure</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">HC [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td align="left">0.63</td>
<td align="left">0.62</td>
</tr>
<tr>
<td align="left">CF [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td align="left">0.70</td>
<td align="left">0.67</td>
</tr>
<tr>
<td align="left">CT [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td align="left">0.75</td>
<td align="left">0.74</td>
</tr>
<tr>
<td align="left">CDC [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td align="left">0.67</td>
<td align="left">0.67</td>
</tr>
<tr>
<td align="left">MCC [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td align="left">0.68</td>
<td align="left">0.67</td>
</tr>
<tr>
<td align="left">AMR [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td align="left">0.78</td>
<td align="left">0.76</td>
</tr>
<tr>
<td align="left">HSIM [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td align="left">-</td>
<td align="left">0.58</td>
</tr>
<tr>
<td align="left">MPF-<italic>l1</italic> [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td align="left">0.83</td>
<td align="left">0.80</td>
</tr>
<tr>
<td align="left">MPF-<italic>l2</italic> [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td align="left">0.80</td>
<td align="left">0.79</td>
</tr>
<tr>
<td align="left">SA-SSAE [ours]</td>
<td align="left"><bold>0.94</bold></td>
<td align="left"><bold>0.93</bold></td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-10"><label>Figure 10</label><caption><title>The relative improvements of the SA-SSAE model compared to HC, CF, CT, CDC, MCC, AMR, HSIM, MPF-<italic>l1</italic>, and MPF-<italic>l2</italic></title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-10.tif"/></fig>
<p><xref ref-type="fig" rid="fig-11">Fig. 11</xref> depicts the comparison findings across several crowd scenario categories (mass movement, street, street-market, station, mall, public pathway, cross-walk, and escalator). The proposed SA-SSAE model performs the best among the other three approaches for all the crowd scenario categories. Moreover, the proposed model provides robust results for all evaluated metrics because it considers the trajectory history for several known frames, which is a significant factor in the model&#x2019;s training. The NMI results for the mass movement category for MCC and CT approaches are very low (close to zero). The SADC approach succeeds in increasing the results of NMI to 0.59. However, SA-SSAE outperforms SADC by 0.38, which proves our model&#x2019;s efficiency.</p>
<fig id="fig-11"><label>Figure 11</label><caption><title>Quantitative comparison according to the scene type (mass movement, street, street-market, station, mall, public-walkway, cross-walk, escalator)</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-11.tif"/></fig>
<p>Similarly, <xref ref-type="fig" rid="fig-12">Figs. 12</xref> and <xref ref-type="fig" rid="fig-13">13</xref> compare the findings across several crowd scenario categories, structured or unstructured, as well as indoor or outdoor, respectively. SA-SSAE excels in all categories.</p>
<fig id="fig-12"><label>Figure 12</label><caption><title>Quantitative comparison according to the crowd movement (structured or unstructured)</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-12.tif"/></fig><fig id="fig-13"><label>Figure 13</label><caption><title>Quantitative comparison according to the crowd location (indoor or outdoor)</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_39479-fig-13.tif"/></fig>
</sec>
<sec id="s4_2_3"><label>4.2.3</label><title>Sparsity Parameter Study</title>
<p>The optimal value of the sparsity proportion parameter for the SA-SSAE model was determined via comparative experiments using the CUHK benchmark. <xref ref-type="table" rid="table-5">Table 5</xref> illustrates that the sparsity proportion gives the best results for all performance metrics when its value is between 0.01 and 0.1. Thus, 0.01 was determined for the sparsity proportion parameter.</p>
<table-wrap id="table-5"><label>Table 5</label><caption><title>Various values of sparsity portion parameter. The sparsity proportion value must be in the range [0, 1]</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Sparsity Proportion</th>
<th align="left">Purity</th>
<th align="left">NMI</th>
<th align="left">RI</th>
<th align="left">Acc</th>
<th align="left">F-measure</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">0.00</td>
<td align="left">0.73</td>
<td align="left">0.26</td>
<td align="left">0.64</td>
<td align="left">0.74</td>
<td align="left">0.75</td>
</tr>
<tr>
<td align="left">0.001</td>
<td align="left">0.73</td>
<td align="left">0.27</td>
<td align="left">0.65</td>
<td align="left">0.75</td>
<td align="left">0.75</td>
</tr>
<tr>
<td align="left"><bold>0.01</bold></td>
<td align="left"><bold>0.96</bold></td>
<td align="left"><bold>0.84</bold></td>
<td align="left"><bold>0.95</bold></td>
<td align="left"><bold>0.94</bold></td>
<td align="left"><bold>0.94</bold></td>
</tr>
<tr>
<td align="left"><bold>0.1</bold></td>
<td align="left"><bold>0.96</bold></td>
<td align="left"><bold>0.84</bold></td>
<td align="left"><bold>0.95</bold></td>
<td align="left"><bold>0.94</bold></td>
<td align="left"><bold>0.94</bold></td>
</tr>
<tr>
<td align="left">0.5</td>
<td align="left">0.96</td>
<td align="left">0.83</td>
<td align="left">0.95</td>
<td align="left">0.94</td>
<td align="left">0.93</td>
</tr>
<tr>
<td align="left">1</td>
<td align="left">0.73</td>
<td align="left">0.26</td>
<td align="left">0.64</td>
<td align="left">0.74</td>
<td align="left">0.75</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Conclusion</title>
<p>This study proposed a deep learning framework for VMS using spatial-angular stacked sparse autoencoders. The framework is meant to aid in proactive stampede avoidance to enhance people&#x2019;s safety in crowd scenes. In this study, it is suggested to decompose the motion tracks in a given crowd scene into groups. Each group contains trajectories with similar behavioral characteristics. We prove that sparse stack autoencoders can be adopted effectively for VMS and provide better results than traditional clustering methods. According to experimental results, the proposed SA-SSAE model is superior to the comparative methods in terms of purity, NMI, RI, accuracy, and F-measure. Future research will focus on methods for integrating our VSM framework with outlier detection. Another crucial element that we will consider for increasing robustness is camera motion. To further improve the performance of the segmentation results, new features will be researched by utilizing other feature descriptors. Moreover, future research can be implemented by investigating the performance of the Kalman filter and YOLO (You Only Look Once) for VMS.</p>
</sec>
</body>
<back>
<ack>
<p>The authors thank the Deputyship of Research &#x0026; Innovation, Ministry of Education in Saudi Arabia, for funding this research work through Project Number 758. The authors also would like to thank the Research Management Center of Universiti Teknologi Malaysia for managing this fund under Vot. No. 4C396.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This research work is supported by the Deputyship of Research &#x0026; Innovation, Ministry of Education in Saudi Arabia (Grant Number 758).</p></sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p></sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Anthwal</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Ganotra</surname></string-name></person-group>, &#x201C;<article-title>An overview of optical flow-based approaches for motion segmentation</article-title>,&#x201D; <source>The Imaging Science Journal</source>, vol. <volume>67</volume>, no. <issue>5</issue>, pp. <fpage>284</fpage>&#x2013;<lpage>294</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Munsif</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Afridi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ullah</surname></string-name>, <string-name><given-names>S. D.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>F. A.</given-names> <surname>Cheikh</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A lightweight convolution neural network for automatic disasters recognition</article-title>,&#x201D; in <conf-name>The 10th European Workshop on Visual Information Processing (EUVIP)</conf-name>, <conf-loc>Lisbon, Portugal</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Al-Dhamari</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Sudirman</surname></string-name> and <string-name><given-names>N. H.</given-names> <surname>Mahmood</surname></string-name></person-group>, &#x201C;<article-title>Abnormal behavior detection using sparse representations through sequential generalization of K-means</article-title>,&#x201D; <source>Turkish Journal of Electrical Engineering and Computer Sciences</source>, vol. <volume>29</volume>, no. <issue>1</issue>, pp. <fpage>152</fpage>&#x2013;<lpage>168</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Raza</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Rafiq</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Awrejcewicz</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Ahmed</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Mohsin</surname></string-name></person-group>, &#x201C;<article-title>Dynamical analysis of coronavirus disease with crowding effect, and vaccination: A study of third strain</article-title>,&#x201D; <source>Nonlinear Dynamics</source>, vol. <volume>107</volume>, no. <issue>4</issue>, pp. <fpage>3963</fpage>&#x2013;<lpage>3982</lpage>, <year>2022</year>; <pub-id pub-id-type="pmid">35002076</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Wu</surname></string-name> and <string-name><given-names>H.</given-names> <surname>San Wong</surname></string-name></person-group>, &#x201C;<article-title>Crowd motion partitioning in a scattered motion field</article-title>,&#x201D; <source>IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)</source>, vol. <volume>42</volume>, no. <issue>1</issue>, pp. <fpage>1443</fpage>&#x2013;<lpage>1454</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Shao</surname></string-name>, <string-name><given-names>C. C.</given-names> <surname>Loy</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Learning scene-independent group descriptors for crowd understanding</article-title>,&#x201D; <source>IEEE Transactions on Circuits and Systems for Video Technology</source>, vol. <volume>27</volume>, no. <issue>6</issue>, pp. <fpage>1290</fpage>&#x2013;<lpage>1303</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Hafeezallah</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Al-Dhamari</surname></string-name> and <string-name><given-names>S. A. R.</given-names> <surname>Abu-Bakar</surname></string-name></person-group>, &#x201C;<article-title>Multi-scale network with integrated attention unit for crowd counting</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>73</volume>, pp. <fpage>3879</fpage>&#x2013;<lpage>3903</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Hafeezallah</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Al-Dhamari</surname></string-name> and <string-name><given-names>S. A. R.</given-names> <surname>Abu-Bakar</surname></string-name></person-group>, &#x201C;<article-title>U-ASD net: Supervised crowd counting based on semantic segmentation and adaptive scenario discovery</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>9</volume>, pp. <fpage>127444</fpage>&#x2013;<lpage>127459</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ali</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Shah</surname></string-name></person-group>, &#x201C;<article-title>A lagrangian particle dynamics approach for crowd flow segmentation and stability analysis</article-title>,&#x201D; in <conf-name>IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Minneapolis, MN, USA</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Shao</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Change Loy</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Scene-independent group profiling in crowd</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Columbus, OH, USA</conf-loc>, pp. <fpage>2219</fpage>&#x2013;<lpage>2226</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Bian</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Tian</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tang</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Tao</surname></string-name></person-group>, &#x201C;<article-title>A survey on trajectory clustering analysis</article-title>,&#x201D; <italic>arXiv Prepr. arXiv1802.06971</italic>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>R. W.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Xiong</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Wu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A dimensionality reduction-based multi-step clustering method for robust vessel trajectory analysis</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>17</volume>, no. <issue>8</issue>, pp. <fpage>1792</fpage>, <year>2017</year>; <pub-id pub-id-type="pmid">28777353</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Irfan</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Munsif</surname></string-name></person-group>, &#x201C;<article-title>Deepdive: A learning-based approach for virtual camera in immersive contents</article-title>,&#x201D; <source>Virtual Reality &#x0026; Intelligent Hardware</source>, vol. <volume>4</volume>, no. <issue>3</issue>, pp. <fpage>247</fpage>&#x2013;<lpage>262</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Ullah</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ullah</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Uzair</surname></string-name></person-group>, &#x201C;<article-title>A hybrid social influence model for pedestrian motion segmentation</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>31</volume>, no. <issue>11</issue>, pp. <fpage>7317</fpage>&#x2013;<lpage>7333</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. M.</given-names> <surname>Nguyen</surname></string-name> and <string-name><given-names>Q. J.</given-names> <surname>Wu</surname></string-name></person-group>, &#x201C;<article-title>A consensus model for motion segmentation in dynamic scenes</article-title>,&#x201D; <source>IEEE Transactions on Circuits and Systems for Video Technology</source>, vol. <volume>26</volume>, no. <issue>12</issue>, pp. <fpage>2240</fpage>&#x2013;<lpage>2249</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Seyedhosseini</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Tasdizen</surname></string-name></person-group>, &#x201C;<article-title>Semantic image segmentation with contextual hierarchical models</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>38</volume>, no. <issue>5</issue>, pp. <fpage>951</fpage>&#x2013;<lpage>964</lpage>, <year>2015</year>; <pub-id pub-id-type="pmid">26336116</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Narayana</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Hanson</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Learned-Miller</surname></string-name></person-group>, &#x201C;<article-title>Coherent motion segmentation in moving camera videos using optical flow orientations</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Int. Conf. on Computer Vision</conf-name>, <conf-loc>Sydney, NSW, Australia</conf-loc>, pp. <fpage>1577</fpage>&#x2013;<lpage>1584</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Meunier</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Badoual</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Bouthemy</surname></string-name></person-group>, &#x201C;<article-title>EM-driven unsupervised learning for efficient motion segmentation</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>45</volume>, no. <issue>4</issue>, pp. <fpage>4462</fpage>&#x2013;<lpage>4473</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Choudhury</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Karazija</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Laina</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Vedaldi</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Rupprecht</surname></string-name></person-group>, &#x201C;<article-title>Guess what moves: Unsupervised video and image segmentation by anticipating motion</article-title>,&#x201D; <italic>arXiv Prepr. arXiv2205.07844</italic>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Mukherjee</surname></string-name> and <string-name><given-names>Q. M. J.</given-names> <surname>Wu</surname></string-name></person-group>, &#x201C;<article-title>Streaming spatio-temporal video segmentation using Gaussian mixture model</article-title>,&#x201D; in <conf-name>IEEE Int. Conf. on Image Processing (ICIP)</conf-name>, <conf-loc>Paris, France</conf-loc>, pp. <fpage>4388</fpage>&#x2013;<lpage>4392</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Milan</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Leal-Taix&#x00E9;</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Schindler</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Reid</surname></string-name></person-group>, &#x201C;<article-title>Joint tracking and segmentation of multiple targets</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Boston, MA, USA</conf-loc>, pp. <fpage>5397</fpage>&#x2013;<lpage>5406</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. C.</given-names> <surname>Shadden</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Lekien</surname></string-name> and <string-name><given-names>J. E.</given-names> <surname>Marsden</surname></string-name></person-group>, &#x201C;<article-title>Definition and properties of lagrangian coherent structures from finite-time lyapunov exponents in two-dimensional aperiodic flows</article-title>,&#x201D; <source>Physica D: Nonlinear Phenomena</source>, vol. <volume>212</volume>, no. <issue>3&#x2013;4</issue>, pp. <fpage>271</fpage>&#x2013;<lpage>304</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Pan</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Pedrycz</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Ren</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Song</surname></string-name></person-group>, &#x201C;<article-title>Viewpoint-based kernel fuzzy clustering with weight information granules</article-title>,&#x201D; <source>IEEE Transactions on Emerging Topics in Computational Intelligence</source>, vol. <volume>7</volume>, no. <issue>2</issue>, pp. <fpage>342</fpage>&#x2013;<lpage>356</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Measuring crowd collectiveness</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>36</volume>, no. <issue>8</issue>, pp. <fpage>1586</fpage>&#x2013;<lpage>1599</lpage>, <year>2014</year>; <pub-id pub-id-type="pmid">26353340</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Tomasi</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Kanade</surname></string-name></person-group>, &#x201C;<article-title>Detection and tracking of point</article-title>,&#x201D; <source>Int. J. Comput. Vis.</source>, vol. <volume>9</volume>, pp. <fpage>137</fpage>&#x2013;<lpage>154</lpage>, <year>1991</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Ye</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Zhao</surname></string-name></person-group>, &#x201C;<article-title>Coherent motion detection with collective density clustering</article-title>,&#x201D; in <conf-name>Proc. of the 23rd ACM Int. Conf. on Multimedia</conf-name>, <conf-loc>Canada</conf-loc>, pp. <fpage>361</fpage>&#x2013;<lpage>370</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Mi</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A diffusion and clustering-based approach for finding coherent motions and understanding crowd scenes</article-title>,&#x201D; <source>IEEE Transactions on Image Processing</source>, vol. <volume>25</volume>, no. <issue>4</issue>, pp. <fpage>1674</fpage>&#x2013;<lpage>1687</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. K.</given-names> <surname>Pai</surname></string-name>, <string-name><given-names>A. K.</given-names> <surname>Karunakar</surname></string-name> and <string-name><given-names>U.</given-names> <surname>Raghavendra</surname></string-name></person-group>, &#x201C;<article-title>Scene-independent motion pattern segmentation in crowded video scenes using spatio-angular density-based clustering</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>145984</fpage>&#x2013;<lpage>145994</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. R.</given-names> <surname>Saufi</surname></string-name>, <string-name><given-names>Z. A.</given-names> <surname>Bin Ahmad</surname></string-name>, <string-name><given-names>M. S.</given-names> <surname>Leong</surname></string-name> and <string-name><given-names>M. H.</given-names> <surname>Lim</surname></string-name></person-group>, &#x201C;<article-title>Low-speed bearing fault diagnosis based on ArSSAE model using acoustic emission and vibration signals</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>46885</fpage>&#x2013;<lpage>46897</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Rodriguez</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Sivic</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Laptev</surname></string-name> and <string-name><given-names>J. -Y.</given-names> <surname>Audibert</surname></string-name></person-group>, &#x201C;<article-title>Data-driven crowd analysis in videos</article-title>,&#x201D; in <conf-name>2011 Int. Conf. on Computer Vision</conf-name>, <conf-loc>Barcelona, Spain</conf-loc>, pp. <fpage>1235</fpage>&#x2013;<lpage>1242</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Tang</surname></string-name></person-group>, &#x201C;<article-title>Understanding collective crowd behaviors: Learning a mixture model of dynamic pedestrian-agents</article-title>,&#x201D; in <conf-name>IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Providence, RI, USA</conf-loc>, pp. <fpage>2871</fpage>&#x2013;<lpage>2878</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M. J.</given-names> <surname>Zaki</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Meira</surname> <suffix>Jr</suffix></string-name></person-group>, <source>Data Mining and Machine Learning: Fundamental Concepts and Algorithms</source>. <publisher-loc>Cambridge, UK:</publisher-loc> <publisher-name>Cambridge University Press</publisher-name>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Nie</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Detecting coherent groups in crowd scenes by multiview clustering</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>42</volume>, no. <issue>1</issue>, pp. <fpage>46</fpage>&#x2013;<lpage>58</lpage>, <year>2018</year>; <pub-id pub-id-type="pmid">30307858</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Ge</surname></string-name>, <string-name><given-names>R. T.</given-names> <surname>Collins</surname></string-name> and <string-name><given-names>R. B.</given-names> <surname>Ruback</surname></string-name></person-group>, &#x201C;<article-title>Vision-based analysis of small groups in pedestrian crowds</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>34</volume>, no. <issue>5</issue>, pp. <fpage>1003</fpage>&#x2013;<lpage>1016</lpage>, <year>2012</year>; <pub-id pub-id-type="pmid">21844622</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Tang</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Coherent filtering: Detecting coherent motions from crowd clutters</article-title>,&#x201D; in <conf-name>European Conf. on Computer Vision</conf-name>, <conf-loc>Florence, Italy</conf-loc>, pp. <fpage>857</fpage>&#x2013;<lpage>871</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Anchor-based group detection in crowd scenes</article-title>,&#x201D; in <conf-name>IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>, <conf-loc>USA</conf-loc>, pp. <fpage>1378</fpage>&#x2013;<lpage>1382</lpage>, <year>2017</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>