<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">47903</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.047903</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Novel Locomotion Rule Rmbedding Long Short-Term Memory Network with Attention for Human Locomotor Intent Classification Using Multi-Sensors Signals</article-title>
<alt-title alt-title-type="left-running-head">A Novel Locomotion Rule Rmbedding Long Short-Term Memory Network with Attention for Human Locomotor Intent Classification Using Multi-Sensors Signals</alt-title>
<alt-title alt-title-type="right-running-head">A Novel Locomotion Rule Rmbedding Long Short-Term Memory Network with Attention for Human Locomotor Intent Classification Using Multi-Sensors Signals</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Shen</surname><given-names>Jiajie</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Wang</surname><given-names>Yan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>wy6868@jlu.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Dongxu</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Key Laboratory of Symbol Computation and Knowledge Engineering, Ministry of Education, Colleague of Computer Science and Technology, Jilin University</institution>, <addr-line>Changchun, 130012</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>College of Software, and Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University</institution>, <addr-line>Changchun, 130012</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yan Wang. Email: <email>wy6868@jlu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>20</day>
<month>6</month>
<year>2024</year></pub-date>
<volume>79</volume>
<issue>3</issue>
<fpage>4349</fpage>
<lpage>4370</lpage>
<history>
<date date-type="received">
<day>21</day>
<month>11</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>11</day>
<month>4</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Shen, Wang and Zhang</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Shen, Wang and Zhang</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_47903.pdf"></self-uri>
<abstract>
<p>Locomotor intent classification has become a research hotspot due to its importance to the development of assistive robotics and wearable devices. Previous work have achieved impressive performance in classifying steady locomotion states. However, it remains challenging for these methods to attain high accuracy when facing transitions between steady locomotion states. Due to the similarities between the information of the transitions and their adjacent steady states. Furthermore, most of these methods rely solely on data and overlook the objective laws between physical activities, resulting in lower accuracy, particularly when encountering complex locomotion modes such as transitions. To address the existing deficiencies, we propose the locomotion rule embedding long short-term memory (LSTM) network with Attention (LREAL) for human locomotor intent classification, with a particular focus on transitions, using data from fewer sensors (two inertial measurement units and four goniometers). The LREAL network consists of two levels: One responsible for distinguishing between steady states and transitions, and the other for the accurate identification of locomotor intent. Each classifier in these levels is composed of multiple-LSTM layers and an attention mechanism. To introduce real-world motion rules and apply constraints to the network, a prior knowledge was added to the network via a rule-modulating block. The method was tested on the ENABL3S dataset, which contains continuous locomotion date for seven steady and twelve transitions states. Experimental results showed that the LREAL network could recognize locomotor intents with an average accuracy of 99.03% and 96.52% for the steady and transitions states, respectively. It is worth noting that the LREAL network accuracy for transition-state recognition improved by 0.18% compared to other state-of-the-art network, while using data from fewer sensors.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Lower-limb prosthetics</kwd>
<kwd>deep neural networks</kwd>
<kwd>motion classification</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>62072212</award-id>
<award-id>62302218</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Development Project of Jilin Province of China</funding-source>
<award-id>20220508125RC</award-id>
<award-id>20230201065GX</award-id>
<award-id>20240101364JC</award-id>
</award-group>
<award-group id="awg3">
<funding-source>National Key R&#x0026;D Program</funding-source>
<award-id>2018YFC2001302</award-id>
</award-group>
<award-group id="awg4">
<funding-source>Jilin Provincial Key Laboratory of Big Data Intelligent Cognition</funding-source>
<award-id>20210504003GH</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Millions of people global suffer from severe disabilities, which can drastically decline the quality of life [<xref ref-type="bibr" rid="ref-1">1</xref>]. Fortunately, with the development of robotics and wearable sensor technology, disabled people can now live relatively comfortably with the help of wearable assistive devices [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>]. To work effectively and ensure the safety of users, devices such as wearable robots or exoskeletons need to recognize the locomotor intent of the user precisely. However, it can be difficult for the devices to distinguish human locomotor intent as it can not be observed directly. Therefore wearable robots identify locomotion by using real-time data from wearable sensors, such as electromyography (EMG) [<xref ref-type="bibr" rid="ref-3">3</xref>], inertial measurement units (IMUs) [<xref ref-type="bibr" rid="ref-4">4</xref>], load cell [<xref ref-type="bibr" rid="ref-5">5</xref>], pressure sensor [<xref ref-type="bibr" rid="ref-6">6</xref>], goniometer (GONIO) [<xref ref-type="bibr" rid="ref-7">7</xref>], and fused signals from multiple sources [<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>Several studies have attempted to enable human locomotor intent detection by robots, via traditional machine learning (ML) approaches, such as support vector machines (SVMs) [<xref ref-type="bibr" rid="ref-9">9</xref>], linear discriminant analysis (LDA) [<xref ref-type="bibr" rid="ref-10">10</xref>], quadratic discriminant analysis (QDA) [<xref ref-type="bibr" rid="ref-11">11</xref>], and artificial neural networks (ANNs) [<xref ref-type="bibr" rid="ref-12">12</xref>] and achieved high prediction accuracies. While these methods can recognize locomotor intent successfully, they require extensive feature engineering in advance. Moreover, when the situation becomes complex or the decision boundary is fuzzy when facing transitions, the traditional ML method performances decline [<xref ref-type="bibr" rid="ref-13">13</xref>]. Compared to classical ML methods, deep learning (DL) methods, such as convolutional neural networks (CNNs) [<xref ref-type="bibr" rid="ref-13">13</xref>], recurrent neural networks (RNNs) [<xref ref-type="bibr" rid="ref-14">14</xref>], gated recurrent units (GRUs) [<xref ref-type="bibr" rid="ref-15">15</xref>], can predict complex locomotion modes with minimal feature engineering. Recently, deep reinforcement learning (DRL) has garnered significant attention an interactive ML paradigm that seamlessly integrates deep neural networks (DNNs) [<xref ref-type="bibr" rid="ref-16">16</xref>] into the well-established conventional reinforcement learning (RL) framework [<xref ref-type="bibr" rid="ref-17">17</xref>]. Many DRL approaches, such as the Markov decision process (MDP) and self-organizing networks have obtained remarkable achievements in several fields [<xref ref-type="bibr" rid="ref-18">18</xref>&#x2013;<xref ref-type="bibr" rid="ref-20">20</xref>], offering new possibilities and advancements in the field of lower-limb prosthetics. However, RL typically requires a large number of interaction samples for training, which can be a challenge in the real world, it still needs further development before being applied to practical scenarios.</p>
<p>Using the time history information on locomotion sequence, long short-term memory (LSTM) networks have demonstrated good results for classifying human locomotion [<xref ref-type="bibr" rid="ref-21">21</xref>]. LSTM concentrates on the long-range information of the previous moments for accurate recognition, especially on steady states. However, this method may not perform as well when facing transitions as they are often situated between two steady states, and most of their historical information would be is similar to that of the adjacent state. As a result, raw LSTM may not be able to focus on crucial time instances to accurately classify transition modes.</p>
<p>Attention mechanism (AM) has been an emerging research direction that has been applied to several fields such as natural language processing (NLP) and computer vision in recent years for its ability to focus on key parts of temporal information [<xref ref-type="bibr" rid="ref-22">22</xref>]. However, AM has not yet been widely used or explored in human locomotion mode prediction. Zhu et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed a knee/ankle joint angle prediction method based on attention-based CNN-LSTM model. They deployed the AM before the LSTM layers to extract detailed information from the CNN while suppressing useless information. As the attention layer in this model is located before LSTM layers, the struct of the latter becomes N-to-1, i.e., the LSTM layers obtain a result directly from the input. Therefore, the AM is mainly used for spatial feature selection rather than temporal characteristics.</p>
<p>Despite their advantages, most DNNs for human locomotor prediction are purely data-driven, and lack an understanding of locomotion rules. This can lead to low model accuracy, especially when meeting complex locomotion modes (e.g., transitions). Rule-based methods such as state machines [<xref ref-type="bibr" rid="ref-24">24</xref>] and decision trees have been applied to address the issue. However, to achieve high accuracies similar to that of DNNs, rule-based methods require a deep understanding of the experimental data to obtain the rules. Rule designing becomes increasingly complex when meeting highly diverse locomotion modes. Considering the respective advantages of DNNs and rule-based methods, it may be helpful to combine a priori knowledge with the former for the optimal utilization of the two.</p>
<p>In response to the existing deficiencies, in this study, we proposed a locomotion rule embedding LSTM network with AM (LREAL), a system to identify human locomotor intent precisely, based on IMU and GONIO sensor signals. We introduced the AM after the LSTM layer to redistribute the attention of different features at different times. This creates an N-to-N LSTM layer struct, which is maintained by its the output. This attention layer placement allows the model to focus on the complete spatiotemporal characteristics. To drive the network by using both data and rules (instead of only data), inspired by rule-embedded neural networks (ReNN) [<xref ref-type="bibr" rid="ref-25">25</xref>], we introduced the locomotion information to the model by embedding the rule in the network. Specifically, the LREAL network is composed of two levels: The first is a steady/transition state classifier and the second is composed of two classifiers to accurately recognize the locomotor intent of the states. Each classifier in these levels is combined with LSTM layers and an attention layer. Initially, continuous raw IMU and GONIO data are processed into fixed-size fragments for input into the structure. After the first level of classification, the fragments from steady locomotion modes are input into the classifier for steady locomotion to classify steady locomotor intent. It further passes a result in the form of a probability vector to a rule modulating block. In this block, the rule will be combined with the probability vector to obtain an encoding vector for transitions. As transitions follow steady locomotion, the features obtained by the rule-modulating block are directly input into the attention layer of the transitions-state classifier for prediction. The proposed model performance vastly improved human locomotor intent recognition compared to CNN, LSTM, and other ML methods. The contributions of this study can be summarized as follows:
<list list-type="bullet">
<list-item>
<p>We introduced AM to the LSTM network, enhancing the ability of the model to capture key instances in time and effectively extract locomotion-related features of the sensor data.</p></list-item>
<list-item>
<p>We combined LSTM layers and an AM to build a classifier (AT-LSTM) that enables complete learning from two IMU and four GONIO datasets.</p></list-item>
<list-item>
<p>We introduced a priori knowledge via a rule modulating block to the network to constrain the decision scope when meeting the transition states, and demonstrate the added effectiveness via ablation studies.</p></list-item>
<list-item>
<p>We constructed the LREAL work with a multi-level architecture composed of AT-LSTM for human locomotor intent and validated its efficacy using ENABL3S dataset [<xref ref-type="bibr" rid="ref-26">26</xref>] through accuracy scores. Our results using fewer sensors were comparable to or better than the existing models.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Review</title>
<p>In this section, we discuss the related works and technologies on: Sensors, locomotor intent classification and ReNNs.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Sensors</title>
<p>Sensors play a crucial role in the design of exoskeletons and wearable robots as they greatly influence reliability and cost. Wearable robots most commonly use IMU sensors which record 3-axis acceleration, angular velocity, and gyroscope signals [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>]. IMUs are preferred as they provide reliable data and are cost-effective. EMG sensors are employed for human activity recognition (HAR) as they can measure electrical signals generated from muscle contractions during physical activities [<xref ref-type="bibr" rid="ref-3">3</xref>]. While EMG data can significantly enhance their performance wearable robots are relatively expensive and inconvenient to wear. Other types of sensors, such as load cells, pressure sensors, (GONIO), are also used. Furthermore, the fusion of signals from multiple sources has also been widely applied to leverage their respective advantages [<xref ref-type="bibr" rid="ref-8">8</xref>]. The selection of sensor types and data processing methods can greatly impact the locomotor intent classifier performance, and finding ways to achieve better results with fewer sensors has been the primary research focal point.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Locomotor Intent Classification</title>
<p>Motivated by their successful application in various fields, DL methods have gained significant popularity in locomotor intent classification. CNNs, known for their effective feature learning in CV, have been applied to locomotor intent classification by similarly processing sensor data [<xref ref-type="bibr" rid="ref-27">27</xref>]. While CNNs offer advantages such as parameter sharing and sparse connections, they require higher computational resources. LSTM networks, known for their ability to process time series data, haves shown promising results in classifying human locomotion [<xref ref-type="bibr" rid="ref-21">21</xref>]. To enhance efficiency, studies have explored CNN-RNN hybrid networks. Wang et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed a CNN-LSTM hybrid approach, achieving a validation accuracy of 95.90% in recognizing six types of activities (walking, lying, sitting, standing, and stair ascent, and stair descent) using a smartphone IMU. Zhu et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed an attention-based CNN-LSTM model for knee/ankle joint angle movement prediction, using AM for feature extraction. In contrast to [<xref ref-type="bibr" rid="ref-23">23</xref>], the AM in LREAL was introduced after the LSTM layer to dynamically redistribute attention across different features at different time steps, aiming for complete spatiotemporal feature selection.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>ReNN</title>
<p>The interpretability of DNNs is often criticized due to the difficulty in explaining inferences as concise interactions among parameters and the network [<xref ref-type="bibr" rid="ref-29">29</xref>]. Unlike human inferences based on experiences and knowledge, most DNNs are purely data-driven. This lack of prior knowledge can result in low interpretability and potential errors in the models [<xref ref-type="bibr" rid="ref-30">30</xref>]. To address these issues, knowledge representation has been utilized to describe the richness of the world in computer systems, allowing artificial intelligence to understand and utilize it for reasoning and inference. Rule-based representation, a common formalism of knowledge representation, has been widely employed in expert systems and achieved excellent results [<xref ref-type="bibr" rid="ref-31">31</xref>]. Considering the aforementioned challenges and inspired by the concept of rules embedded in neural networks [<xref ref-type="bibr" rid="ref-25">25</xref>], we propose a rule-modulating block that incorporates locomotion rules into a knowledge representation and combines it with DNNs. This approach aims to enhance interpretability and improve the model performance by leveraging both data-driven and rule-based learning.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Dataset and Preprocessing Steps</title>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset</title>
<p>We used a the ENABL3S public benchmark dataset to validate the proposed method. The dataset consists of data from IMU sensors (placed on the thigh and shank of the subjects), GONIO sensors (placed on knee and ankle of the subjects) and EMG sensors, collected from 10 able-bodied subjects (seven males and three females) with an average age of (25.5 &#x00B1; 2) y, height of (174 &#x00B1; 12) cm, and weight of (70 &#x00B1; 14) kg. The subjects move according to the required mode [sit (S), stand (ST), ground-level walking (LW), ramp ascending/descending (RA/RD), and stair ascent/descent (SA/SD)]. The ramps have a 10&#x00B0; slope slopes of 10&#x00B0; and the stairs consist of four steps. The odd-numbered moving sequence is: S &#x2192; ST &#x2192; LW &#x2192; SA &#x2192; LW &#x2192; RD &#x2192; LW &#x2192; ST &#x2192; S; and the even-numbered moving sequence is: S &#x2192; ST &#x2192; LW &#x2192; RA &#x2192; LW &#x2192; SD &#x2192; LW &#x2192; ST &#x2192; S. There are seven kinds of steady locomotion modes in ENABL3S with recorded true labels. Each data sample contains the signals from five IMU sensors recording six data channels (3 axis acceleration and angular velocity) and four GONIO sensors. By using a key fob, the start and end of the steady locomotion modes are recorded and the data is labeled according to the true moving mode. Although only labels of steady locomotion modes are available in ENABL3S, the information on the transition modes can be obtained by extracting the data between the steady modes as ENABL3S marks the time of the beginning and ends of the steady states.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Data Preprocessing</title>
<p>The IMU data underwent low-pass filtering at 20 Hz to eliminate high-frequency noise. In the ENABL3S dataset, the IMU and GONIO data are sampled at 500 Hz (i.e., sampling once every 2 ms). Every 250 samples (2 ms per sample) were segmented into a 500 ms analysis window with a sliding window of 50 ms. We used the locomotion mode at the end of each gait event to obtain the label of each analysis window. If the beginning and end of the gait event maintained the same locomotion mode, analysis windows for the gait event were directly labeled as &#x201C;steady locomotion intent&#x201D; (<xref ref-type="fig" rid="fig-1">Fig. 1</xref>). Gait events with different modes can experience two different situations. If the analysis window was located (i) fully within a transition gait event, it was labeled according to the starting and ending mode and (ii) partially within a transition gait event of transitions, it was labeled based on the majority of its location (&#x003E;50%; <xref ref-type="table" rid="table-1">Table 1</xref>). We utilized the acceleration and angular velocity data of only two IMU sensors on the upper and lower right leg in 3 axes (X, Y and Z axes) and GONIO data, so the initial input of the network can be denoted as <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mn>500</mml:mn><mml:mo>&#x2217;</mml:mo><mml:mn>12</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The above data was divided as follow: 90% for training and 10% for testing. Additionally, 20% of the training set was also used for validation to prevent the network from over-fitting.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Sequence extraction visualization and labeling: The yellow windows refer to the steady LW locomotor intent, as they are located fully within a steady gait event; the blue windows refer to the LW-RA transition, as they are located fully within a transition gait event; and the green window refers to steady RA locomotor intent as the start and end of the window fails in the same locomotor mode</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-1.tif"/>
</fig><table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Analysis windows extracted for each label from the ENABL3S dataset; each window is composed of 250 samples and each sample has 12 features</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Label</th>
<th>Number</th>
<th>Label</th>
<th>Number</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>S</bold></td>
<td>29089</td>
<td>SA-LW</td>
<td>1158</td>
</tr>
<tr>
<td><bold>LW</bold></td>
<td>45741</td>
<td>LW-RD</td>
<td>1805</td>
</tr>
<tr>
<td><bold>RA</bold></td>
<td>15346</td>
<td>RD-LW</td>
<td>1420</td>
</tr>
<tr>
<td><bold>RD</bold></td>
<td>17730</td>
<td>LW-ST</td>
<td>1556</td>
</tr>
<tr>
<td><bold>SA</bold></td>
<td>5641</td>
<td>ST-S</td>
<td>2065</td>
</tr>
<tr>
<td><bold>SD</bold></td>
<td>5499</td>
<td>LW-RA</td>
<td>1778</td>
</tr>
<tr>
<td><bold>ST</bold></td>
<td>6552</td>
<td>RA-LW</td>
<td>1410</td>
</tr>
<tr>
<td><bold>S-ST</bold></td>
<td>8282</td>
<td>LW-SD</td>
<td>1638</td>
</tr>
<tr>
<td><bold>ST-LW</bold></td>
<td>4045</td>
<td>SD-LW</td>
<td>1392</td>
</tr>
<tr>
<td><bold>LW-SA</bold></td>
<td>1390</td>
<td></td>
<td></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Methods</title>
<sec id="s4_1">
<label>4.1</label>
<title>Network Architecture</title>
<p>The proposed LREAL network is composed of two levels (<xref ref-type="fig" rid="fig-2">Fig. 2</xref>). The firstlevel, i.e., the steady/transition classifier, is responsible for initially classifying the input sequence and dividing it into the two main states. This classifier treats both steady and transition states equally. Its purpose is to serve as a pre-classification step for the subsequent level.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Schematic representation LREAL network for human locomotor intent recognition</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-2.tif"/>
</fig>
<p>The second level consists of two distinct classifiers and a rule-modulating block. The input sequence divided by the first-level classifier is directed to the appropriate classifier based on the assigned state label. The steady-state classifier receives the fragments from the steady-state portion of the sequence as its input. The output of this classifier is stored and passed onto the rule-modulating block as auxiliary information, which helps in preparing for the recognition of transitions.</p>
<p>As mentioned previously, transitions occur between steady states, i.e., there always exists a preceding steady state that significantly influences the range of post-transition behavior. There also exist specific rules governing the motion transformations, such as considering the previous locomotion mode when encountering a transition [<xref ref-type="bibr" rid="ref-32">32</xref>]. Incorporating these rules and prior information from the classifier for steady locomotion into the transition-state classifier does not affect practical applications or add computational cost. Therefore, the rule-modulating block is designed to introduce motion transformation rules to the network. It receives transmission information from the steady-state classifier to obtain the preceding condition for each window for transitions. This information is combined with the motion transformation rules to assist the transition-state classifier in recognition. This process embeds the provided motion conversion rules with auxiliary information from the preceding steady state, resulting in an output vector that represents refined information.</p>
<p>Finally, the transition-state classifier uses the sequence of recognized transitions from the first-level classifier and refined information from the rule-modulating block to generate the final output. Note that the classification of steady states solely relies on the data, while the transition recognition combines the data with rule-embedded information. This integration of data and rule-based information enhances the of transition recognition accuracy.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>AM Design</title>
<p>As different samples within a specific analysis window may contribute differently to the recognition of various activities, the proposed AT-LSTM model employs an AM to capture the relationship between the samples and locomotion classes (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>). The attention layer first derives a vector representation for the features of each sample and maps to a probability value indicating the likelihood of the analysis window belonging to a certain locomotion class. Formally, we constructed the attention layer as follows.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Attention layer structure in the proposed LREAL model. Here, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denotes the input vector of the <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>k</mml:mi></mml:math></inline-formula>th sample containing <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>n</mml:mi></mml:math></inline-formula> features</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-3.tif"/>
</fig>
<p>We first formulated the attention importance score, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, for each feature in the <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>k</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> sample by multiplying the input vector, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, with a weight matrix, <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and normalizing the result to a probability which represents the degree of relevance between the <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>i</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> feature and the locomotion mode of the <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>k</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> sample. It can be expressed as
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>To concentrate on the feature instead of focusing on a specific sample, each feature of a sample is assigned a separate importance score instead of a fixed score. Then we can denote the newly-constructed feature vectors of <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:math></inline-formula> feature, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, as
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula>which indicates the degree of relevance between the <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>i</mml:mi></mml:math></inline-formula>th feature of <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>k</mml:mi></mml:math></inline-formula>th sample and the locomotion class. By flattening <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> ,we then combine the different features of the sample within the feature vector, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, as
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>In conclusion, the attention layer helps the network to better predict the locomotion classes by focusing on the key sample of the analysis window.</p>
<p>After obtaining <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the probability vector, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, for analysis window classification can be formulated as
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">f</mml:mi><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">m</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>w</italic><sub><italic>p</italic></sub> and <italic>b</italic><sub><italic>p</italic></sub> are the weight matrix and bias vector, respectively.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Attention-Based LSTM Design</title>
<p>The input of the attention-based LSTM network (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>) is an analysis window of dimension 250 &#x00D7; 12. The three LSTM layers have 128, 64, and 32 units, respectively, and the LSTM layers will keep the structure of N-to-N, which means the output of the final LSTM layer remains a sequence. The attention layer takes the sequence as the input (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>). Finally, we add a dense layer with 100 units and a softmax layer that obtains the final classification result. The size of the softmax layer is designed based on tasks. Specifically, the layer size is set to (i) 2 when the classifier is designed to distinguish the steady and transition states, and (ii) 7/12 when recognizing the steady and transition states in detail.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>AT-LSTM structure of the attention-based LSTM. The dense layer size d is fixed based on the actual problem</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-4.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Rule-Modulating Block Design</title>
<p>The rule-modulating block was designed based on the locomotion transformation information from the ENABL3S dataset. It combines this information with that from the steady-state classifier and passes it to the transition-state classifier for recognition. The rule was designed based on the previous locomotion mode before the transition, as we know that each steady state can only have limited options for subsequent transition states (e.g., for the S steady state, only the S-ST transition is possible; and for the ST steady state, only the ST-S and ST-LW transitions are possible). All rules were obtained based on the actual locomotion mode in ENABL3S and encoded as within the <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> matrix in the network (<xref ref-type="table" rid="table-2">Table 2</xref>). The rows of the matrix represent the previous steady locomotion state and the columns represent the subsequent transition. If the transformation from the previous locomotion to the subsequent transition is valid(invalid), its value in the matrix is set to 1(0).</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Locomotion mode conversion rules in the ENABL3S dataset. The rows represent the previous steady locomotion and the columns represent the subsequent transition; &#x2018;&#x221A;&#x2019; represents valid conversions</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th></th>
<th>S-ST</th>
<th>ST-LW</th>
<th>LW-SA</th>
<th>SA-LW</th>
<th>LW-RD</th>
<th>RD-LW</th>
<th>LW-ST</th>
<th>ST-S</th>
<th>LW-RA</th>
<th>RA-LW</th>
<th>LW-SD</th>
<th>SD-LW</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>S</bold></td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td><bold>LW</bold></td>
<td></td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td>&#x221A;</td>
<td></td>
</tr>
<tr>
<td><bold>RA</bold></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td></td>
</tr>
<tr>
<td><bold>RD</bold></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td><bold>SA</bold></td>
<td></td>
<td></td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td><bold>SD</bold></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td>&#x221A;</td>
</tr>
<tr>
<td><bold>ST</bold></td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>After obtaining the conversion rules, the rule-modulating block combines it with the probability vector, <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, of the previous analysis window from the steady-state classifier as
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>7</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mspace width="1em" /><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>12</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula></p>
<p>The obtained encoded information vector <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2217;</mml:mo><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, which represents the probability of the next transition, is passed on to the transition-state classifier as input. To prevent the encoded information from being diluted, it is directly input into the attention layer along with the hidden layer output vector from the previous LSTM layer to construct the final transition-state classifier (<xref ref-type="fig" rid="fig-5">Fig. 5</xref>).</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>AT-LSTM structure for transitions</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Setup and Discussions</title>
<p>In this section, we present the conducted experiments and obtained results to validate the efficacy of the proposed LREAL method. All experiments were conducted using the Keras API, with the training performed on a system with an NVIDIA RTX 3080 and AMD Ryzen 9 5900 HX processor with 32 GB RAM. We evaluated the overall network performance when handling mixed data, as well as that of each classifier in LREAL when dealing with pure steady-state or transition-state scenarios. Additionally, we compared our results with those of the existing classification methods to demonstrate its accuracy.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Experimental Setup</title>
<p>We utilized an initial set of hyperparameters and investigated their impact on the model to narrow down the range for the subsequent grid search (<xref ref-type="table" rid="table-3">Table 3</xref>). A grid search with a 5-fold cross-validation was then performed on the hyperparameters to maximize validation accuracy (<xref ref-type="table" rid="table-4">Table 4</xref>; a detailed discussion on the influence of the hyperparameters can be found in <xref ref-type="sec" rid="s5_3_6">Section 5.3.6</xref>). The configuration shown in <xref ref-type="table" rid="table-4">Table 4</xref> was selected as the optimal choice for subsequent experiments involving the classifier. To deal with the imbalanced distribution of the different classes in the ENABL3S dataset (<xref ref-type="table" rid="table-1">Table 1</xref>), we adjusted the weight of each class for training inversely proportional to class frequency in the dataset as</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Hyperparameters values</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Epoch</td>
<td>200</td>
</tr>
<tr>
<td>Learning rate</td>
<td>0.0001</td>
</tr>
<tr>
<td>Batch size</td>
<td>256</td>
</tr>
<tr>
<td>Optimizer</td>
<td>Adam [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
</tr>
<tr>
<td>Loss function</td>
<td>Cross-entropy</td>
</tr>
<tr>
<td>Class-weighting</td>
<td>True</td>
</tr>
<tr>
<td>Early stopping</td>
<td>10 epochs if no improvement on validation loss</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Grid search results</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Classifier</th>
<th>Learning rate</th>
<th>Optimizer</th>
<th>Batch size</th>
</tr>
</thead>
<tbody>
<tr>
<td>Steady/transition</td>
<td>0.001</td>
<td>Adam</td>
<td>64</td>
</tr>
<tr>
<td>Steady state</td>
<td>0.001</td>
<td>Adam</td>
<td>64</td>
</tr>
<tr>
<td>Transition</td>
<td>0.0005</td>
<td>RMSprop</td>
<td>64</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mi mathvariant="italic">W</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">i</mml:mi><mml:mi mathvariant="italic">g</mml:mi><mml:mi mathvariant="italic">h</mml:mi><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">i</mml:mi><mml:mi mathvariant="italic">n</mml:mi><mml:mi mathvariant="italic">g</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">m</mml:mi><mml:mi mathvariant="italic">p</mml:mi><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">s</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">s</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">s</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">m</mml:mi><mml:mi mathvariant="italic">p</mml:mi><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">s</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">s</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>To validate the effectiveness of the proposed method, we devised two distinct experimental scenarios. The first scenario aimed to evaluate the classification performance of each LREAL component individually. Thus, the performance of the three classifiers was tested separately in classifying pure data for steady/transitions states. In the second scenario, we introduced mixed data and examined the performance of LREAL when considering the mutual influence of various components. Furthermore, in both scenarios, the networks were trained and tested on user-dependent and user-independent base.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Performance Evaluation</title>
<p>We used the accuracy (ACC) score to measure the overall performance of our proposed LREAL network; it can be expressed as</p>
<p><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">t</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></disp-formula>where <italic>N</italic><sub><italic>correct</italic></sub> is the number of correct classifications and <italic>N</italic><sub><italic>total</italic></sub> is the total number of data points. We also evaluated the network performance in two bases: (1) Data from one out of ten subjects was used to train and test a user-dependent network; (2) data from nine out of ten subjects was used as the training set, while the remaining subject was used as test set to construct a user-independent network. The processes were repeated 10 times until all subjects were used.</p>
<p>Furthermore, when evaluating the overall performance, the number of sensors utilized serves as a crucial metric. In terms of practical applications, achieving comparable results with fewer sensors signifies superior model performance.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Discussion</title>
<sec id="s5_3_1">
<label>5.3.1</label>
<title>Single Classifier Evaluation</title>
<p>To evaluate and analyze the performance of the three types of classifiers in the proposed network (<xref ref-type="fig" rid="fig-2">Fig. 2</xref>), we computed their average accuracies on the user-dependent and user-independent bases. The results of individual classifiers did not consider their mutual influence to evaluate their actual capacity for recognizing locomotor intent (<xref ref-type="table" rid="table-5">Table 5</xref>).</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance of the three types of LREAL classifiers</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Recognition classifier</th>
<th>User-dependent ACC</th>
<th>User-independent</th>
</tr>
<tr>
<th/>
<th>ACC (%)</th>
<th>ACC (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Steady/transition</td>
<td>98.56 &#x00B1; 0.35</td>
<td>96.36 &#x00B1; 2.12</td>
</tr>
<tr>
<td>Steady state</td>
<td>99.32 &#x00B1; 0.18</td>
<td>96.18 &#x00B1; 2.25</td>
</tr>
<tr>
<td>Transition</td>
<td>98.96 &#x00B1; 0.32</td>
<td>92.51 &#x00B1; 3.16</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The first-level steady/transition state classifier achieved an accuracy of (98.56% &#x00B1; 0.35%) and (96.42% &#x00B1; 2.12%) on the user-dependent and user-independent bases, respectively. A confusion matrix is a table that visualizes the performance of a classifier, comprising data of the true and predicted labels that the model evaluated. We obtained and analyzed the confusion matrix of the user-dependent classifier (<xref ref-type="fig" rid="fig-6">Fig. 6</xref>). Transition class obtained an accuracy of 97.50% much lower than that for the steady state class (99.61%). A previous study has shown that steady RA and RD locomotion can introduce a high error rate for the similarities between ramp and ground-level walking, as well as the transitions between these steady states [<xref ref-type="bibr" rid="ref-32">32</xref>]. The performance of the steady/transition classifier largely affects the end-stage classification accuracy and the overall network performance.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Confusion matrix for the steady/transition classifier on the user-dependent</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-6.tif"/>
</fig>
<p>A previous study developed a recognition system based on a Gaussian SVM and kinematic data which was able to classify five steady states and eight transition states [<xref ref-type="bibr" rid="ref-34">34</xref>], and the steady/transition classifier in it had an accuracy of 93.30%. Our framework was more accurate (&#x002B;5.26%) while our classifiers utilized the kinematic signals with the data from GONIO sensors. Although our steady/transition classifier demonstrated higher accuracy, further analysis is required to improve its effectiveness.</p>
<p>The steady-state classifier was effective on both user-independent (ACC &#x003D; 99.32%) and user-dependent (ACC &#x003D; 96.18%) bases. However, the transitions-state classifier was more accurate in the user-independent (ACC &#x003D; 98.96%) basis compared to the user-dependent (ACC &#x003D; 92.51%) basis. The errors were mainly due to the transitions between RD/RA and LW. The differences in individual exercise methods and the complexity of transitions are expected to explain this observation.</p>
<p>We also compared the user-dependent classifier with the existing method when facing pure transition or steady states data (<xref ref-type="table" rid="table-6">Tables 6</xref> and <xref ref-type="table" rid="table-7">7</xref>). Although our approach was slightly less effective in recognizing steady states compared to a multi-level SVM [<xref ref-type="bibr" rid="ref-34">34</xref>], it was more versatile (recognized seven types of steady states compared to the five types identified by the SVM) with similar efficacy. Furthermore, our approach considered more transition states and achieved a much higher accuracy than the multi-level SVM (&#x002B;3.06%). Compared to context-based Bayesian [<xref ref-type="bibr" rid="ref-35">35</xref>], central pattern generator (CPG) [<xref ref-type="bibr" rid="ref-36">36</xref>], and CNN-LSTM networks [<xref ref-type="bibr" rid="ref-37">37</xref>], our AT-LSTM network surpassed in terms of accuracy for both steady and transition-state recognition when considering more classes. These results validated the effectiveness of AT-LSTM in recognizing human locomotor intent.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparison of various methods for steady-state recognition</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Number of recognizable classes</th>
<th>Accuracy (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>SVM (multi-level) [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>5</td>
<td>99.4</td>
</tr>
<tr>
<td>Context-based Bayesian [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>7</td>
<td>98.3</td>
</tr>
<tr>
<td>CPG [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>5</td>
<td>99.03</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>7</td>
<td>99.1</td>
</tr>
<tr>
<td>AT-LSTM</td>
<td>7</td>
<td>99.32</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparison of various methods for transition-state recognition</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Number of recognizable classes</th>
<th>Accuracy (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>SVM (multi-level) [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>8</td>
<td>95.9</td>
</tr>
<tr>
<td>Context-based Bayesian [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>12</td>
<td>98.8</td>
</tr>
<tr>
<td>CPG [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>8</td>
<td>98.9</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>12</td>
<td>96.9</td>
</tr>
<tr>
<td>AT-LSTM</td>
<td>12</td>
<td><bold>98.96</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3_2">
<label>5.3.2</label>
<title>Overall Network Evaluation</title>
<p>The overall performance of the LREAL network was further evaluated by calculating the recognition accuracy across the entire network, i.e., considering the impact of all classifiers. The network achieves an average accuracy of 98.60% and 92.78% on the user-dependent and user-independent base, respectively. Additionally, the user-dependent (user-independent) network obtained an average accuracy of 99.03% (94.25%) and 96.52% (86.52%) for the steady and transition states, respectively. From the confusion matrix of the user-dependent LREAL network (<xref ref-type="fig" rid="fig-7">Fig. 7</xref>), it can be seen that the errors mainly belong to the recognition of transitions, including misclassification between transitions and between transition and steady states. Certain steady states were also misclassified as transition states. This can be attributed to the similarity between transition and steady states and partial locomotion.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Confusion matrix for the user-dependent LREAL network</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-7.tif"/>
</fig>
<p>On the other hand, it is worth noting that the scope of misclassification is narrow, as the rule-embedded information restrains the classification range of each class in the case of correct first level classification. However, our approach took a longer time to recognize locomotor intents as there were three classifiers working together. The average running time for the recognition of one sequence when classifying 10,000 sequences was (112.06 &#x00B1; 6.23) ms for both user-dependent and user-independent base.</p>
<p>Furthermore, we compared the performance of the LREAL network with previous methods using ENABL3S according to the recognizable classes, the number of used sensors, and accuracy (<xref ref-type="table" rid="table-8">Table 8</xref>). The results demonstrated that the LREAL network outperformed LDA [<xref ref-type="bibr" rid="ref-8">8</xref>], CNN-LSTM [<xref ref-type="bibr" rid="ref-37">37</xref>], CNN [<xref ref-type="bibr" rid="ref-38">38</xref>], and Light-Weight Artificial Neural Network (LWANN) [<xref ref-type="bibr" rid="ref-39">39</xref>] in terms of accuracy for steady-state recognition, and surpassed LIR-Net [<xref ref-type="bibr" rid="ref-27">27</xref>], LDA, CNN, LWANN, and CNN-LSTM in terms of accuracy for transition recognition. While LIR-Net achieves better overall performance and accuracy for steady-state recognition compared to LREAL, it only recognizes five steady states and eight types of transitions compared to the seven steady states and twelve types of transitions recognized by the LREAL network. Additionally, LIR-Net utilizes a larger number of sensors compared to LREAL. LREAL demonstrated superior performance in handling transitions, albeit with a slightly lower overall accuracy than that of LIR-Net due to fewer sensors (two IMU and four GONIO) and a larger number of recognizable locomotor intent classes. In summary, LREAL used fewer sensors to achieve a higher accuracy for a higher number of classes.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Performance comparison of the LREAL network with other methods on ENABL3S</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Sensors</th>
<th colspan="2" align="center">Classes</th>
<th colspan="3" align="center">Accuracy (%)</th>
</tr>
<tr>
<th/>
<th/>
<th>Number of steady states</th>
<th>Number of transitions</th>
<th>Steady states</th>
<th>Transitions</th>
<th>Overall</th>
</tr>
</thead>
<tbody>
<tr>
<td>LIR-Net [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>14 EMG, 4 IMU, 4 GONIO</td>
<td>5</td>
<td>8</td>
<td><bold>99.46</bold></td>
<td>96.34</td>
<td>98.89</td>
</tr>
<tr>
<td>LDA [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>14 EMG, 4 IMU, 4 GONIO</td>
<td>7</td>
<td>12</td>
<td>98.75</td>
<td>94.06</td>
<td>98.48</td>
</tr>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>14 EMG, 5IMU 4 GONIO</td>
<td>5</td>
<td>8</td>
<td>NULL</td>
<td>NULL</td>
<td>96.3</td>
</tr>
<tr>
<td>LWANN [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>5 IMU 4GONIO</td>
<td>7</td>
<td>12</td>
<td>NULL</td>
<td>NULL</td>
<td>98.60</td>
</tr>
<tr>
<td>CNN-LSTM [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>2 IMU</td>
<td>7</td>
<td>12</td>
<td>96.41</td>
<td>94.32</td>
<td>95.82</td>
</tr>
<tr>
<td>LREAL</td>
<td>2 IMU 4GONIO</td>
<td>7</td>
<td>12</td>
<td>99.03</td>
<td><bold>96.52</bold></td>
<td>98.60</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3_3">
<label>5.3.3</label>
<title>AM Visualization</title>
<p>To better understand and explain the internal operations of the AM the locomotor intent recognition task, we visualized its performance from two aspects.</p>
<p>Firstly, we analyzed the performance of AM on the features of a sample. We extracted the activations of the attention layer of the steady-state classifier as an example. Let us denote the simple average of hidden vectors from the previous LSTM layer as <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and their weighted sum after the attention layer as <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. To make the visualization clearer, we normalized the activations after obtaining them. In <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, each column represents a feature dimension, and the rows correspond to the feature vectors before and after the action of the attention layer. The attention layer modifies <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> by redistributing the weight of each feature.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Visualization of feature vectors before and after the action of the attention layer</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-8.tif"/>
</fig>
<p>We also analyzed the ability of the AM to focus on key instances of the time sequence for the steady states and transitions. Specifically, using the attention importance score for each feature, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, in the <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>k</mml:mi></mml:math></inline-formula>th sample, we calculated the average importance score, <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, for the <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>k</mml:mi></mml:math></inline-formula>th sample to study the preference of the AM for the overall analysis window. The attention layer exhibits different preferences for the tasks of steady and transition-state recognition (<xref ref-type="fig" rid="fig-9">Figs. 9a</xref> and <xref ref-type="fig" rid="fig-9">9b</xref>). The attention layer focused more on the intermediate fragment of the analysis window for steady-state recognition (<xref ref-type="fig" rid="fig-9">Fig. 9a</xref>), whereas it concentrates on both ends of the analysis window for transitions (<xref ref-type="fig" rid="fig-9">Fig. 9b</xref>). This indicated that the information that helps identify steady states is concentrated in the middle of the analysis window and the signals that dictate transitions are located on either of its end segments. Although previous studies reached similar findings for transitions by visualizing the activations of CNNs [<xref ref-type="bibr" rid="ref-40">40</xref>], they considered only the activations for transitions. Our work showed relied on the importance scores of both steady and transition states, which validates the capacity of the AM to capture valuable temporal information when facing different tasks. The finding is also critical for recognizing locomotor intent and is worth further study.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Relationship of the attention weights of each sample over the complete analysis window for (a) steady-state and (b) transition recognition</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-9.tif"/>
</fig>
</sec>
<sec id="s5_3_4">
<label>5.3.4</label>
<title>Impact of a Priori Knowledge</title>
<p>To analyze the impact of a priori knowledge in LREAL, we conducted tests without a priori knowledge using pure transition data. The resulting confusion matrix clearly illustrated the higher tendency for misrecognition of reversed transition states without a priori knowledge (<xref ref-type="fig" rid="fig-10">Fig. 10</xref>). For instance, the LW-RA transitions were easily misclassified as RA-LW transitions and vice versa. The incorporation of a priori knowledge imposed constraints on the classifiers and reduced the occurrence of such misjudgments by enabling the model to learn from previous states. It enabled the model to gain a better understanding of movement rules, leading to improved performance.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Confusion matrix for LREAL without a priori knowledge when facing transitions</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-10.tif"/>
</fig>
</sec>
<sec id="s5_3_5">
<label>5.3.5</label>
<title>Ablation Study</title>
<p>To investigate the effectiveness of each component in the LREAL network, we conducted ablation analyses on its second level (constructed by incorporating the AM into LSTM and connecting the two classifiers with a rule-modulating block) based on the ENABL3S dataset. We compared the results with the following model:
<list list-type="order">
<list-item>
<p>Raw LSTM: Two LSTM classifiers without an attention layer or a rule-modulating block between them.</p></list-item>
<list-item>
<p>Rule-Embedded (RE)-LSTM: Two LSTM classifiers with a rule-modulating block to deliver rule-embedded knowledge between them.</p></list-item>
<list-item>
<p>AT-LSTM: Two classifiers constructed using AT-LSTM, but without a rule-modulating block to input locomotion knowledge between them.</p></list-item>
</list></p>
<p>From the ablation study results (<xref ref-type="table" rid="table-9">Table 9</xref>), we observed that the addition of a rule-modulating block greatly improved the accuracy of recognizing transitions compared to the network without it. In summary, LREAL indeed outperformed all other variants.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Ablation study comparison for LREAL with other variant models</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Method</th>
<th colspan="2" align="center">Accuracy (%)</th>
</tr>
<tr>
<th>Steady states</th>
<th>Transitions</th>
</tr>
</thead>
<tbody>
<tr>
<td>Raw LSTM</td>
<td>94.16</td>
<td>92.22</td>
</tr>
<tr>
<td>RE-LSTM</td>
<td>96.88</td>
<td>98.37</td>
</tr>
<tr>
<td>AT-LSTM</td>
<td>99.32</td>
<td>95.63</td>
</tr>
<tr>
<td>LREAL</td>
<td><bold>99.32</bold></td>
<td><bold>98.96</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3_6">
<label>5.3.6</label>
<title>Parameter Analysis</title>
<p>To achieve the best performance for the hybrid network, the choice of parameters should be considered carefully and the parameters of each sub-classifier must be tuned and optimized. To assess the influence of the different hyperparameters (<xref ref-type="table" rid="table-3">Table 3</xref>) on the model and streamline the grid search process, a series of ablation experiments were conducted on the individual sub-classifiers and the hybrid network.</p>

<p>We conducted ablation experiments by considering the significant impact of the learning rate and optimizer on classifier performance, and recognizing that each optimizer has its own optimal learning rate. Specifically, we evaluated three optimizers, namely Adam, stochastic gradient descent (SGD) with momentum, and root mean squared propagation (RMSprop). Each optimizer was tested for a set of learning rates (0.01, 0.005, 0.001, 0.0005, and 0.0001). Analyzing the results (<xref ref-type="fig" rid="fig-11">Fig. 11</xref>), we observed that for the steady-state classifier and steady/transition classifier, the Adam optimizer with a learning rate of 0.001 outperformed the other combinations. However, for the transition-state classifier, the RMSprop optimizer with a learning rate of 0.0005 achieved the best performance compared to the other combinations (<xref ref-type="fig" rid="fig-11">Fig. 11</xref>).</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Impact of optimizers and learning rates on the sub-classifier performance</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-11.tif"/>
</fig>
<p>Finding a suitable batch size is crucial as it can effectively reduce memory usage and enhance the model generalize ability. We evaluated the impact of different batch sizes (32, 64, 128, 256, and 512) on the overall performance of the network (<xref ref-type="fig" rid="fig-12">Fig. 12</xref>). Our findings indicated that LREAL achieved the best overall effect for a batch size of 64.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Effect of batch size on LREAL performance</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_47903-fig-12.tif"/>
</fig>
</sec>
<sec id="s5_3_7">
<label>5.3.7</label>
<title>Limitations and Future Outlook</title>
<p><list list-type="simple">
<list-item><label>1)</label><p><italic>Rule-modulating block design:</italic> Although we designed a simple rule to incorporate into the network, its design of rule is based on the specific locomotor circuit in the ENABL3S dataset. When the situation becomes much more complex, e.g., considering more daily locomotion the rule will become diverse and the rule-embedded module may not be very effective. The method of combining a rule-based system with DL remains relatively simple. Therefore, a more comprehensive study of such combinations may improve the impact of rule-embedded networks and model interpretability.</p></list-item>
<list-item><label>2)</label><p><italic>First-level classifier accuracy:</italic> The performance of the first-level steady/transition classifier largely affected the overall network effectiveness as its classification errors propagate to the final classification. The steady/transition classifier achieved an average accuracy of 97.50% when facing transition states (<xref ref-type="fig" rid="fig-6">Fig. 6</xref>). This might be due to the manner of extracting the analysis window containing steady-state information before the transition. The accuracy may be improved by a further study of extracting transition-state analysis windows. Moreover, the unbalanced distribution of data may also contribute to this issue. A balanced dataset containing more transition state data could enhance the classifier performance.</p>
</list-item>
<list-item><label>3)</label><p><italic>Practical application requirements:</italic> To meet the needs of practical applications and reduce user discomfort, the latency of recognition should be &#x003C;300 ms [<xref ref-type="bibr" rid="ref-41">41</xref>]; our method can currently classify a sequence in 112 ms. However, when considering the 50 ms sliding window for sequence extraction and the time taken to preprocess the data, the total recognition time increases greatly (&#x003C;300 ms). As for a microcomputer on a real prosthesis, it might take a longer time to process data and recognize locomotor intent. Therefore, further research is required to reduce the system running time to meet the needs of practical applications.</p></list-item>
</list></p>
</sec>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>In this study, we propose a locomotion rule embedding LSTM network with an AM (LREAL) based on IMU and GONIO sensors to recognize human locomotor intent. The performance of the overall network and each classifier was validated on the public ENABL3S dataset and compared to other methods. The results show that the locomotion recognition of seven steady and twelve transition states achieved an average accuracy of 99.03% and 96.52%, respectively, on the ENABL3S dataset, i.e., comparable to or better than the other methods. We also analyzed the impact of the AM from two aspects and verified the effectiveness of adding the rule-modulating block by ablation studies. To summarize, this study demonstrates the excellent human locomotor intent recognition capabilities of AM and hence, the proposed LREAL network, which is important for the further study and development of assistive robotic.</p>
</sec>
</body>
<back>
<ack><p>We would like to express our gratitude to Haoming Da for his invaluable assistance with the experiments, and to Hui Yang for her significant contributions to language polishing and editing. Their expertise and insights have greatly enhanced the quality of this work.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This research was funded by the National Natural Science Foundation of China (Nos. 62072212, 62302218), the Development Project of Jilin Province of China (Nos. 20220508125RC, 20230201065GX, 20240101364JC), National Key R&#x0026;D Program (No. 2018YFC2001302), and the Jilin Provincial Key Laboratory of Big Data Intelligent Cognition (No. 20210504003GH).</p>
</sec>
<sec><title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Software, writing&#x2014;original draft preparation, writing&#x2014;review and editing: J. Shen; formal analysis, project administration and supervision: Y. Wang; investigation and validation: D. Zhang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>We using an open-source dataset to test our method. The ENABL3S dataset can be found in <ext-link ext-link-type="uri" xlink:href="https://figshare.com/articles/dataset/Benchmark_datasets_for_bilateral_lower_limb_neuromechanical_signals_from_wearable_sensors_during_unassisted_locomotion_in_able-bodied_individuals/5362627">https://figshare.com/articles/dataset/Benchmark_datasets_for_bilateral_lower_limb_neuromechanical_signals_from_wearable_sensors_during_unassisted_locomotion_in_able-bodied_individuals/5362627</ext-link> (accessed on 18 December 2021).</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Verghese</surname></string-name>, <string-name><given-names>A.</given-names> <surname>LeValley</surname></string-name>, <string-name><given-names>C. B.</given-names> <surname>Hall</surname></string-name>, <string-name><given-names>M. J.</given-names> <surname>Katz</surname></string-name>, <string-name><given-names>A. F.</given-names> <surname>Ambrose</surname></string-name> and <string-name><given-names>R. B.</given-names> <surname>Lipton</surname></string-name></person-group>, &#x201C;<article-title>Epidemiology of gait disorders in community-residing older adults</article-title>,&#x201D; <source>J. Am. Geriatr. Soc.</source>, vol. <volume>54</volume>, no. <issue>2</issue>, pp. <fpage>255</fpage>&#x2013;<lpage>261</lpage>, <year>2006</year>. doi: <pub-id pub-id-type="doi">10.1111/j.1532-5415.2005.00580.x</pub-id>; <pub-id pub-id-type="pmid">16460376</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>L. J.</given-names> <surname>Hargrove</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Dou</surname></string-name>, <string-name><given-names>D. R.</given-names> <surname>Rogers</surname></string-name> and <string-name><given-names>K. B.</given-names> <surname>Englehart</surname></string-name></person-group>, &#x201C;<article-title>Continuous locomotion-mode identification for prosthetic legs based on neuromuscular-mechanical fusion</article-title>,&#x201D; <source>IEEE Trans. BioMed. Eng.</source>, vol. <volume>58</volume>, no. <issue>10</issue>, pp. <fpage>2867</fpage>&#x2013;<lpage>2875</lpage>, <year>Oct. 2011</year>. doi: <pub-id pub-id-type="doi">10.1109/TBME.2011.2161671</pub-id>; <pub-id pub-id-type="pmid">21768042</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. X.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>E. M.</given-names> <surname>Gutierrez-Farewik</surname></string-name></person-group>, &#x201C;<article-title>Joint kinematics, kinetics and muscle synergy patterns during transitions between locomotion modes</article-title>,&#x201D; <source>IEEE Trans. Biomed. Eng.</source>, vol. <volume>70</volume>, no. <issue>3</issue>, pp. <fpage>1062</fpage>&#x2013;<lpage>1071</lpage>, <year>Mar. 2023</year>. doi: <pub-id pub-id-type="doi">10.1109/TBME.2022.3208381</pub-id>; <pub-id pub-id-type="pmid">36129869</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>L. R. B.</given-names> <surname>Schomaker</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Carloni</surname></string-name></person-group>, &#x201C;<chapter-title>IMU-based deep neural networks for locomotor intention prediction</chapter-title>,&#x201D; in <source>2020 IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS)</source>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>2020</year>, pp. <fpage>4134</fpage>&#x2013;<lpage>4139</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Mai</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Commuri</surname></string-name></person-group>, &#x201C;<chapter-title>Gait identification for an intelligent prosthetic foot</chapter-title>,&#x201D; in <source>2011 IEEE Int. Symp. Intell. Control</source>, <publisher-loc>Denver, CO, USA</publisher-loc>, <year>2011</year>, pp. <fpage>1341</fpage>&#x2013;<lpage>1346</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Human motion state recognition based on flexible, wearable capacitive pressure sensors</article-title>,&#x201D; <source>Micromachines</source>, vol. <volume>12</volume>, no. <issue>10</issue>, pp. <fpage>1219</fpage>, <year>Oct. 2021</year>. doi: <pub-id pub-id-type="doi">10.3390/mi12101219</pub-id>; <pub-id pub-id-type="pmid">34683270</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Camargo</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Flanagan</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Csomay-Shanklin</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Kanwar</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Young</surname></string-name></person-group>, &#x201C;<article-title>A machine learning strategy for locomotion classification and parameter estimation using fusion of wearable sensors</article-title>,&#x201D; <source>IEEE Trans. Biomed. Eng.</source>, vol. <volume>68</volume>, no. <issue>5</issue>, pp. <fpage>1569</fpage>&#x2013;<lpage>1578</lpage>, <year>May 2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TBME.2021.3065809</pub-id>; <pub-id pub-id-type="pmid">33710951</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Rouse</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Hargrove</surname></string-name></person-group>, &#x201C;<article-title>Fusion of bilateral lower-limb neuromechanical signals improves prediction of locomotor activities</article-title>,&#x201D; <source>Front. Rob. AI</source>, vol. <volume>5</volume>, pp. <fpage>78</fpage>&#x2013;<lpage>78</lpage>, <year>Jun. 2018</year>. doi: <pub-id pub-id-type="doi">10.3389/frobt.2018.00078</pub-id>; <pub-id pub-id-type="pmid">33500957</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Ding</surname></string-name></person-group>, &#x201C;<article-title>Real-time recognition of human lower-limb locomotion based on exponential coordinates of relative rotations</article-title>,&#x201D; <source>Sci. China Technol. Sci.</source>, vol. <volume>64</volume>, pp. <fpage>1423</fpage>&#x2013;<lpage>1435</lpage>, <year>Jul. 2021</year>. doi: <pub-id pub-id-type="doi">10.1007/s11431-020-1802-2</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>On-board training strategy for IMU-based real-time locomotion recognition of transtibial amputees with robotic prostheses</article-title>,&#x201D; <source>Front. Neurorobot.</source>, vol. <volume>14</volume>, no. <issue>47</issue>, pp. <fpage>608</fpage>, <year>Oct. 2020</year>. doi: <pub-id pub-id-type="doi">10.3389/fnbot.2020.00047</pub-id>; <pub-id pub-id-type="pmid">33192430</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Mai</surname></string-name>, and <string-name><given-names>Q.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Real-time on-board recognition of continuous locomotion modes for amputees with robotic transtibial prostheses</article-title>,&#x201D; <source>IEEE Trans. Neural Syst. Rehabil. Eng.</source>, vol. <volume>26</volume>, no. <issue>10</issue>, pp. <fpage>2015</fpage>&#x2013;<lpage>2025</lpage>, <year>Oct. 2018</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSRE.2018.2870152</pub-id>; <pub-id pub-id-type="pmid">30334741</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. H.</given-names> <surname>Moon</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kim</surname></string-name>, and <string-name><given-names>Y. D.</given-names> <surname>Hong</surname></string-name></person-group>, &#x201C;<article-title>Development of a single leg knee exoskeleton and sensing knee center of rotation change for intention detection</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>19</volume>, no. <issue>18</issue>, pp. <fpage>3960</fpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.3390/s19183960</pub-id>; <pub-id pub-id-type="pmid">31540298</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Narayan</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<chapter-title>A deep learning based end-to-end locomotion mode detection method for lower limb wearable robot control</chapter-title>,&#x201D; in <source>2020 IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS)</source>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>2020</year>, pp. <fpage>4091</fpage>&#x2013;<lpage>4097</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>P. C.</given-names> <surname>Dixon</surname></string-name>, <string-name><given-names>J. V.</given-names> <surname>Jacobs</surname></string-name>, <string-name><given-names>J. T.</given-names> <surname>Dennerlein</surname></string-name>, and <string-name><given-names>J. M.</given-names> <surname>Schiffman</surname></string-name></person-group>, &#x201C;<article-title>Machine learning algorithms based on signals from a single wearable inertial sensor can detect surface- and age-related differences in walking</article-title>,&#x201D; <source>J. Biomech.</source>, vol. <volume>71</volume>, pp. <fpage>37</fpage>&#x2013;<lpage>42</lpage>, <year>Apr. 2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.jbiomech.2018.01.005</pub-id>; <pub-id pub-id-type="pmid">29452755</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Bruinsma</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Carloni</surname></string-name></person-group>, &#x201C;<article-title>IMU-based deep neural networks: Prediction of locomotor and transition intentions of an osseointegrated transfemoral amputee</article-title>,&#x201D; <source>IEEE Trans. Neural Syst. Rehabil. Eng.</source>, vol. <volume>29</volume>, pp. <fpage>1079</fpage>&#x2013;<lpage>1088</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSRE.2021.3086843</pub-id>; <pub-id pub-id-type="pmid">34097612</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Arulkumaran</surname></string-name>, <string-name><given-names>M. P.</given-names> <surname>Deisenroth</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Brundage</surname></string-name>, and <string-name><given-names>A. A.</given-names> <surname>Bharath</surname></string-name></person-group>, &#x201C;<article-title>Deep reinforcement learning: A brief survey</article-title>,&#x201D; <source>IEEE Signal Process. Mag.</source>, vol. <volume>34</volume>, no. <issue>6</issue>, pp. <fpage>26</fpage>&#x2013;<lpage>38</lpage>, <year>Nov. 2017</year>. doi: <pub-id pub-id-type="doi">10.1109/MSP.2017.2743240</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Fran&#x00E7;ois-Lavet</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Henderson</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Islam</surname></string-name>, <string-name><given-names>M. G.</given-names> <surname>Bellemare</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Pineau</surname></string-name></person-group>, &#x201C;<article-title>An introduction to deep reinforcement learning</article-title>,&#x201D; <source>Found Trends Mach. Learn.</source>, vol. <volume>11</volume>, no. <issue>3&#x2013;4</issue>, pp. <fpage>219</fpage>&#x2013;<lpage>354</lpage>, <year>Dec. 2018</year>. doi: <pub-id pub-id-type="doi">10.1561/2200000071</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Ahn</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Lee</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Sentis</surname></string-name></person-group>, &#x201C;<article-title>Data-efficient and safe learning for humanoid locomotion aided by a dynamic balancing model</article-title>,&#x201D; <source>IEEE Robot. Autom. Lett.</source>, vol. <volume>5</volume>, no. <issue>3</issue>, pp. <fpage>4376</fpage>&#x2013;<lpage>4383</lpage>, <year>Jul. 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/LRA.2020.2990743</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhao</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<article-title>End-to-end high-level control of lower-limb exoskeleton for human performance augmentation based on deep reinforcement learning</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>11</volume>, pp. <fpage>102340</fpage>&#x2013;<lpage>102351</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3317183</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liang</surname></string-name>, <string-name><given-names>K. I. K.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>L. T.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Jin</surname></string-name></person-group>, &#x201C;<article-title>Deep-learning-enhanced human activity recognition for internet of healthcare things</article-title>,&#x201D; <source>IEEE Internet Things J.</source>, vol. <volume>7</volume>, no. <issue>7</issue>, pp. <fpage>6429</fpage>&#x2013;<lpage>6438</lpage>, <year>July 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/JIOT.2020.2985082</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Huang</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>sEMG-based continuous estimation of grasp movements by long-short term memory network</article-title>,&#x201D; <source>Biomed. Signal Process. Control</source>, vol. <volume>59</volume>, pp. <fpage>101774</fpage>, <year>May 2020</year>. doi: <pub-id pub-id-type="doi">10.1016/j.bspc.2019.101774</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Guo</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Attention mechanisms in computer vision: A survey</article-title>,&#x201D; <source>Comput. Vis. Media</source>, vol. <volume>8</volume>, no. <issue>3</issue>, pp. <fpage>331</fpage>&#x2013;<lpage>368</lpage>, <year>Sep. 2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s41095-022-0271-y</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Meng</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Ai</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Xie</surname></string-name></person-group>, &#x201C;<chapter-title>An attention-based CNN-LSTM model with limb synergy for joint angles prediction</chapter-title>,&#x201D; in <source>2021 IEEE/ASME Int. Conf. Adv. Intell. Mechatron. (AIM)</source>, <publisher-loc>Delft, Netherlands</publisher-loc>, <year>2021</year>, pp. <fpage>747</fpage>&#x2013;<lpage>752</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. A.</given-names> <surname>Varol</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Sup</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Goldfarb</surname></string-name></person-group>, &#x201C;<article-title>Multiclass real-time intent recognition of a powered lower limb prosthesis</article-title>,&#x201D; <source>IEEE Trans. Biomed. Eng.</source>, vol. <volume>57</volume>, no. <issue>3</issue>, pp. <fpage>542</fpage>&#x2013;<lpage>551</lpage>, <year>Mar. 2010</year>. doi: <pub-id pub-id-type="doi">10.1109/TBME.2009.2034734</pub-id>; <pub-id pub-id-type="pmid">19846361</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<chapter-title>ReNN: Rule-embedded neural networks</chapter-title>,&#x201D; in <source>2018 24th Int. Conf. Pattern Recogn. (ICPR)</source>, <publisher-loc>Beijing, China</publisher-loc>, <year>2018</year>, pp. <fpage>824</fpage>&#x2013;<lpage>829</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Rouse</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Hargrove</surname></string-name></person-group>, &#x201C;<article-title>Benchmark datasets for bilateral lower-limb neuromechanical signals from wearable sensors during unassisted locomotion in able-bodied individuals</article-title>,&#x201D; <source>Front. Rob. AI</source>, vol. <volume>5</volume>, pp. <fpage>14</fpage>, <year>Feb. 2018</year>. doi: <pub-id pub-id-type="doi">10.3389/frobt.2018.00014</pub-id>; <pub-id pub-id-type="pmid">33500901</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>U. H.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Bi</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Patel</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Fouhey</surname></string-name>, and <string-name><given-names>E.</given-names> <surname>Rouse</surname></string-name></person-group>, &#x201C;<article-title>Image transformation and CNNs: A strategy for encoding human locomotor intent for autonomous wearable robots</article-title>,&#x201D; <source>IEEE Robot. Autom. Lett.</source>, vol. <volume>5</volume>, no. <issue>4</issue>, pp. <fpage>5440</fpage>&#x2013;<lpage>5447</lpage>, <year>Oct. 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/LRA.2020.3007455</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Wearable sensor-based human activity recognition using hybrid deep learning techniques</article-title>,&#x201D; <source>Secur. Commun. Netw.</source>, vol. <volume>2020</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>Jul. 2020</year>. doi: <pub-id pub-id-type="doi">10.1155/2020/2132138</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Sun</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Pan</surname></string-name></person-group>, &#x201C;<chapter-title>LSTM with uniqueness attention for human activity recognition</chapter-title>,&#x201D; in <source>Artificial Neural Networks and Machine Learning&#x2013;ICANN 2019</source>, <publisher-loc>Munich, Germany</publisher-loc>, <year>2019</year>, pp. <fpage>498</fpage>&#x2013;<lpage>509</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bengio</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Hardt</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Recht</surname></string-name>, and <string-name><given-names>O.</given-names> <surname>Vinyals</surname></string-name></person-group>, &#x201C;<article-title>Understanding deep learning (still) requires rethinking generalization</article-title>,&#x201D; <source>Commun. ACM.</source>, vol. <volume>64</volume>, pp. <fpage>107</fpage>&#x2013;<lpage>115</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1145/3446776</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Winter</surname></string-name></person-group>, &#x201C;<article-title><italic>xia2</italic>: An expert system for macromolecular crystallography data reduction</article-title>,&#x201D; <source>J. Appl. Crystallogr.</source>, vol. <volume>43</volume>, pp. <fpage>186</fpage>&#x2013;<lpage>190</lpage>, <year>Feb. 2010</year>. doi: <pub-id pub-id-type="doi">10.1107/S0021889809045701</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. J.</given-names> <surname>Young</surname></string-name> and <string-name><given-names>L. J.</given-names> <surname>Hargrove</surname></string-name></person-group>, &#x201C;<article-title>A classification method for user-independent intent recognition for transfemoral amputees using powered lower limb prostheses</article-title>,&#x201D; <source>IEEE Trans. Neural Syst. Rehabil. Eng.</source>, vol. <volume>24</volume>, no. <issue>2</issue>, pp. <fpage>217</fpage>&#x2013;<lpage>225</lpage>, <year>Feb. 2016</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSRE.2015.2412461</pub-id>; <pub-id pub-id-type="pmid">25794392</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>D. P.</given-names> <surname>Kingma</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Ba</surname></string-name></person-group>, &#x201C;<chapter-title>Adam: A method for stochastic optimization</chapter-title>,&#x201D; in <source>3rd Int. Conf. Learn. Represent. (ICLR)</source>, <publisher-loc>San Diego, USA</publisher-loc>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Figueiredo</surname></string-name>, <string-name><given-names>S. P.</given-names> <surname>Carvalho</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Goncalve</surname></string-name>, <string-name><given-names>J. C.</given-names> <surname>Moreno</surname></string-name>, and <string-name><given-names>C. P.</given-names> <surname>Santos</surname></string-name></person-group>, &#x201C;<article-title>Daily locomotion recognition and prediction: A kinematic data-based machine learning approach</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>33250</fpage>&#x2013;<lpage>33262</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2020.2971552</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>U.</given-names> <surname>Martinez-Hernandez</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Meng</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Rubio-Solis</surname></string-name></person-group>, &#x201C;<chapter-title>Towards a context-based Bayesian recognition of transitions in locomotion activities</chapter-title>,&#x201D; in <source>2020 29th IEEE Int. Conf. Robot Hum. Interact. Commun. (RO-MAN)</source>, <publisher-loc>Naples, Italy</publisher-loc>, <year>2020</year>, pp. <fpage>677</fpage>&#x2013;<lpage>682</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Jabban</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Sui</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Motion intention prediction and joint trajectories generation toward lower limb prostheses using EMG and IMU signals</article-title>,&#x201D; <source>IEEE Sens. J.</source>, vol. <volume>22</volume>, no. <issue>11</issue>, pp. <fpage>10719</fpage>&#x2013;<lpage>10729</lpage>, <year>Jun. 2022</year>. doi: <pub-id pub-id-type="doi">10.1109/JSEN.2022.3167686</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Marcos Mazon</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Groefsema</surname></string-name>, <string-name><given-names>L. R.</given-names> <surname>Schomaker</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Carloni</surname></string-name></person-group>, &#x201C;<article-title>IMU-based classification of locomotion modes, transitions, and gait phases with convolutional recurrent neural networks</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>22</volume>, no. <issue>22</issue>, pp. <fpage>8871</fpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.3390/s22228871</pub-id>; <pub-id pub-id-type="pmid">36433469</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>C. W.</given-names> <surname>de Silva</surname> </string-name>, and <string-name><given-names>C.</given-names> <surname>Fu</surname></string-name></person-group>, &#x201C;<article-title>Unsupervised cross-subject adaptation for predicting human locomotion intent</article-title>,&#x201D; <source>IEEE Trans. Neural Syst. Rehabil. Eng.</source>, vol. <volume>28</volume>, no. <issue>3</issue>, pp. <fpage>646</fpage>&#x2013;<lpage>657</lpage>, <year>Mar. 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSRE.2020.2966749</pub-id>; <pub-id pub-id-type="pmid">31944980</pub-id></mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. A.</given-names> <surname>Mohamed</surname></string-name> and <string-name><given-names>U.</given-names> <surname>Martinez-Hernandez</surname></string-name></person-group>, &#x201C;<article-title>A light-weight artificial neural network for recognition of activities of daily living</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>23</volume>, no. <issue>13</issue>, pp. <fpage>5854</fpage>, <year>Jul. 2023</year>. doi: <pub-id pub-id-type="doi">10.3390/s23135854</pub-id>; <pub-id pub-id-type="pmid">37447703</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B. Y.</given-names> <surname>Su</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A CNN-based method for intent recognition using inertial measurement units and intelligent lower limb prosthesis</article-title>,&#x201D; <source>IEEE Trans. Neural Syst. Rehabil. Eng.</source>, vol. <volume>27</volume>, no. <issue>5</issue>, pp. <fpage>1032</fpage>&#x2013;<lpage>1042</lpage>, <year>May 2019</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSRE.2019.2909585</pub-id>; <pub-id pub-id-type="pmid">30969928</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Englehart</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Hudgins</surname></string-name></person-group>, &#x201C;<article-title>A robust, real-time control scheme for multifunction myoelectric control</article-title>,&#x201D; <source>IEEE Trans. Biomed. Eng.</source>, vol. <volume>50</volume>, no. <issue>7</issue>, pp. <fpage>848</fpage>&#x2013;<lpage>854</lpage>, <year>July 2003</year>. doi: <pub-id pub-id-type="doi">10.1109/TBME.2003.813539</pub-id>; <pub-id pub-id-type="pmid">12848352</pub-id></mixed-citation></ref>
</ref-list>
</back></article>