<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">25106</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2022.025106</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Development of Voice Control Algorithm for Robotic Wheelchair Using NIN and LSTM Models</article-title>
<alt-title alt-title-type="left-running-head">Development of Voice Control Algorithm for Robotic Wheelchair Using MIN and LSTM Models</alt-title>
<alt-title alt-title-type="right-running-head">Development of Voice Control Algorithm for Robotic Wheelchair Using MIN and LSTM Models</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Bakouri</surname><given-names>Mohsen</given-names></name><xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="aff" rid="aff-2">2</xref><email>m.bakouri@mu.edu.sa</email>
</contrib>
<aff id="aff-1"><label>1</label><institution>Department of Medical Equipment Technology, College of Applied Medical Science, Majmaah University</institution>, <addr-line>Majmaah City, 11952</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Physics, College of Arts, Fezzan University</institution>, <addr-line>Traghen, 71340</addr-line>, <country>Libya</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Mohsen Bakouri. Email: <email>m.bakouri@mu.edu.sa</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-06-14"><day>14</day>
<month>06</month>
<year>2022</year></pub-date>
<volume>73</volume>
<issue>2</issue>
<fpage>2441</fpage>
<lpage>2456</lpage>
<history>
<date date-type="received"><day>12</day><month>11</month><year>2021</year></date>
<date date-type="accepted"><day>09</day><month>2</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Bakouri</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Bakouri</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_25106.pdf"></self-uri>
<abstract>
<p>In this work, we developed and implemented a voice control algorithm to steer smart robotic wheelchairs (SRW) using the neural network technique. This technique used a network in network (NIN) and long short-term memory (LSTM) structure integrated with a built-in voice recognition algorithm. An Android Smartphone application was designed and configured with the proposed method. A Wi-Fi hotspot was used to connect the software and hardware components of the system in an offline mode. To operate and guide SRW, the design technique proposed employing five voice commands (yes, no, left, right, no, and stop) via the Raspberry Pi and DC motors. Ten native Arabic speakers trained and validated an English speech corpus to determine the method&#x2019;s overall effectiveness. The design method of SRW was evaluated in both indoor and outdoor environments in order to determine its time response and performance. The results showed that the accuracy rate for the system reached 98.2&#x0025; for the five-voice commends in classifying voices accurately. Another interesting finding from the real-time test was that the root-mean-square deviation (RMSD) for indoor/outdoor maneuvering nodes was 2.2&#x002A;10<sup>&#x2013;5</sup> (for latitude), while that for longitude coordinates was a whopping 2.4&#x002A;10<sup>&#x2013;5</sup> (for latitude).</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Network in network</kwd>
<kwd>long short-term memory</kwd>
<kwd>voice recognition</kwd>
<kwd>wheelchair</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Disabled people in public places have a complex time maneuvering wheelchairs. Those people also depend on others to assist them in moving their wheelchairs [<xref ref-type="bibr" rid="ref-1">1</xref>]. According to [<xref ref-type="bibr" rid="ref-2">2</xref>], People with limited mobility make up 40&#x0025; of those unable to steer and maneuver wheelchairs adequately, compared to the 9&#x0025;&#x2013;10&#x0025; who have been taught to operate power wheelchairs. Furthermore, clinical studies indicate that nearly half of 40&#x0025; of disabled with impaired mobility cannot control an electric wheelchair. More than 10&#x0025; of those disabled who use electric wheelchairs have had an accident within the first four months [<xref ref-type="bibr" rid="ref-3">3</xref>]. Accordingly, to provide a better quality of life for wheelchair users, it has been developed with various technologies fitted with a navigation and sensor system that works automatically [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>]. These wheelchairs are known as Smart Robotic Wheelchairs (SRW) due to the introduction of more choices for controlling the chair, improved safety, and comfort over conventional wheelchairs [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Generally, to perform autonomous activities, the SRW must be capable of navigating safely, avoiding obstacles, and passing through doorways or any other confined space [<xref ref-type="bibr" rid="ref-10">10</xref>]. In SRW, which is controlled via a joystick intelligent control system unit, the operation of this system has proven to be the most significant development [<xref ref-type="bibr" rid="ref-11">11</xref>]. However, people with disabilities in their upper extremities will have difficulty using the joystick smoothly. As a result, situations requiring quick response could result in tragic incidents [<xref ref-type="bibr" rid="ref-12">12</xref>]. Therefore, several researchers have started to develop SRW based on human physiological signals. For example, the human-computer interface (HCI) operates a wheelchair by the use of physiological signals such as the electrooculogram (EOG), the electromyogram (EMG), and the electroencephalogram (EEG) [<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>]. On the other hand, brain-computer interfaces (BCIs) have advantages for translating brain signals into action for wheelchair control [<xref ref-type="bibr" rid="ref-16">16</xref>]. The hybrid BCI (hBCI) approach, which integrates EEG and EOG, increased wheelchair accuracy and speed. However, the technology of EEG-BCI has several limitations in terms of low resolution and signal-to-noise ratio (SNR). Additionally, hBCI encounters difficulties simultaneously controlling speed and direction [<xref ref-type="bibr" rid="ref-17">17</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>In general, different researchers have significantly enhanced the development of SRWs with autonomous functions via voice recognition technology. The strategy described in [<xref ref-type="bibr" rid="ref-20">20</xref>] illustrated the result of an intelligent wheelchair system using a voice recognition technique in conjunction with a GPS tracking model. By using a Wi-Fi module, voice commands were transformed into hexadecimal number data and used to drive the wheelchair in three different speed phases. Additionally, the system utilized an infrared (IR) sensor to identify barriers and a mobile application to determine the patient&#x2019;s location. A similar work conducted by Raiyan et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] utilizes an Arduino and Easy VR3 with a voice recognition module to drive an autonomous wheelchair system. This study demonstrated that the implemented system robustly guided the wheelchair with less complex data processing and without wearable sensors. A different novel study employs an adaptive neuro-fuzzy to steer a motorized wheelchair [<xref ref-type="bibr" rid="ref-22">22</xref>]. The study was created and executed using real-time control signals supplied by voice instructions via a classification unit. This architecture&#x2019;s proposed system for tracking the wheelchair uses a wireless sensor network [<xref ref-type="bibr" rid="ref-22">22</xref>]. Despite the highly advanced methodologies presented by researchers in this field, the high cost, and precision required for distinguishing, categorizing, and identifying the patient&#x2019;s voice continue to be the primary obstacles.</p>
<p>Recently, numerous researchers have employed the convolutional neural network (CNN) technology to overcome the inaccuracy of identifying and classifying patients&#x2019; speech [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]. This technology converts speech commands into spectrogram visuals, then fed into CNN. This technique has been shown to improve the accuracy of speech recognition. In this context, Sharifuddin et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] introduced an inelegant design using CNN to steer SRW based on four voice commands. The method used data collected from the google website and applied Mel-frequency cepstral coefficient (MFCC) to extract voice commands. Authors claim that the results of the vice commend classification using CNN have an accuracy of 95.30&#x0025;. Similarly, Sutikno et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] developed a voice-controlled wheelchair-using CNN and long short-term memory (LSTM) based on five commends. The developed method used data obtained from recording several subjects using sound recorder pro and sox sound exchange. This method demonstrated that the vice commands classification using CNN and LSTM has accuracy above 97.80&#x0025;. Although many of the research results that have been conducted are significantly high, computers are still used in these methods to perform complex operations on CNN.</p>
<p>However, due to the extensive calculations required to attain high accuracy, employing CNN in smartphones is still developing [<xref ref-type="bibr" rid="ref-26">26</xref>]. This article proposes to design a voice control method for robotic wheelchairs by using CNNs and LSTM models [<xref ref-type="bibr" rid="ref-27">27</xref>]. The system used a smartphone to build an interactive user interface that can be controlled easily by delivering a voice command to the system&#x2019;s motherboard via the mobile application. The objective of this study was accomplished by developing and implementing a mobile application, a voice recognition model, a CNN model, and an LSTM model. Additionally, all safety considerations were taken into account while driving and navigating in both indoor and outdoor environments. The results indicated that the built system was remarkably resilient in terms of response time and correct execution of all orders without delay.</p>
</sec>
<sec id="s2"><label>2</label><title>Materials and Methods</title>
<sec id="s2_1"><label>2.1</label><title>Architecture of the System</title>
<p>The implementation of the proposed architecture system is divided into two stages, as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. In the first stage, an Android mobile application was developed using the Flutter programming language [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>]. Six steps were used to create and program the application (named voice control), as shown in <xref ref-type="fig" rid="fig-2">Figs. 2a</xref> and <xref ref-type="fig" rid="fig-2">2b</xref>. The voice control application appears in the application list when accessed. After granting the application access to the microphone, it attempts to recognize the words and highlights them in the interface recognition, as depicted in <xref ref-type="fig" rid="fig-2">Fig. 2c</xref>. The second stage consists of assembling hardware devices, including mechanical parts and a control unit, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The mechanical parts are composed of a standard wheelchair, two motor pairs (3.13.6 LST10 24v DC 120&#x2005;rpm), and an NP7&#x2013;12 12v 7ah lead acid battery. At the same time, the control unit includes Raspberry pi4 (GPU: Broadcom Video Core VI, Networking: 2.4&#x2005;GHz, RAM: 4 GB LPDDR4 SDRAM, Bluetooth 5.0, microSD), and Relay Module (5&#x2005;V 4-channel relay interface board).</p>
<fig id="fig-1">
<label>Figure 1</label><caption><title>Overall system architecture</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_25106-fig-1.png"/>
</fig>
<fig id="fig-2">
<label>Figure 2</label><caption><title>Android app interface created for controlling powered wheelchair and its voice command prediction ratio: (a) Steps of creating the application (a) Main of voice control app, (b) Mode screen (Right, Left, No, Yes, and stop)</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_25106-fig-2.png"/>
</fig>
<fig id="fig-3">
<label>Figure 3</label><caption><title>Mechanical assembly of wheelchair</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_25106-fig-3.png"/>
</fig>
</sec>
<sec id="s2_2"><label>2.2</label><title>Development of Voice Recognition Model</title>
<p>Feature extraction is used to produce a frequency map for each audio file, which displays how the signal evolves over time. As a result, speech analysis systems used the Mel-frequency cepstral coefficients (MFCC) coefficients to extract this information [<xref ref-type="bibr" rid="ref-30">30</xref>]. An important part of character extraction is preventing numerical instability by putting it through a finite impulse response filter (FIR), which is a single-coefficient, digital filter as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BE;</mml:mi><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are the output filter and the original voice signal respectively. The number of sampling is denoted by <italic>n</italic>, and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>&#x03BE;</mml:mi></mml:math></inline-formula> given as <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>&#x03BE;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>.</p>
<p>It is necessary to use framing and windowing [<inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>w</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>] in order to maintain the samples within frames and reduce signal discontinuities as:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>w</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is a constant, and <italic>N</italic> represent the number of frames.</p>
<p>In this method, the spectral analysis is achieved using fast Fourier transform (FFT) to calculate the magnitude spectrum for each frame as:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mfrac><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>j</mml:mi><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mi>N</mml:mi></mml:mfrac></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></disp-formula></p>
<p>The spectrum is subsequently processed according to MFCC using a bank of filters; where Mel-filter-bank can be written as:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>It is possible to write the boundary points (<inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>f</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>) by taking the lowest (<inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), and highest (<inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) values on the filter-bank in terms of hertz and frequency as:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>B</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>m</mml:mi><mml:mfrac><mml:mrow><mml:mi>B</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>B</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>M</italic>, and <italic>N</italic> are the number of filters and the size of the FFT respectivily. The term <italic>B</italic> is representing the Mel-scale which calculated by:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>B</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>1125</mml:mn><mml:mi>ln</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mfrac><mml:mi>f</mml:mi><mml:mrow><mml:mn>700</mml:mn></mml:mrow></mml:mfrac><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>In this work, we used an approximate homomorphic transform to remove noise and spectral estimation errors as:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>S</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mi>ln</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>M</mml:mi></mml:math></disp-formula></p>
<p>In the final step of MFCC processing, Cosine Transformer (DCT) are employed to provide high decorrelation properties as:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:msqrt><mml:mfrac><mml:mn>2</mml:mn><mml:mi>M</mml:mi></mml:mfrac></mml:msqrt><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>n</mml:mi><mml:mi>&#x03C0;</mml:mi></mml:mrow><mml:mi>M</mml:mi></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>M</mml:mi><mml:mspace width="thinmathspace" /></mml:math></disp-formula></p>
<p>The first and second derivatives of <xref ref-type="disp-formula" rid="eqn-8">(8)</xref> are used to obtain the feature map:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msup><mml:mi>P</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msup><mml:mi>P</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>The database was thus developed and used by CNN, and it applies to all recordings that have been recorded.</p>
</sec>
<sec id="s2_3"><label>2.3</label><title>Development of CNN Model</title>
<p>This study used the NIN structure as the core architecture for developing mobile applications [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. NIN is a CNN technique that does not employ fully connected (FC) layers. Instead, NIN uses global pooling rather than fixed-size pools to take images of any size as inputs. This technique is helpful for mobile applications because it allows users to fine-tune the speed-accuracy trade-off without compromising network weights.</p>
<p>In order to develop CNN, we employ a multi-threading technique. The smartphone used in this technique has four CPU cores, making it simple to divide a kernel matrix into four sub-matrices and divide a row into four sub-matrices. To obtain the output feature maps, it is necessary to conduct four generalized matrix multiplication (GEMM) operations simultaneously. The cascaded cross channel parametric pooling (CCCPP) technique was also utilized to compensate for the loss of the FC layers. Because of this, our CNN model comprises input and output, twelve convolution layers, and two succeeding layers.</p>
</sec>
<sec id="s2_4"><label>2.4</label><title>Development of LSTM Model</title>
<p>We adopted LSTM model as a vanilla structure [<xref ref-type="bibr" rid="ref-32">32</xref>]. The architecture of this model consists of a set of recurrently connected sub-networks, known as memory blocks. In specific, the model is composed of a cell, an input gate, an output gate, and a forget gate. The mechanism of the LSTM model start works by identifying and eliminating of last outputs data (<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) and current inputs data (<inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) using the sigmoid function (<inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula>). This step is achieved by forgetting gate (<inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) function, which is given by:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are weight matrices and bias weight vector respectively. In the second step, the model will store and update the new input data in the cell state (<inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) using the sigmoid layer and tanh layer. The sigmoid layer decides to update or neglect the new data using (1 or 0), while the tanh layer gives weighs to the passing data using (1 or &#x2212;1). Then, the old memory data (<inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) added to the new memory of the cell state as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In the final step, the output value (<inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) is calculated based on the sigmoid gate (<inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and the new values created by <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and tanh layer as:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates the diagram of the proposed CNN with LSTM neural network.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>The CNN with LSTM neural network</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_25106-fig-4.png"/></fig>
<p>The flowchart in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> depicts the signal flow for controlling wheelchair system.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Flowchart of the system</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_25106-fig-5.png"/></fig>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Experimental Procedure</title>
<p>In the Health and Basic Sciences Research Center at Majmaah University, the English speech corpus of isolated words was used to evaluate the proposed system. Ten native Arabic speakers were selected to pronounce five words with a total of 2,000 utterances. The data was recorded using a 20 kHz sample rate and 16-bit resolution. Then, using the reinforcement method, this data set was supplemented with additional audio cues. The supplementary dataset contains 2,000 speech altered in pitch, velocity, dynamic range, noise, and forward and backward time shift. The new data set (original and supplemented) contains 4000 utterances and is divided into two parts: a training set (training and validation) containing about 80&#x0025; of the samples (3200), and a test set containing the remaining 20&#x0025; of the sample (800).</p>
<p>To quantify the predictive accuracy and quality for the proposed system, we compute the F-score as:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>F</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mi mathvariant="normal">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">R</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where P and R indicate for precision and recall, respectively, and are defined as follows::
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>here, <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are false positive, false negative, and true positive respectively.</p>
<p>During the classification, the percentage difference (&#x0025;<italic>d</italic>) equation was employed to measure the accuracy of each voice command prediction as:
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x2217;</mml:mo><mml:mn>100</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></disp-formula>where <italic>V<sub>1</sub> and V<sub>2</sub></italic> represent the first and second observations during the comparison process. Indoor/outdoor navigational performance is also assessed in real time using this methodology. With vocal commands, the wheelchair navigated around and inside a mosque at 24.893374, 46.614728 coordinates.</p>
</sec>
<sec id="s4"><label>4</label><title>Results</title>
<p><xref ref-type="table" rid="table-1 table-2 table-3">Tabs. 1&#x2013;3</xref> represent the steps of voice recognition model development. <xref ref-type="table" rid="table-1">Tab. 1</xref> shows the five audio wave shapes with training time.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>The five audio wave shapes</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center">
<inline-graphic xlink:href="CMC_25106-inline-1"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>The five singles of long-term spectrum</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_25106-inline-2"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The five singles of time-frequency spectrogram</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_25106-inline-3"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The next step is to convert the audio file waves into its frequency domain by using Fourier analysis as shown in <xref ref-type="table" rid="table-2">Tab. 2</xref>.</p>
<p>Then, these frequency domain waves were converted into spectrograms (See <xref ref-type="table" rid="table-3">Tab. 3</xref>) and used as input in NIN model then to LSTM model respectively.</p>
<p><xref ref-type="table" rid="table-4">Tab. 4</xref> illustrated the screen shoot of mobile application (voice command prediction ratio). Additionally, the application displays the user&#x2019;s expected word weight. It is usually a single-voice order with greater weight than other words, indicating that no incorrect classification judgment can be made during the classification process.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>The five words of screen shoot for mobile app</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_25106-inline-4"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The confusion matrix was generated using the initial results, as shown in <xref ref-type="table" rid="table-5">Tab. 5</xref>. The average accuracy was around 82.6&#x0025; of the accurate forecast for five-voice commands. We used the phrases true positives, true negatives, false positives, and false negatives to describe the classification activity. The computations for the voice-command prediction ratio, accuracy, and precision are shown in <xref ref-type="table" rid="table-6">Tabs. 6</xref> and <xref ref-type="table" rid="table-7">7</xref>. In terms of calculating the percentage difference between two commands when comparing them, the example compares &#x201C;STOP&#x201D; to other commands. This suggests a minor risk of selecting an inaccurate classification option. On the other hand, the difference between accurate and erroneous predictions is quite significant, indicating a negligible risk of making incorrect predictions. The difference exceeded 187 percent, as shown in <xref ref-type="table" rid="table-7">Tab. 7</xref>.</p>
<table-wrap id="table-5"><label>Table 5</label><caption><title>Normalized of confusion matrix</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="center" colspan="7">Actual voice command</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="6">Prediction ratio &#x0025;</td>
<td align="left">Class</td>
<td align="left">Yes</td>
<td align="left">No</td>
<td align="left">Left</td>
<td align="left">Right</td>
<td align="left">Stop</td>
</tr>
<tr>
<td align="left">Yes</td>
<td align="left">87&#x0025;</td>
<td align="left">3&#x0025;</td>
<td align="left">3&#x0025;</td>
<td align="left">4&#x0025;</td>
<td align="left">3&#x0025;</td>
</tr>
<tr>
<td align="left">No</td>
<td align="left">2&#x0025;</td>
<td align="left">90&#x0025;</td>
<td align="left">3&#x0025;</td>
<td align="left">2&#x0025;</td>
<td align="left">3&#x0025;</td>
</tr>
<tr>
<td align="left">Left</td>
<td align="left">4&#x0025;</td>
<td align="left">1&#x0025;</td>
<td align="left">92&#x0025;</td>
<td align="left">2&#x0025;</td>
<td align="left">1&#x0025;</td>
</tr>
<tr>
<td align="left">Right</td>
<td align="left">5&#x0025;</td>
<td align="left">4&#x0025;</td>
<td align="left">2&#x0025;</td>
<td align="left">87&#x0025;</td>
<td align="left">2&#x0025;</td>
</tr>
<tr>
<td align="left">Stop</td>
<td align="left">3&#x0025;</td>
<td align="left">5&#x0025;</td>
<td align="left">3&#x0025;</td>
<td align="left">3&#x0025;</td>
<td align="left">86&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-6"><label>Table 6</label><caption><title>Accuracy, precision, recall, and F-score for voice commands</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Class</th>
<th align="left">Accuracy</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">F-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Yes</td>
<td align="left">95&#x0025;</td>
<td align="left">0.73</td>
<td align="left">0.74</td>
<td align="left">0.735</td>
</tr>
<tr>
<td align="left">No</td>
<td align="left">96.3&#x0025;</td>
<td align="left">0.75</td>
<td align="left">0.75</td>
<td align="left">0.75</td>
</tr>
<tr>
<td align="left">Left</td>
<td align="left">98.2&#x0025;</td>
<td align="left">0.77</td>
<td align="left">0.75</td>
<td align="left">0.76</td>
</tr>
<tr>
<td align="left">Right</td>
<td align="left">94.8&#x0025;</td>
<td align="left">0.75</td>
<td align="left">0.73</td>
<td align="left">0.74</td>
</tr>
<tr>
<td align="left">Stop</td>
<td align="left">93.5&#x0025;</td>
<td align="left">0.71</td>
<td align="left">0.69</td>
<td align="left">0.70</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-7"><label>Table 7</label><caption><title>Calculation of percentage difference for stop command</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Voice command</th>
<th align="left">Yes</th>
<th align="left">No</th>
<th align="left">Left</th>
<th align="left">Right</th>
<th align="left">Stop</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Prediction ratio</td>
<td align="left">3&#x0025;</td>
<td align="left">5&#x0025;</td>
<td align="left">3&#x0025;</td>
<td align="left">3&#x0025;</td>
<td align="left">86&#x0025;</td>
</tr>
<tr>
<td align="left">Percentage difference</td>
<td align="left">187&#x0025;</td>
<td align="left">178&#x0025;</td>
<td align="left">187&#x0025;</td>
<td align="left">187&#x0025;</td>
<td align="left">----</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>An evaluation of indoor/outdoor navigation for the wheelchair was obtained to test the real-time performance of the designed system within the public area. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> depicts the intended route navigation in comparison to the actual path. <xref ref-type="table" rid="table-8">Tab. 8</xref> shows the coordinate nodes of the intended and actual pathways while traversing. The root means square deviation (RMSD) was used to represent the difference between the planned and actual nodes in this experiment. Figures show that the RMSD for latitude and longitude coordinates are 2.2 &#x002A; 10<sup>&#x2013;5</sup> and 2.4 &#x002A; 10<sup>&#x2013;5</sup>, respectively.</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>Navigation planned route <italic>vs.</italic> actual route</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_25106-fig-6.png"/></fig>
<table-wrap id="table-8"><label>Table 8</label><caption><title>Coordinates of outdoor navigation</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Planned longitude</th>
<th align="left">Planned latitude</th>
<th align="left">Actual longitude</th>
<th align="left">Actual latitude</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">24.89347</td>
<td align="left">46.614989</td>
<td align="left">24.893469</td>
<td align="left">46.614989</td>
</tr>
<tr>
<td align="left">24.89348</td>
<td align="left">46.61502</td>
<td align="left">24.89348</td>
<td align="left">46.61505</td>
</tr>
<tr>
<td align="left">24.89349</td>
<td align="left">46.615048</td>
<td align="left">24.893498</td>
<td align="left">46.615048</td>
</tr>
<tr>
<td align="left">24.89351</td>
<td align="left">46.61508</td>
<td align="left">24.893508</td>
<td align="left">46.615092</td>
</tr>
<tr>
<td align="left">24.8935</td>
<td align="left">46.615083</td>
<td align="left">24.8935</td>
<td align="left">46.615083</td>
</tr>
<tr>
<td align="left">24.89349</td>
<td align="left">46.615056</td>
<td align="left">24.893497</td>
<td align="left">46.615076</td>
</tr>
<tr>
<td align="left">24.89346</td>
<td align="left">46.61499</td>
<td align="left">24.89347</td>
<td align="left">46.614991</td>
</tr>
<tr>
<td align="left">24.89345</td>
<td align="left">46.614936</td>
<td align="left">24.893448</td>
<td align="left">46.614966</td>
</tr>
<tr>
<td align="left">24.89346</td>
<td align="left">46.614928</td>
<td align="left">24.893457</td>
<td align="left">46.614958</td>
</tr>
<tr>
<td align="left">24.89347</td>
<td align="left">46.614915</td>
<td align="left">24.893496</td>
<td align="left">46.614918</td>
</tr>
<tr>
<td align="left">24.89354</td>
<td align="left">46.614881</td>
<td align="left">24.893541</td>
<td align="left">46.614891</td>
</tr>
<tr>
<td align="left">24.8936</td>
<td align="left">46.614847</td>
<td align="left">24.893642</td>
<td align="left">46.614847</td>
</tr>
<tr>
<td align="left">24.8936</td>
<td align="left">46.614847</td>
<td align="left">24.893602</td>
<td align="left">46.614867</td>
</tr>
<tr>
<td align="left">24.8936</td>
<td align="left">46.614818</td>
<td align="left">24.893597</td>
<td align="left">46.614819</td>
</tr>
<tr>
<td align="left">24.89348</td>
<td align="left">46.614572</td>
<td align="left">24.893495</td>
<td align="left">46.614572</td>
</tr>
<tr>
<td align="left">24.89348</td>
<td align="left">46.614572</td>
<td align="left">24.893483</td>
<td align="left">46.614592</td>
</tr>
<tr>
<td align="left">24.89351</td>
<td align="left">46.614557</td>
<td align="left">24.893558</td>
<td align="left">46.614557</td>
</tr>
<tr>
<td align="left">24.89351</td>
<td align="left">46.614557</td>
<td align="left">24.893509</td>
<td align="left">46.614578</td>
</tr>
<tr>
<td align="left">24.89353</td>
<td align="left">46.614578</td>
<td align="left">24.89354</td>
<td align="left">46.614578</td>
</tr>
<tr>
<td align="left">24.89362</td>
<td align="left">46.614776</td>
<td align="left">24.893619</td>
<td align="left">46.614779</td>
</tr>
<tr>
<td align="left">24.89367</td>
<td align="left">46.614877</td>
<td align="left">24.893698</td>
<td align="left">46.614897</td>
</tr>
<tr>
<td align="left">24.89367</td>
<td align="left">46.614877</td>
<td align="left">24.893675</td>
<td align="left">46.614882</td>
</tr>
<tr>
<td align="left">24.89366</td>
<td align="left">46.614884</td>
<td align="left">24.89368</td>
<td align="left">46.614894</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5"><label>5</label><title>Discussion</title>
<p>The results of this work indicate that the average response time for processing the command signal 0.5 s in order to avoid any accidents. The study also shows that the smart wheelchair program can be used and applied without an internet connection. Moreover, the proposed program achieves a significant results in presence of external noise. As shown in <xref ref-type="table" rid="table-6">Tab. 6</xref>, the accuracy, precision, recall, and F-Score for the implemented system have achieved an adequate results in comparing with the previous study in [<xref ref-type="bibr" rid="ref-26">26</xref>]. The experiment results revealed a statistically significant difference in the percentage values of the different categories, indicating a low risk of making incorrect predictions. For example, <xref ref-type="table" rid="table-7">Tab. 7</xref> shows that the difference between true and false predictions was about 187&#x0025;. For the evaluation of the performance of indoor and outdoor navigation, the results indicated that the wheelchair was able to accurately maneuvering and RMSD was significantly low.</p>
<p>Although, this study enhances the system&#x2019;s suitability for a variety of users. However, wheelchairs require additional research in the static, motion, and moment of inertia domains. Additionally, the existing model of voice recognition omitted a speaker identification mechanism. By identifying a speaker, wheelchair users can only take particular directions from an authorized individual by increasing their safety. When comparing this study with other studies regarding efficacy, dependability, and cost, we believe that our design overcomes numerous complexities. For instance, in a recent study conducted by Abdulghani et al. [<xref ref-type="bibr" rid="ref-22">22</xref>], an adaptive neuro-fuzzy control was constructed and tested to track motorized wheelchairs using voice recognition. To achieve a high level of precision, the design must incorporate a wireless network in which the wheelchair is treated as a node. In another study, a wheelchair was driven using an eye and voice. In this study, the authors used a voice-controlled mode in conjunction with a web camera in order to make the system more congenial and reliable [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
<p>Despite this work has different merits; however some limitations need to be treated and updated in the following stages. For example, the system needs to be equipped with a variable controller and GPS to make it more efficient and meet the needs of users. In addition, the mechanical design of the wheelchair needs to be modified to change the torque of inertia. This change will alleviate the sudden jump of a wheelchair during the initial start or stop.</p>
</sec>
<sec id="s6"><label>6</label><title>Conclusions</title>
<p>This research developed a voice-controlled wheelchair utilizing a low-cost and reliable technology. This technology uses a built-in voice recognition model combined with the CNN and LSTM models to train and classify five spoken commands. The design method used an Android smartphone (Flutter-based) app that connects with microcontrollers over an offline Wi-Fi hotspot. For the design and implementation of the experiment, ten native Arabic speakers produced a total of 2000 utterances of five words. The precision and usability of both indoor and outdoor navigation were tested using a range of disturbances. All voice commands have been given a normalized confusion matrix, precision, recall, and F-score. Voice recognition commands and wheelchair moves were demonstrated to be reliable in real-world testing. During indoor/outdoor maneuvering, it was also discovered that the RMSD estimated between the planned and real nodes was accurate. The ease of use, low cost, independence, and security are only a few of the benefits of the actual prototype. The device also has an emergency push-button as an additional safety element.</p>
</sec>
<sec id="s7"><label>7</label><title>Future Work</title>
<p>The system can be enhanced with GPS technology, allowing users to design their own routes. The system also can be equipped with ultrasonic sensors for added safety, as it will operate and ignore user commands if the chair gets too close to an obstacle that could cause an accident. Additional research could be done to see if users prefer voice control interfaces over brain control interfaces. The voice recognition model can be improved using the speaker identification algorithm to protect the disabled person&#x2019;s safety by accepting commands from only one user.</p>
</sec>
</body>
<back>
<ack>
<p>The author extend their appreciation to Deanship of Scientific Research, Majmaah University for supporting this work under Project Number (R-2022&#x2013;77).</p>
</ack>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> This research was funded by the deputyship for Research and Innovation, Ministry of Education, Saudi Arabia, Grant Number IFP-2020&#x2013;31.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Vignier</surname></string-name>, <string-name><given-names>J. F.</given-names> <surname>Ravaud</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Winance</surname></string-name>, <string-name><given-names>F. X.</given-names> <surname>Lepoutre</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Ville</surname></string-name></person-group>, &#x201C;<article-title>Demographics of wheelchair users in France: Results of national community-based handicaps-incapacit&#x00E9;s-d&#x00E9;pendance surveys</article-title>,&#x201D; <source>Journal of Rehabilitation Medicine</source>, vol. <volume>40</volume>, no. <issue>3</issue>, pp. <fpage>231</fpage>&#x2013;<lpage>239</lpage>, <year>2008</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Zeng</surname></string-name>, <string-name><given-names>C. L.</given-names> <surname>Teo</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Rebsamen</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Burdet</surname></string-name></person-group>, &#x201C;<article-title>A collaborative wheelchair system</article-title>,&#x201D; <source>IEEE Transactions on Neural Systems and Rehabilitation Engineering</source>, vol. <volume>16</volume>, no. <issue>2</issue>, pp. <fpage>161</fpage>&#x2013;<lpage>170</lpage>, <year>2008</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Carlson</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Demiris</surname></string-name></person-group>, &#x201C;<article-title>Collaborative control for a robotic wheelchair: Evaluation of performance, attention, and workload</article-title>,&#x201D; <source>IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)</source>, vol. <volume>42</volume>, no. <issue>3</issue>, pp. <fpage>876</fpage>&#x2013;<lpage>888</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Pineau</surname></string-name>, <string-name><given-names>R.</given-names> <surname>West</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Atrash</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Villemure</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Routhier</surname></string-name></person-group>, &#x201C;<article-title>On the feasibility of using a standardized test for evaluating a speech-controlled smart wheelchair</article-title>,&#x201D; <source>International Journal of Intelligent Control and Systems</source>, vol. <volume>16</volume>, no. <issue>2</issue>, pp. <fpage>124</fpage>&#x2013;<lpage>131</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sharmila</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Saini</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Choudhary</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yuvaraja</surname></string-name> and <string-name><given-names>S. G.</given-names> <surname>Rahul</surname></string-name></person-group>, &#x201C;<article-title>Solar powered multi-controlled smart wheelchair for disabled: Development and features</article-title>,&#x201D; <source>Journal of Computational and Theoretical Nanoscience</source>, vol. <volume>16</volume>, no. <issue>11</issue>, pp. <fpage>4889</fpage>&#x2013;<lpage>4900</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Hartman</surname></string-name> and <string-name><given-names>V. K.</given-names> <surname>Nandikolla</surname></string-name></person-group>, &#x201C;<article-title>Human-machine interface for a smart wheelchair</article-title>,&#x201D; <source>Journal of Robotics</source>, vol. <volume>2019</volume>, pp. <fpage>11</fpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Hu</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>Towards BCI-actuated smart wheelchair system</article-title>,&#x201D; <source>Biomedical Engineering Online</source>, vol. <volume>17</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>22</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Leaman</surname></string-name> and <string-name><given-names>H. M.</given-names> <surname>La</surname></string-name></person-group>, &#x201C;<article-title>A comprehensive review of smart wheelchairs: Past, present, and future</article-title>,&#x201D; <source>IEEE Transactions on Human-Machine Systems</source>, vol. <volume>47</volume>, no. <issue>4</issue>, pp. <fpage>486</fpage>&#x2013;<lpage>499</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Bourhis</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Horn</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Habert</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Pruski</surname></string-name></person-group>, &#x201C;<article-title>An autonomous vehicle for people with motor disabilities</article-title>,&#x201D; <source>IEEE Robotics &#x0026; Automation Magazine</source>, vol. <volume>8</volume>, no. <issue>1</issue>, pp. <fpage>20</fpage>&#x2013;<lpage>28</lpage>, <year>2001</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. C.</given-names> <surname>Simpson</surname></string-name></person-group>, &#x201C;<article-title>Smart wheelchairs: A literature review</article-title>,&#x201D; <source>Journal of Rehabilitation Research and Development</source>, vol. <volume>42</volume>, no. <issue>4</issue>, pp. <fpage>423</fpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Desai</surname></string-name>, <string-name><given-names>S. S.</given-names> <surname>Mantha</surname></string-name> and <string-name><given-names>V. M.</given-names> <surname>Phalle</surname></string-name></person-group>, &#x201C;<article-title>Advances in smart wheelchair technology</article-title>,&#x201D; in <conf-name>Int. Conf. on Nascent Technologies in Engineering (ICNTE)</conf-name>, <conf-loc>Navi Mumbai, India</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Rabhi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Mrabet</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Fnaiech</surname></string-name></person-group>, &#x201C;<article-title>Intelligent control wheelchair using a new visual joystick</article-title>,&#x201D; <source>Journal of Healthcare Engineering</source>, vol. <volume>2018</volume>, pp. <fpage>20</fpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yathunanthan</surname></string-name>, <string-name><given-names>L. U.</given-names> <surname>Chandrasena</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Umakanthan</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Vasuki</surname></string-name> and <string-name><given-names>S. R.</given-names> <surname>Munasinghe</surname></string-name></person-group>, &#x201C;<article-title>Controlling a wheelchair by use of EOG signal</article-title>,&#x201D; in <conf-name>4th Int. Conf. on Information and Automation for Sustainability</conf-name>, <conf-loc>Sri Lanka</conf-loc>, pp. <fpage>283</fpage>&#x2013;<lpage>288</lpage>, <year>2008</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Wieczorek</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kukla</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Rybarczyk</surname></string-name> and <string-name><given-names>&#x0141;</given-names> <surname>Wargu&#x0142;a</surname></string-name></person-group>, &#x201C;<article-title>Evaluation of the biomechanical parameters of human-wheelchair systems during ramp climbing with the use of a manual wheelchair with anti-rollback devices</article-title>,&#x201D; <source>Applied Sciences</source>, vol. <volume>10</volume>, no. <issue>23</issue>, pp. <fpage>8757</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C. S. L.</given-names> <surname>Tsui</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Jia</surname></string-name>, <string-name><given-names>J. Q.</given-names> <surname>Gan</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Hu</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Yuan</surname></string-name></person-group>, &#x201C;<article-title>EMG-Based hands-free wheelchair control with EOG attention shift detection</article-title>,&#x201D; in <conf-name>2007 IEEE Int. Conf. on Robotics and Biomimetics (ROBIO)</conf-name>, <conf-loc>Sanya, China</conf-loc>, pp. <fpage>1266</fpage>&#x2013;<lpage>1271</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Pan</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>A hybrid BCI system combining p300 and SSVEP and its application to wheelchair control</article-title>,&#x201D; <source>IEEE Transactions on Biomedical Engineering</source>, vol. <volume>60</volume>, no. <issue>11</issue>, pp. <fpage>3156</fpage>&#x2013;<lpage>3166</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. M.</given-names> <surname>Hosni</surname></string-name>, <string-name><given-names>H. A.</given-names> <surname>Shedeed</surname></string-name>, <string-name><given-names>M. S.</given-names> <surname>Mabrouk</surname></string-name> and <string-name><given-names>M. F.</given-names> <surname>Tolba</surname></string-name></person-group>, &#x201C;<article-title>EEG-EOG based virtual keyboard: Toward hybrid brain computer interface</article-title>,&#x201D; <source>Neuroinformatics</source>, vol. <volume>17</volume>, no. <issue>3</issue>, pp. <fpage>323</fpage>&#x2013;<lpage>341</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. D.</given-names> <surname>Olesen</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Das</surname></string-name>, <string-name><given-names>M. D.</given-names> <surname>Olsson</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Puthusserypady</surname></string-name></person-group>, &#x201C;<article-title>Hybrid EEG-EOG-based BCI system for vehicle control</article-title>,&#x201D; in <conf-name>9th Int. Winter Conf. on Brain-Computer Interface (BCI)</conf-name>, <conf-loc>Gangwon, South Korea</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. T.</given-names> <surname>Al-Qays</surname></string-name>, <string-name><given-names>B. B.</given-names> <surname>Zaidan</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Zaidan</surname></string-name> and <string-name><given-names>M. S.</given-names> <surname>Suzani</surname></string-name></person-group>, &#x201C;<article-title>A review of disability EEG based wheelchair control system: Coherent taxonomy, open challenges and recommendations</article-title>,&#x201D; <source>Computer Methods and Programs in Biomedicine</source>, vol. <volume>1</volume>, no. <issue>164</issue>, pp. <fpage>221</fpage>&#x2013;<lpage>237</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Aktar</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Jaharr</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Lala</surname></string-name></person-group>, &#x201C;<article-title>Voice recognition based intelligent wheelchair and GPS tracking system</article-title>,&#x201D; in <conf-name>Int. Conf. on Electrical, Computer and Communication Engineering (ECCE)</conf-name>, <conf-loc>Cox&#x0027;s Bazar, Babgladesh</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Raiyan</surname></string-name>, <string-name><given-names>M. S.</given-names> <surname>Nawaz</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Adnan</surname></string-name> and <string-name><given-names>M. H.</given-names> <surname>Imam</surname></string-name></person-group>, &#x201C;<article-title>Design of an arduino based voice-controlled automated wheelchair</article-title>,&#x201D; in <conf-name>IEEE Region 10 Humanitarian Technology Conf. (R10-HTC)</conf-name>, <conf-loc>Dhaka, Babgladesh</conf-loc>, pp. <fpage>267</fpage>&#x2013;<lpage>270</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. M.</given-names> <surname>Abdulghani</surname></string-name>, <string-name><given-names>K. M.</given-names> <surname>Al-Aubidy</surname></string-name>, <string-name><given-names>M. M.</given-names> <surname>Ali</surname></string-name> and <string-name><given-names>Q. J.</given-names> <surname>Hamarsheh</surname></string-name></person-group>, &#x201C;<article-title>Wheelchair neuro fuzzy control and tracking system based on voice recognition</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>20</volume>, no. <issue>10</issue>, pp. <fpage>2872</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sutikno Anam</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Saleh</surname></string-name></person-group>, &#x201C;<article-title>Voice controlled wheelchair for disabled patients based on CNN and LSTM</article-title>,&#x201D; in <conf-name>4th Int. Conf. on Informatics and Computational Sciences (ICICoS)</conf-name>, <conf-loc>Semarang Indonesia</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>O.</given-names> <surname>Abdel-Hamid</surname></string-name>, <string-name><given-names>A. R.</given-names> <surname>Mohamed</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Penn</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Convolutional neural networks
for speech recognition</article-title>,&#x201D; <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>, vol. <volume>22</volume>, no. <issue>10</issue>, pp. <fpage>1533--1545</fpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. S. I.</given-names> <surname>Sharifuddin</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Nordin</surname></string-name> and <string-name><given-names>A. M.</given-names> <surname>Ali</surname></string-name></person-group>, &#x201C;<article-title>Comparison of CNNs and SVM for voice control wheelchair</article-title>,&#x201D; <source>IAES International Journal of Artificial Intelligence</source>, vol. <volume>9</volume>, no. <issue>3</issue>, pp. <fpage>387</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Bakouri</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Alsehaimi</surname></string-name>, <string-name><given-names>H. F.</given-names> <surname>Ismail</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Alshareef</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Ganoun</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Steering a robotic wheelchair based on voice recognition system using convolutional neural networks</article-title>,&#x201D; <source>Electronics</source>, vol. <volume>11</volume>, no. <issue>1</issue>, pp. <fpage>168</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Alaeddine</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Jihene</surname></string-name></person-group>, &#x201C;<article-title>Deep network in network</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>33</volume>, pp. <fpage>1453</fpage>&#x2013;<lpage>1465</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Kuzmin</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Ignatiev</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Grafov</surname></string-name></person-group>, &#x201C;<chapter-title>Experience of developing a mobile application using flutter</chapter-title>,&#x201D; in <source>Information Science and Applications</source>, <publisher-loc>Singapore: Springer</publisher-loc>, pp. <fpage>571</fpage>&#x2013;<lpage>575</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M. L.</given-names> <surname>Napoli</surname></string-name></person-group>, <source>Beginning Flutter: A Hands on Guide to App Development</source>, <publisher-loc>New Jersey, USA</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons</publisher-name>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. S. B. A.</given-names> <surname>Ghaffar</surname></string-name>, <string-name><given-names>U. S.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Iqbal</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Rashid</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Hamza</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Improving classification performance of four class FNIRS-BCI using Mel Frequency Cepstral Coefficients (MFCC)</article-title>,&#x201D; <source>Infrared Physics &#x0026; Technology</source>, vol. <volume>112</volume>, pp. <fpage>103589</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. J.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>J. P.</given-names> <surname>Bae</surname></string-name>, <string-name><given-names>J. W.</given-names> <surname>Chung</surname></string-name>, <string-name><given-names>D. K.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>K. G.</given-names> <surname>Kim</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>New polyp image classification technique using transfer learning of network-in-network structure in endoscopic images</article-title>,&#x201D; <source>Scientific Reports</source>, vol. <volume>11</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G. V.</given-names> <surname>Houdt</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Mosquera</surname></string-name> and <string-name><given-names>G.</given-names> <surname>N&#x00E1;poles</surname></string-name></person-group>, &#x201C;<article-title>A review on the long short-term memory model</article-title>,&#x201D; <source>Artificial Intelligence Review</source>, vol. <volume>53</volume>, no. <issue>8</issue>, pp. <fpage>5929</fpage>&#x2013;<lpage>5955</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Anwer</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Waris</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Sultan</surname></string-name>, <string-name><given-names>S. I.</given-names> <surname>Butt</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Zafar</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Eye and voice-controlled human machine interface system for wheelchairs using image gradient approach</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>20</volume>, no. <issue>19</issue>, pp. <fpage>5510</fpage>, <year>2020</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>