<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">28631</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.028631</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Multilayer Neural Network Based Speech Emotion Recognition for&#x00A0;Smart&#x00A0;Assistance</article-title>
<alt-title alt-title-type="left-running-head">Multilayer Neural Network Based Speech Emotion Recognition for Smart Assistance</alt-title>
<alt-title alt-title-type="right-running-head">Multilayer Neural Network Based Speech Emotion Recognition for Smart Assistance</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Kumar</surname><given-names>Sandeep</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Haq</surname><given-names>MohdAnul</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Jain</surname><given-names>Arpit</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Jason</surname><given-names>C. Andy</given-names></name><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Moparthi</surname><given-names>Nageswara Rao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Mittal</surname><given-names>Nitin</given-names></name><xref ref-type="aff" rid="aff-5">5</xref></contrib>
<contrib id="author-7" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Alzamil</surname><given-names>Zamil S.</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>z.alzamil@mu.edu.sa</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science and Engineering, Koneru Lakshmaiah Education Foundation</institution>, <addr-line>Vaddeswaram, AP, 522502</addr-line>, <country>India</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Science, College of Computer and Information Sciences, Majmaah University</institution>, <addr-line>11952, Al-Majmaah</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Science and Engineering, Teerthanker Mahaveer University</institution>, <addr-line>Moradabad, Uttar Pradesh, 244001</addr-line>, <country>India</country></aff>
<aff id="aff-4"><label>4</label><institution>Department of Electronics and Communication Engineering, Sreyas Institute of Engineering and Technology</institution>, <addr-line>Hyderabad, 500068</addr-line>, <country>India</country></aff>
<aff id="aff-5"><label>5</label><institution>University Centre for Research and Development, Chandigarh University</institution>, <addr-line>Mohali, 140413, Punjab</addr-line>, <country>India</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Zamil S. Alzamil. Email: <email>z.alzamil@mu.edu.sa</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-08-16"><day>16</day>
<month>08</month>
<year>2022</year></pub-date>
<volume>74</volume>
<issue>1</issue>
<fpage>1523</fpage>
<lpage>1540</lpage>
<history>
<date date-type="received"><day>14</day><month>2</month><year>2022</year></date>
<date date-type="accepted"><day>24</day><month>5</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Kumar et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Kumar et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_28631.pdf"></self-uri>
<abstract>
<p>Day by day, biometric-based systems play a vital role in our daily lives. This paper proposed an intelligent assistant intended to identify emotions via voice message. A biometric system has been developed to detect human emotions based on voice recognition and control a few electronic peripherals for alert actions. This proposed smart assistant aims to provide a support to the people through buzzer and light emitting diodes (LED) alert signals and it also keep track of the places like households, hospitals and remote areas, etc. The proposed approach is able to detect seven emotions: worry, surprise, neutral, sadness, happiness, hate and love. The key elements for the implementation of speech emotion recognition are voice processing, and once the emotion is recognized, the machine interface automatically detects the actions by buzzer and LED. The proposed system is trained and tested on various benchmark datasets, i.e., Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) database, Acoustic-Phonetic Continuous Speech Corpus (TIMIT) database, Emotional Speech database (Emo-DB) database and evaluated based on various parameters, i.e., accuracy, error rate, and time. While comparing with existing technologies, the proposed algorithm gave a better error rate and less time. Error rate and time is decreased by 19.79&#x0025;, 5.13 s. for the RAVDEES dataset, 15.77&#x0025;, 0.01 s for the Emo-DB dataset and 14.88&#x0025;, 3.62 for the TIMIT database. The proposed model shows better accuracy of 81.02&#x0025; for the RAVDEES dataset, 84.23&#x0025; for the TIMIT dataset and 85.12&#x0025; for the Emo-DB dataset compared to Gaussian Mixture Modeling(GMM) and Support Vector Machine (SVM) Model.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Speech emotion recognition</kwd>
<kwd>classifier implementation</kwd>
<kwd>feature extraction and selection</kwd>
<kwd>smart assistance</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>In everyday life, Speech Emotion Recognition (SER) based devices play an essential role [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. The SER technology has been expanded in sports, e-learning, voice search, and even aircraft cockpit call centers. The main aim of the SER system is to understand the individual emotions of humans [<xref ref-type="bibr" rid="ref-3">3</xref>]. A highly significant feature of the voice recognition system is the reliance on the human voice, i.e., age, language, culture, temperament, environment, etc. The primary issue during speech recognition is more than one emotion expressed in the same vocabulary, so identifying the emotions is very critical [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. There are several traditional methods for feature extraction and classifier implementation models, i.e., KNN (K-Nearest Neighbour), GMM, ANN (Artificial neural network), DNN(Deep neural network), etc. which is used to recognize the emotions from speech but still not optimized [<xref ref-type="bibr" rid="ref-6">6</xref>]. To recognize the emotions from speech, we have proposed a novel SER model. The proposed SER model can detect seven emotions: worry, surprise, neutral, sadness, happiness, hate and love [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>]. While evaluating the results, the benchmark datasets can be divided into testing and training. Each speech dataset is transferred through the pre-processing function to extract the necessary function vector for features from the data. The vector training set is passed on to the correct classifier and the classifier then forecasts the emotion to validate a model [<xref ref-type="bibr" rid="ref-9">9</xref>]. The identification of speech emotions is carried out in four significant steps to generate speech-based output, i.e., acquisition, processing, output generation and the application of the extracted voice feature. The proposed module is mainly essential for the people treated with social distancing so that the people who are treating them can recognize their emotional condition and treat them well [<xref ref-type="bibr" rid="ref-10">10</xref>]. The proposed methodology also tends to boost the accuracy for better performance of speech emotion detection [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<p>The rest of the paper is organized as follows: Section 2 discusses the work related to emotion recognition through voice. Further, Section 3 concisely discusses the proposed method and the dataset details. Further, in Section 4, results obtained by the proposed method have been discussed and compared with the existing state-of-the-art methods. The last section concludes along with the future course of action. Let us know more about SER existing systems with their various feature extraction selection methods and classification algorithms in the following literature survey.</p>
</sec>
<sec id="s2"><label>2</label><title>Literature Work</title>
<p>Much research has been done in the area of emotion recognition through speech. There are many traditional methodologies and evaluation parameters, i.e., accuracy, error, time taken, etc., used by the SER system to recognize the emotion from speech but still exiting methods are the more complex and computational time taken were high. In this section, the literature survey is done based on the various parameters, i.e., approach based, evaluation based, results, databases, and conclusions of various models of SER systems as shown in <xref ref-type="table" rid="table-1">Tab. 1</xref>.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Approach, evaluation, and classifiers of existing SER system</title></caption>
 
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">S. No.</th>
<th align="center" rowspan="2">Author</th>
<th align="center" colspan="2">Methodology</th>
<th align="left" rowspan="2">Results</th>
<th align="left" rowspan="2">Conclusions</th>
<th align="left" rowspan="2">Data-bases</th>
</tr>
<tr>
<th align="left">Approach</th>
<th align="left">Evaluation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">Yogesh Kumar et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-1">1</xref>], 2019</td>
<td align="left">MFCC, LPCC, DELTA, FFT, PLP and RASTA</td>
<td align="left">KNN, SVM, Convolution neural network, Naive Bayes and RNN</td>
<td align="left">The qualified model proposed is tested with a test accuracy of 76.97&#x0025; in the entire file classification.</td>
<td align="left">Compared to other models, DNN models offer the best results.</td>
<td align="left">Emo-DB and LDC emotional prosody<break/>speech database</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">K. Prasada Rao et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-2">2</xref>], 2019</td>
<td align="left">SSR, PR</td>
<td align="left">MFCC and MSER</td>
<td align="left">The average efficiency of the group is 77&#x0025;. Happiness and disgust (78&#x0025;) are feelings with the highest awareness rate.</td>
<td align="left">Provides better efficiency than the uni-modal system.</td>
<td align="left">Emo-DB and Indian Face<break/>Database</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Maryam Imani et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-3">3</xref>], 2019</td>
<td align="left">Signal Processing and gesture recognition</td>
<td align="left">E-Learning Algorithms</td>
<td align="left">It can be used as a reference for the emotion detection of successful tutoring programs.</td>
<td align="left">It can be used as a reference for the emotion detection of successful tutoring programs.</td>
<td align="left">Science Direct database</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Wei Jiang et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-4">4</xref>], 2019</td>
<td align="left">IS10, MFCCs, eGemaps</td>
<td align="left">Heterogeneous Unification Module</td>
<td align="left">Compared to current cutting-edge solutions, 64&#x0025; of the proposed architecture improves the recognition efficiency.</td>
<td align="left">To improve classification efficiency, use the best multiple and heterogeneous features.</td>
<td align="left">IEMOCAP dataset</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Nithya Roopa S. et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-5">5</xref>], 2018</td>
<td align="left">Deep Learning</td>
<td align="left">Inception Net v3</td>
<td align="left">The precision rate is reached by approximately 38 percent.</td>
<td align="left">The highest accuracy rate during data validation is 0.8</td>
<td align="left">IEMOCAP datasets</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">Youddha Beer Singh et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-6">6</xref>], 2018</td>
<td align="left">MFCC</td>
<td align="left">HMM, KNN, SVM, ANN and GMM</td>
<td align="left">The highest precision of 79.6&#x0025; for classifier SVM and the lowest accuracy of 54.3&#x0025; for classification ELM.</td>
<td align="left">Even with different datasets, SVM has recorded the highest precision.</td>
<td align="left">Emo-DB, IITKGP-SESC and Wizard of Oz databases</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">Praseetha etal. [<xref ref-type="bibr" rid="ref-7">7</xref>], 2018</td>
<td align="left">FFT, MFCC</td>
<td align="left">DNN, RNN</td>
<td align="left">The accuracy of the DNN model is 89.96&#x0025;, and the accuracy of the GRU model is 95.82&#x0025;.</td>
<td align="left">The GRU model works very well for dynamic grouping compared to the DNN model.</td>
<td align="left">IA database</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">Rahhal Errattahi et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-8">8</xref>], 2018</td>
<td align="left">Evaluation methods of ASR</td>
<td align="left">ASR errors detection and correction techniques</td>
<td align="left">A data set consisting of five separate English articles of about 100 words read by five distinct speakers represented about 2.4&#x0025; of the proposed technique&#x2019;s error rate.</td>
<td align="left">Further work on the automatic correction of ASR failures is needed, and performance and reliability issues should be addressed.</td>
<td align="left">Nil</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">Sneha Lukose et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-9">9</xref>], 2017</td>
<td align="left">MFCC and End Point Detection</td>
<td align="left">SVM, GMM</td>
<td align="left">Provides 76.31&#x0025; positive performance with the GMM model and 81.57&#x0025; accuracy with the SVM model.</td>
<td align="left">SVM offers more incredible accuracy to extract the emotion from the speech signal.</td>
<td align="left">Emo-DB database</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">David Griol et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-10">10</xref>], 2017</td>
<td align="left">Feature Extraction and Selection Methods</td>
<td align="left">SVM, PNN and Naive Bayes</td>
<td align="left">The results show that all hypotheses feature classifiers, including recognition and fusion, can be used at every stage.</td>
<td align="left">Concerning precision and training time, ELM delivered the best results.</td>
<td align="left">Images descriptions, UAH and Let&#x2019;s Go corpus</td>
</tr>
<tr>
<td align="left">11</td>
<td align="left">Seyed H. Mohammadi et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-11">11</xref>], 2017</td>
<td align="left">STPK</td>
<td align="left">MLSA</td>
<td align="left">The average score for similarities for top submissions was defined correctly by about 70&#x0025;.</td>
<td align="left">Some health tests are best incorporated to eliminate listeners who perform below a minimum or inconsistently output threshold.</td>
<td align="left">CMU arctic speech database</td>
</tr>
<tr>
<td align="left">12</td>
<td align="left">S. Lugovi&#x0107; et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-12">12</xref>], 2016</td>
<td align="left">HCI</td>
<td align="left">SER Models</td>
<td align="left">Possibly monitor emotions and actions in different social groups by using emotional recognition in speech.</td>
<td align="left">It improves the efficiency of social technology structures and the benefit to the cost ratio is high as per computers.</td>
<td align="left">DES, BES and SUSAS databases</td>
</tr>
<tr>
<td align="left">13</td>
<td align="left">Haihua Jiang et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-13">13</xref>], 2017</td>
<td align="left">(KNN, GMM, SVM) &#x002B;INT, PIC, REA</td>
<td align="left">UDD, STEDD</td>
<td align="left">A higher level of precision of 80.30 percent for men and 75.96&#x0025; for women and an acceptable 75.00&#x0025; for men and 77.36&#x0025; for women. for women, a good sensitivity/specific ratio of 75.00&#x0025;</td>
<td align="left">The highest rating result was shown and both men and women had the best stability.</td>
<td align="left">INT, PIC, IEA databases</td>
</tr>
<tr>
<td align="left">14</td>
<td align="left">Isidoros Perikos et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-14">14</xref>], 2017</td>
<td align="left">FP, ECV, ESV, MEC</td>
<td align="left">Ensemble classifier</td>
<td align="left">The obtained findings suggest satisfactory results concerning the ability to perceive the role of emotions in the text and the emotional polarity of the text.</td>
<td align="left">Ensemble methodology is an effective way to combine different classification algorithms with helping classify textual emotions.</td>
<td align="left">WordNet database</td>
</tr>
<tr>
<td align="left">15</td>
<td align="left">Pavol Partila et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-15">15</xref>], 2014</td>
<td align="left">MFCC</td>
<td align="left">GMM, KNN and ANN</td>
<td align="left">Increased accuracy after training</td>
<td align="left">These three classifiers have shown the highest understanding of emotional indignation.</td>
<td align="left">Emo-DB database</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Kumar (2019)&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-1">1</xref>] obtained an accuracy of 76.98&#x0025; over the entire classification and concluded that the DNN model provides the best performance compared to other models. In the same year, Prasadarao&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-2">2</xref>], Imani&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-3">3</xref>] and Jiang&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-4">4</xref>] also worked on SSR signal processing, gesture recognition and MFCC models with various evaluation parameters based on MSER (Maximally stable extremal regions) e-learning algorithms resulting in 77&#x0025; overall classification success on joy as well as disgust emotions. In 2018, Errattahi&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-6">6</xref>], Singh&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-7">7</xref>] and Praseetha&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-8">8</xref>] and worked on MFCC (Mel-frequency cepstral coefficients), FFT(fast fourier transform) and while evaluating the methodology, the highest accuracy achieved 79.6&#x0025; and the lowest accuracy got 54.3&#x0025;.Another evaluation parameter, i.e., the error rate, was evaluated using the proposed method on a benchmark database. In contrast with existing state-of-the-art solutions, architecture proposed in other strategies improves recognition effectiveness by 64&#x0025;. In 2017 Lukose&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-9">9</xref>], Griol&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-10">10</xref>], Mohammadi&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-12">12</xref>] used MFCC, endpoint detection feature selection methods i.e., SVM, ANN, Naive Bayes for SER modules. Finally, 76.31&#x0025; of devices used the GM model and overall accuracy improved by 1.57&#x0025; using SVM models. In 2016, Lugovic&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-13">13</xref>] tested SCR models using HCI (Human-Computer Interaction).</p>
<p>Many authors have worked to improve the performance of the SER based models and but still there is a room of improvement [<xref ref-type="bibr" rid="ref-16">16</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]. The author concluded the possibility to monitor emotions and actions in different groups by using the emotion recognition module [<xref ref-type="bibr" rid="ref-24">24</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>]. This literature survey formulated the proposed approach with improved overall accuracy.</p>
</sec>
<sec id="s3"><label>3</label><title>Proposed Methodology</title>
<p>The proposed methodology is wholly based on a multilayer neural network SER system [<xref ref-type="bibr" rid="ref-14">14</xref>&#x2013;<xref ref-type="bibr" rid="ref-17">17</xref>]. The central aspect of the SER system is to recognize the speech emotions where the speech is given as an input to the system. After that, the multilayer neural network automatically makes feature extraction and selection. The entire proposed work of the speech recognition system is discussed step by step in the proposed algorithm, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>Flow chart of smart assistant based on SER</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-1.png"/></fig>
<sec id="s3_1"><label>3.1</label><title>Voice Acquisition</title>
<p>In the first step of speech recognition, is voice sample has been taken from benchmark datasets for further process.</p>
</sec>
<sec id="s3_2"><label>3.2</label><title>Feature Extraction and Selection</title>
<p>Feature extraction is a process of extracting the characteristics of the input sample to perform the classification task. In this model, the novel algorithm has been proposed and Mel Frequency Cepstral Coefficients is used in the speech recognition system for feature extraction. MFCC generates a discreet cosine transformation (DCT) of a natural short-term energy logarithm on the Mel frequency scale as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> and also specifies no of output samples that are considered for trainable parameters from a dataset. There are some advantages of choosing the functionalities and benefits of performing role selection before designing the data of MFCC.
<fig id="fig-2"><label>Figure 2</label><caption><title>Features extracted based on NN</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-2.png"/></fig>
<list list-type="bullet">
<list-item><p>Eliminates over fittings.</p></list-item>
<list-item><p>Enhances Accuracy: Modeling accuracy increases with less misleading results.</p></list-item>
<list-item><p>Reduces training time: Fewer data points increase the algorithm complexity and learn quicker algorithms.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mtext>C</mml:mtext></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>n</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mo>&#x2227;</mml:mo></mml:msup><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>&#x2227;</mml:mo></mml:msup><mml:mo stretchy="false">[</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:math></disp-formula></p></list-item>
</list>
where C[n] is real ceptron and x[n] is the real input signal</p>
<p>The extracted features are considered validation points to calculate the SER model&#x2019;s accuracy, time, and error rate. The Classifier is implemented to classify the emotions in speech after feature extraction and collection of voice samples so that the functions of the classifier understand the feelings. The emotions are categorized and remembered by training the dataset of various audio files for SER.</p>
</sec>
<sec id="s3_3"><label>3.3</label><title>Classification</title>
<p>The key and most important classification aspect is the Multilayer Perceptron (MLP) classification, which we used for our SER module. Multilayer Perceptron is an artificial neural feed-forward (ARN) type that consists of 3 node layers: one input, one hidden and one output layer, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. According to <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the input layer values are denoted as x<sub>1</sub>, x<sub>2</sub>&#x2026;&#x2026;x<sub>n</sub> output layer values as y<sub>1</sub>, y<sub>2</sub>&#x2026;&#x2026;y<sub>n</sub>, and hidden layer values as h<sub>1</sub>, h<sub>2</sub>&#x2026;&#x2026;h<sub>n</sub>. Each layer is fully connected to the next with the activation function forward feeding. We evaluate it based on the following equations. For each training sample, d, do: Propagate the input forward through the network and calculate network output for d&#x2019;s input values. Propagate the errors backward through a neural network and for each network output unit j
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Neural network of MLP</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-3.png"/></fig>
<p>For each hidden unit j
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mo>&#x2211;</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Update weights (w<sub>ji</sub>) by back-propagating error and using learning rule
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>w</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>;</mml:mo><mml:mspace width="thickmathspace" /><mml:mrow><mml:mtext mathvariant="italic">where</mml:mtext></mml:mrow><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:msub><mml:mo>.</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:mspace width="thickmathspace" /><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where &#x0394;w&#x003D;Predicted desired output, d&#x003D;The learning rate is usually less than 1, &#x03B7;&#x003D;Input data. After summation of the input values substitute the value in the sigmoid function
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mtext>Sigmoid  f</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>MLP uses a supervised training method for backpropagation characterized by linear perceptron multi-layered and non-linear activation functions. MLP cannot linearly separate distinguishable information, but this process is evaluated in real-time. The training data is stored in the target folder as an audio sample format compared with the input data. MLP recognizes the emotion based on the previous outcomes to improve the accuracy by learning rate (0.9) from the previous data, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. The algorithm recognizes the emotion as previous when similar binary data is found. If not, it stores the input data and learns from it. In our proposed methodology, we trained our dataset based on voice modulations of the speaker, voice input and prediction sentences, as shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>Training of the input values based on MLP</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-4.png"/></fig>
<fig id="fig-5"><label>Figure 5</label><caption><title>PSEUDO code</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-5.png"/></fig>
<p>In this way, the MLP-based neural network predicts emotion based on the input audio signal and a few output samples on various emotions are shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. The neural network&#x2019;s performance depends on the number of layers concealed in the network. As the number of hidden layers&#x2019; increases, predicting emotions will increase and help the neural network recognize emotions more accurately. Our neural network consists of 17 hidden layers that analyze speech characteristics based on the extraction and selection process. After the classification is done, we can recognize the following seven emotions based on the ANN by giving samples of voice messages, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>Sample images of the outputs of recognizing the emotions (Sadness, Worry, Happiness, Surprise, Love, Neutral, Hate) based on the audio signal where the x-axis represents the period taken to record the audio signal and the y-axis represents the frequency of speech signal in decibels (DB)</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-6.png"/></fig>
</sec>
<sec id="s3_4"><label>3.4</label><title>Hardware Module</title>
<p>The proposed module receives the input data through a microphone, and output is observed in the form of alerts created by a buzzer and LED&#x2019;s to the user, as shown in <xref ref-type="fig" rid="fig-7">Figs. 7a</xref> and <xref ref-type="fig" rid="fig-7">7b</xref>. Firstly, hardware components are used as raspberry pi3 model B&#x002B;, which looks like a small card-sized electronic board, monitor/system as a screen to show efficiency [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]. A compilation of code &#x0026; display results, statements &#x0026; to start the process, connecting cables to connect the system to raspberry pi3 &#x0026; buzzer and LED and interface with raspberry, mic for the real-time speech recognition with high quality, wires, raspberry and for a system power supply should be connected or if laptop no need of extra power supplier.</p>
<fig id="fig-7"><label>Figure 7</label><caption><title>(a) Complete hardware setup (b) Smart assistant kit</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-7.png"/></fig>
<p>Here we used python version 3.8.3 software as it was advanced and better. This software was inbuilt in the raspberry PI and is worked by using a VNC viewer (Virtual Network Computing). The program of speech and recognition will run in this VNC viewer software. This VNC viewer gives a cryptographic representation of your pi and we use SSH (Secure Shell), which is standard to support encrypted data transfer between two computers and gives access to the pi via terminal [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Experimental Setup and Database</title>
<p>All tests were performed on Python 3.8 using Intel Core I5 system specification, 8 GB RAM, RADEON Graphics Card, etc. This software was inbuilt in the raspberry PI and is worked by using a VNC viewer (Virtual Network Computing). The program of speech and recognition will run in this VNC viewer software. The entire proposed module works on speech emotion recognition where the input is given through a microphone and output is observed in the form of alerts created by peripherals as assistants for the user. The main component in the module is raspberry Pi which carries the complete control of the module and will assign work to each peripheral and give a response according to the speech input is given to the system.</p>
<sec id="s4_1"><label>4.1</label><title>Standard Datasets</title>
<p>For evaluation of the proposed work, three benchmark datasets were used, i.e., RAVDNESS, TIMIT Corpus and Emo-DB and description of all three datasets as shown in <xref ref-type="table" rid="table-2">Tab. 2</xref>.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Details of benchmark datasets</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Sr. No.</th>
<th align="left">Dataset Name</th>
<th align="left">Remarks</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">RAVDNESS</td>
<td align="left">RAVDESS contains 7356 files (total size: 24.8 GB). Datasets consist of 24 professional actors (12 females and 12 males).<break/>Speech includes expressions of calm, happiness, sad, anger, fear, surprise, and disgust expressions.</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">TIMIT Corpus</td>
<td align="left">TIMIT contains broadband recordings of 630 speakers of eight significant dialects of American English, each reading ten phonetically rich sentences.<break/>The TIMIT corpus includes time-aligned orthographic, phonetic and word transcriptions, and a 16-bit, 16&#x2005;kHz speech waveform file for each utterance.</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Emo-DB</td>
<td align="left">EMO-DB is another database used in our project for comparing accuracy with our database.<break/>It contains about 500 utterances spoken by actors in a happy, angry, anxious, fearful, bored and disgusted way, as well as a neutral version.<break/>You can choose utterances from 10 different actors and ten different texts.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2"><label>4.2</label><title>Evaluation Parameter</title>
<p>As we already mentioned, the proposed module evaluated the various parameters, i.e., Accuracy, Error Rate and Time Taken, on three benchmark datasets.</p>
<sec id="s4_2_1"><label>4.2.1</label><title>Error Rate</title>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mo>&#x2219;</mml:mo><mml:mrow><mml:mtext>Word Error Rate</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>WER</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mi>C</mml:mi><mml:mi>N</mml:mi></mml:mfrac><mml:mspace width="thickmathspace" /><mml:mrow><mml:mtext>for discrete speech,</mml:mtext></mml:mrow></mml:math></disp-formula>
<disp-formula id="ueqn-1">
<mml:math id="mml-ueqn-1" display="block"><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>D</mml:mi><mml:mo>+</mml:mo><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mi>N</mml:mi></mml:mfrac><mml:mspace width="thickmathspace" /><mml:mrow><mml:mtext>for continuous speech</mml:mtext></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula>
<p><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>Authentication Accuracy</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>AA</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>S</mml:mi><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>where SA is the number of successful authentication attempts and TA is the total several authentication attempts.</p>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>Command Error rate</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>CER</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:mfrac><mml:mspace width="thickmathspace" /><mml:mrow><mml:mtext>for commands with no attribute</mml:mtext></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mfrac><mml:mrow><mml:mi>C</mml:mi><mml:mi>C</mml:mi><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>C</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:mfrac><mml:mspace width="thickmathspace" /><mml:mrow><mml:mtext>for commands with attribute</mml:mtext></mml:mrow></mml:math></disp-formula>
<p>where, TC is the total number of commands issued.</p>
<p>CC is the number of correctly carried out command&#x2019;s beta software.</p>
<p>CA is the total number of attributes correctly interpreted by the software.</p>
<p>TA is the total number of attributes issued.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>Rejection rate</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>RR</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>R</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where NRR number of rejected unwanted sounds.</p>
<p>TRR total number of unwanted sounds that should be rejected.
<list list-type="bullet">
<list-item><p>Accuracy testing for varying sound pressure level values (SPL)
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>S</mml:mi><mml:mi>P</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">variance</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>A</mml:mi><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>A</mml:mi><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p></list-item>
</list>
where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:math></inline-formula> is the accuracy-test at a greater distance.</p>
<p><inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:math></inline-formula> is the accuracy-test at a short distance.</p>
<p>&#x2022; Accuracy testing for wearing signal to noise ratios (SNR)
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>S</mml:mi><mml:mi>N</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">variance</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>A</mml:mi><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>A</mml:mi><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:math></inline-formula> is the accuracy test with background noise.</p>
<p><inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi></mml:math></inline-formula> is the accuracy test without background noise.</p>
</sec>
<sec id="s4_2_2"><label>4.2.2</label><title>Accuracy</title>
<p>
<list list-type="bullet">
<list-item><p>SPAB (Speech Processing Accuracy Benchmark)
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>A</mml:mi><mml:mo>+</mml:mo><mml:mi>S</mml:mi><mml:mi>P</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">variance</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>S</mml:mi><mml:mi>N</mml:mi><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">variance</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p></list-item>
</list></p>
</sec>
<sec id="s4_2_3"><label>4.2.3</label><title>Time Taken</title>
<p>Time Taken is calculated based on the output generated after the compilation of the proposed methodology is shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>.</p>
<fig id="fig-8"><label>Figure 8</label><caption><title>Emotion recognized within time</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-8.png"/></fig>
</sec>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Results and Discussion</title>
<sec id="s5_1"><label>5.1</label><title>Results</title>
<p>The total no of speech samples acquired by the standard databases consists of various male and female voice samples. The MLP based proposed work classified them based on emotions: worry, surprise, neutral, sadness, happiness, hate, and love, as shown in <xref ref-type="table" rid="table-3">Tab. 3</xref>. Here we have used all three datasets for training and testing. The K-fold cross-validation technique was used in our module. During testing of the module, the trained data is saved in .csv format, which is 70&#x0025; and testing data which is 30&#x0025; of total samples of all datasets based on test train and split validation technique.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>No of samples based on TIMIT, Emo-DB and RAVDESS datasets</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Emotions</th>
<th align="center" colspan="3">Number of Speech Samples</th>
</tr>
<tr>
<th/>
<th align="left">Emo-DB</th>
<th align="left">TIMIT</th>
<th align="left">RAVDESS</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Hate/Anger</td>
<td align="left">127</td>
<td align="left">91</td>
<td align="left">1052</td>
</tr>
<tr>
<td align="left">Worry</td>
<td align="left">81</td>
<td align="left">63</td>
<td align="left">1073</td>
</tr>
<tr>
<td align="left">Love</td>
<td align="left">46</td>
<td align="left">76</td>
<td align="left">1093</td>
</tr>
<tr>
<td align="left">Surprise</td>
<td align="left">69</td>
<td align="left">118</td>
<td align="left">1085</td>
</tr>
<tr>
<td align="left">Happiness</td>
<td align="left">71</td>
<td align="left">102</td>
<td align="left">1021</td>
</tr>
<tr>
<td align="left">Sadness</td>
<td align="left">62</td>
<td align="left">83</td>
<td align="left">1029</td>
</tr>
<tr>
<td align="left">Neutral</td>
<td align="left">79</td>
<td align="left">97</td>
<td align="left">1003</td>
</tr>
<tr>
<td align="left">Total</td>
<td align="left">535</td>
<td align="left">630</td>
<td align="left">7356</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5_1_1"><label>5.1.1</label><title>Software Part</title>
<p>In our proposed methodology, using MLP classifier, the overall efficiency increased inaccuracy, time taken and the error rate reduced. The recognition rate was acquired as 81.02&#x0025; for the RAVADEES dataset, 86.71&#x0025; for Emo-DB and 84.23&#x0025; for the TIMIT dataset as shown in <xref ref-type="table" rid="table-4">Tab. 4</xref> and graphical representation of our proposed work on three benchmark datasets are shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. The error rate is calculated using equation six and the result acquired is a decrease of error rate by 19.79&#x0025; for the RAVDEES dataset, 15.77&#x0025; for the TIMIT dataset, and 14.88&#x0025; for the Emo-DB dataset. Similarly, the time taken is also evaluated from the output sample as shown in <xref ref-type="table" rid="table-5">Tab. 5</xref>, which 5.13 s decreases for the RAVDEES dataset, 3.62 s for the TIMIT dataset and 0.01 s for the Emo-DB dataset. After the precise analysis of the results, it is established that the proposed model is giving better results than the existing state-of-art-methodology and could be a big boon for human life.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>SER accuracy based on the SPAB score in (&#x0025;)</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left"/>
<th align="left">WER</th>
<th align="left">AA</th>
<th align="left">CER</th>
<th align="left">RR</th>
<th align="left">SPL</th>
<th align="left">SNR</th>
<th align="left">SPAB</th>
<th align="left">Accuracy (&#x0025;)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">EMO-DB</td>
<td align="left">0.9642</td>
<td align="left">1</td>
<td align="left">0.966</td>
<td align="left">0.87</td>
<td align="left">0.942</td>
<td align="left">0.959</td>
<td align="left">5.173</td>
<td align="left">86.2</td>
</tr>
<tr>
<td align="left">TIMIT</td>
<td align="left">0.9464</td>
<td align="left">1</td>
<td align="left">0.9</td>
<td align="left">0.75</td>
<td align="left">0.912</td>
<td align="left">0.955</td>
<td align="left">5.124</td>
<td align="left">85.43</td>
</tr>
<tr>
<td align="left">RAVDESS</td>
<td align="left">0.9136</td>
<td align="left">0.933</td>
<td align="left">0.889</td>
<td align="left">0.625</td>
<td align="left">0.862</td>
<td align="left">0.891</td>
<td align="left">4.812</td>
<td align="left">80.21</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-9"><label>Figure 9</label><caption><title>Comparison of accuracy with the state-of-art- methods (&#x0025;)</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_28631-fig-9.png"/></fig>
<table-wrap id="table-5"><label>Table 5</label><caption><title>Comparison of accuracy with the state-of-art- methods (&#x0025;)</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Dataset</th>
<th align="left">Reference</th>
<th align="left">Methodology</th>
<th align="left">Error Rate (&#x0025;)</th>
<th align="left">Time (sec)</th>
<th align="left">Accuracy (&#x0025;)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="3">RAVDESS</td>
<td align="left">[<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td align="left">GMM</td>
<td align="left">61</td>
<td align="left">9.31</td>
<td align="left">39</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td align="left">SVM</td>
<td align="left">23.68</td>
<td align="left">7.35</td>
<td align="left">76.32</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="left">MLP based on ANN</td>
<td align="left">19.79</td>
<td align="left">5.13</td>
<td align="left">81.02</td>
</tr>
<tr>
<td align="left" rowspan="3">TIMIT Corpus</td>
<td align="left">[<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td align="left">GMM</td>
<td align="left">45.39</td>
<td align="left">6.82</td>
<td align="left">54.61</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-1">1</xref>]</td>
<td align="left">SVM</td>
<td align="left">22.07</td>
<td align="left">4.32</td>
<td align="left">77.93</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="left">MLP based on ANN</td>
<td align="left">15.77</td>
<td align="left">3.62</td>
<td align="left">84.23</td>
</tr>
<tr>
<td align="left" rowspan="3">Emo-DB</td>
<td align="left">[<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td align="left">GMM</td>
<td align="left">23.18</td>
<td align="left">0.32</td>
<td align="left">76.82</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td align="left">SVM</td>
<td align="left">17.7</td>
<td align="left">0.19</td>
<td align="left">82.30</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="left">MLP based on ANN</td>
<td align="left">14.88</td>
<td align="left">0.01</td>
<td align="left">86.71</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_1_2"><label>5.1.2</label><title>Hardware Part</title>
<p>After recognition, when the output comes under these recognitions, i.e., sadness, worry, hate and love the buzzer is used as an alert assistant. The recognition of emotions is categorized based on the number of beep sounds made by the buzzer. If the emotion is recognized as love, the buzzer sounds one beep, two beeps for worry, three beeps for hate, and four for love. When the emotion is recognized as happiness and surprise, the green LED glows, and the white LED glows to detect neutral emotion. This hardware component helps the people who are kept alone for social distancing to avoid contagious infections and also the staff who are treating the patient could monitor the patient&#x2019;s emotional condition and treat him well. In this way, buzzers and LEDs alert the following emotions, as shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. When the proposed module cannot recognize the emotion from speech, the system will display a message that it could not recognize speech then automatically; the system goes to recording mode.</p>
</sec>
</sec>
<sec id="s5_2"><label>5.2</label><title>Discussion</title>
<p>From the deep analysis of the proposed module, results show that the GMM model produces an accuracy of 36&#x0025;, 54.61&#x0025; and 76.82&#x0025; and when it comes to SVM Model, which gives 76.32&#x0025;, 77.93&#x0025; and 82.30&#x0025; for RAVDEES, TIMIT and Emo-DB datasets. The other evaluation parameters, error rate and time, is taken were mentioned in <xref ref-type="table" rid="table-5">Tab. 5</xref>. Many authors have worked on the same dataset and used various evaluation parameters per the analysis. In the RAVDESS dataset GMM method has produced a higher error and the proposed method has made minor errors comparatively, which is 4&#x0025; less than the existing one. At the same time, execution time is also reduced by 2.22 s and accuracy is improved by 4.70&#x0025;, which is a good improvement in accuracy and time. Comparative analysis is performed on the TIMIT Corpus dataset and again, GMM has produced the highest error rate of 45.39&#x0025; and at the same time, the proposed method has reduced the error rate by 6.30&#x0025; execution time is reduced by 0.70 s which does not show the significant difference between the existing and proposed methodology. But at the same time, accuracy is improved by 6.30&#x0025;. Emo-DB is also used for the performance analysis of the proposed methodology and once again GMM method doesn&#x2019;t perform well and produces a higher error rate. The error rate is decreased by 2.82&#x0025; using the proposed methodology. Execution time is not improved much through the proposed methodology, but still, it shows some improvement. The accuracy is also enhanced by 4.41&#x0025; and it&#x2019;s a good improvement. Finally, we can conclude that our proposed work outperforms all three benchmark datasets. As emotion recognition through speech is a growing and exciting area of research, it could be helpful for human beings in many ways. The proposed method assists the people through LED alerts and buzzers. The key elements for the implementation of Speech Emotion Recognition (SER) are voice processing, and once the emotion is recognized, the machine interface automatically detects the actions by Buzzer and LED. But still, there is a scope of improvement in terms of accuracy because, as of now, we have achieved the highest accuracy of 86.71&#x0025;. Hence accuracy can be improved in future work.</p>
</sec>
</sec>
<sec id="s6"><label>6</label><title>Conclusion and Future Scope</title>
<p>In this paper, proposed methodology provides a better result for the speech emotion recognition system over the seven emotions by MLP classification for all benchmark datasets which are considered for the research. While analyzing the results, it is observed that the MLP classifier has a high accuracy rate as compared to other state-of-art-method for detecting emotions from the speech signal. This proposed methodology leads us to conclude that speech recognition plays a vital role in better supporting individuals than other speech aids. In addition, the proposed module can help us to track our loved ones and at the same time we can alert them. This will be a small contribution to us to our present situation of corona virus and its victims, where it could monitor the health condition of the people by maintaining social distance and provide better support to the people, alerts them through buzzer and LEDs which should be located in the audible and visible range.</p>
<p>In future, model could be trained and tested with real-time datasets to improve the performance of the system. Including camera modules and other peripherals could make the proposed methodology more efficient to serve people and it could also use in the defense for security purpose.</p>
</sec>
</body>
<back>
<ack>
<p>Zamil S. Alzamil would like to thank Deanship of Scientific Research at Majmaah University for supporting this work under Project No. R-2022-166.</p>
</ack>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> Zamil S. Alzamil would like to thank Deanship of Scientific Research at Majmaah University for supporting this work under Project No. R-2022-166.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Mahajan</surname></string-name></person-group>, &#x201C;<article-title>Machine learning-based speech emotions recognition system</article-title>,&#x201D; <source>International Journal of Scientific &#x0026; Technology Research</source><italic>,&#x201D;</italic> vol. <volume>8</volume>, no. <issue>7</issue>, pp. <fpage>722</fpage>&#x2013;<lpage>729</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. K.</given-names> <surname>Prasada</surname></string-name>, <string-name><given-names>M. S.</given-names> <surname>Rao</surname></string-name> and <string-name><given-names>N. H.</given-names> <surname>Chowdary</surname></string-name></person-group>, &#x201C;<article-title>An integrated approach to emotion recognition and gender classification</article-title>,&#x201D; <source>Journal of Visual Communication and Image Representation</source>, vol. <volume>60</volume>, no. <issue>1</issue>, pp. <fpage>339</fpage>&#x2013;<lpage>345</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Imani</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Ali Montazer</surname></string-name></person-group>, &#x201C;<article-title>A survey of emotion recognition methods with emphasis on e-learning environments</article-title>,&#x201D; <source>Journal of Network and Computer Applications</source>, vol. <volume>147</volume>, no. <issue>2</issue>, pp. <fpage>102423</fpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J. S.</given-names> <surname>Jin</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Han</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Speech emotion recognition with heterogeneous feature unification of deep neural network</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>19</volume>, no. <issue>12</issue>, pp. <fpage>2730</fpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. N.</given-names> <surname>Roop</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Prabhakaran</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Betty</surname></string-name></person-group>, &#x201C;<article-title>Speech emotion recognition using deep learning</article-title>,&#x201D; <source>International Journal of Recent Technology and Engineering</source>, vol. <volume>7</volume>, no. <issue>4</issue>, pp. <fpage>247</fpage>&#x2013;<lpage>250</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Errattahi</surname></string-name>, <string-name><given-names>A. E.</given-names> <surname>Hannani</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Ouahmane</surname></string-name></person-group>, &#x201C;<article-title>Automatic speech recognition errors detection and correction: A review</article-title>,&#x201D; <source>International Conference on Natural Language and Speech Processing</source>, vol. <volume>128</volume>, pp. <fpage>34</fpage>&#x2013;<lpage>37</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y. B.</given-names> <surname>Singh</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Goel</surname></string-name></person-group>, &#x201C;<article-title>Survey on human emotion recognition: Speech database, features and classification</article-title>,&#x201D; in <conf-name>Int. Conf. on Advances in Computing, Communication Control and Networking</conf-name>, <conf-loc>Greater Noida, India</conf-loc>, pp. <fpage>298</fpage>&#x2013;<lpage>301</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V. M.</given-names> <surname>Praseetha</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Vadivel</surname></string-name></person-group>, &#x201C;<article-title>Deep learning models for speech emotion recognition</article-title>,&#x201D; <source>Journal of Computer Science</source>, vol. <volume>14</volume>, no. <issue>11</issue>, pp. <fpage>1577</fpage>&#x2013;<lpage>1587</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Lukoseand</surname></string-name> and <string-name><given-names>S. S.</given-names> <surname>Upadhya</surname></string-name></person-group>, &#x201C;<article-title>Music player based on emotion recognition of voice signals</article-title>,&#x201D; in <conf-name>Int. Conf. on Intelligent Computing, Instrumentation and Control Technologies</conf-name>, <conf-loc>Kerala, India</conf-loc>, pp. <fpage>1751</fpage>&#x2013;<lpage>1754</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Griol</surname></string-name>, <string-name><given-names>J. M.</given-names> <surname>Molina</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Callejas</surname></string-name></person-group>, &#x201C;<article-title>Combining speech-based and linguistic classifiers to recognize emotion in user spoken utterances</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>326</volume>, no. <issue>1</issue>, pp. <fpage>132</fpage>&#x2013;<lpage>140</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yan</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Investigation of different speech types and emotions for detecting depression using different classifiers</article-title>,&#x201D; <source>Speech Communication</source>, vol. <volume>90</volume>, no. <issue>2</issue>, pp. <fpage>39</fpage>&#x2013;<lpage>46</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. H.</given-names> <surname>Mohammadi</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Kain</surname></string-name></person-group>, &#x201C;<article-title>An overview of voice conversion systems</article-title>,&#x201D; <source>Speech Communication</source>, vol. <volume>88</volume>, no. <issue>1</issue>, pp. <fpage>65</fpage>&#x2013;<lpage>82</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Lugovi&#x0107;</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Dun&#x0111;er</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Horvat</surname></string-name></person-group>, &#x201C;<article-title>Techniques and applications of emotion recognition in speech</article-title>,&#x201D; in <conf-name>39th Int. Convention on Information and Communication Technology, Electronics and Microelectronics</conf-name>, <conf-loc>Opatija, Croatia</conf-loc>, pp. <fpage>1551</fpage>&#x2013;<lpage>1556</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Perikos</surname></string-name>, and <string-name><given-names>I.</given-names> <surname>Hatzilygeroudis</surname></string-name></person-group>, &#x201C;<article-title>Recognizing emotions in text using ensemble of classifiers</article-title>,&#x201D; <source>Engineering Applications of Artificial Intelligence</source>, vol. <volume>51</volume>, no. <issue>1</issue>, pp. <fpage>191</fpage>&#x2013;<lpage>201</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Partila</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Tovarek</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Voznak</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Safarik</surname></string-name></person-group>, &#x201C;<article-title>Classification methods accuracy for speech emotion recognition system</article-title>,&#x201D; <source>Nostradamus 2014: Prediction, Modeling and Analysis of Complex Systems Prediction</source>, vol. <volume>289</volume>, no. <issue>3</issue>, pp. <fpage>439</fpage>&#x2013;<lpage>447</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. A.</given-names> <surname>Jason</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>An appraisal on speech and emotion recognition technologies based on machine learning</article-title>,&#x201D; <source>International Journal of Recent Technology and Engineering</source>, vol. <volume>8</volume>, no. <issue>5</issue>, pp. <fpage>211</fpage>&#x2013;<lpage>228</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Singh</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Gender classification using machine learning with multi-feature method</article-title>,&#x201D; in <conf-name>IEEE 9th Annual Computing and Communication Workshop and Conference</conf-name>, <conf-loc>Las Vegas, NV, USA</conf-loc>, pp. <fpage>0648</fpage>&#x2013;<lpage>0653</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Singh</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Live detection of face using machine learning with multi-feature method</article-title>,&#x201D; <source>Wireless Personal Communications</source>, vol. <volume>103</volume>, no. <issue>3</issue>, pp. <fpage>2353</fpage>&#x2013;<lpage>2375</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name>,<string-name><given-names>S.</given-names> <surname>Singh</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Automatic live facial expression detection using genetic algorithm with haar wavelet features and SVM</article-title>,&#x201D; <source>Wireless Personal Communications</source>, vol. <volume>103</volume>, no. <issue>3</issue>, pp. <fpage>2435</fpage>&#x2013;<lpage>2453</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Singh</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<chapter-title>Multiple face detection using hybrid features with SVM classifier</chapter-title>,&#x201D; in <source>Data and Communication Networks, Data and Communication Networks</source>, <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>, pp. <fpage>253</fpage>&#x2013;<lpage>265</lpage>, <year>2019</year>. Online Available: <uri xlink:href="https://link.springer.com/chapter/10.1007/978-981-13-2254-9_23">https://link.springer.com/chapter/10.1007/978-981-13-2254-9_23</uri>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Cao</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Mao</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Xu</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Speech emotion recognition based on feature selection and extreme learning machine decision tree</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>273</volume>, no. <issue>2</issue>, pp. <fpage>253</fpage>&#x2013;<lpage>265</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. A. R.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>M. K.</given-names> <surname>Jain</surname></string-name></person-group>, &#x201C;<article-title>Feature point detection for repacked android apps</article-title>,&#x201D; <source>Intelligent Automation &#x0026; Soft Computing</source>, vol. <volume>26</volume>, no. <issue>6</issue>, pp. <fpage>1359</fpage>&#x2013;<lpage>1373</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Binti</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ahmad</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Mahmoud</surname></string-name> and <string-name><given-names>R. M.</given-names> <surname>Mehmood</surname></string-name></person-group>, &#x201C;<article-title>A pursuit of sustainable privacy protection in big data environment by an optimized clustered-purpose based algorithm</article-title>,&#x201D; <source>Intelligent Automation &#x0026; Soft Computing</source>, vol. <volume>26</volume>, no. <issue>6</issue>, pp. <fpage>1217</fpage>&#x2013;<lpage>1231</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Shilpa</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Lakhwani</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Three-dimensional wireframe model of medical and complex images using cellular logic array processing techniques</article-title>,&#x201D; in <conf-name>Int. Conf. on Soft Computing and Pattern Recognition</conf-name>, <conf-loc>Switzerland</conf-loc>, pp. <fpage>196</fpage>&#x2013;<lpage>207</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Shilpa</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Lakhwani</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Three dimensional objects recognition &#x0026; pattern recognition technique; related challenges: A review</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>81</volume>, no. <issue>2</issue>, pp. <fpage>17303</fpage>&#x2013;<lpage>17346</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X. R.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>W. F.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>X. M.</given-names> <surname>Sun</surname></string-name> and <string-name><given-names>S. K.</given-names> <surname>Jha</surname></string-name></person-group>, &#x201C;<article-title>A robust 3-D medical watermarking based on wavelet transform for data protection</article-title>,&#x201D; <source>Computer Systems Science &#x0026; Engineering</source>, vol. <volume>41</volume>, no. <issue>3</issue>, pp. <fpage>1043</fpage>&#x2013;<lpage>1056</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X. R.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>X. M.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Sun</surname></string-name> and <string-name><given-names>S. K.</given-names> <surname>Jha</surname></string-name></person-group>, &#x201C;<article-title>Robust reversible audio watermarking scheme for telemedicine and privacy protection</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>71</volume>, no. <issue>2</issue>, pp. <fpage>3035</fpage>&#x2013;<lpage>3050</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Choudhary</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Lakhwani</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Agrwal</surname></string-name></person-group>, &#x201C;<article-title>An efficient hybrid technique of feature extraction for facial expression recognition using AdaBoost classifier</article-title>,&#x201D; <source>International Journal of Engineering Research &#x0026; Technology</source><italic>, v</italic>ol. <volume>8</volume>, no. <issue>1</issue>, pp. <fpage>30</fpage>&#x2013;<lpage>41</lpage>, <year>2012</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>