<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">18406</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2021.018406</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Mental Illness Disorder Diagnosis Using Emotion Variation Detection from Continuous English Speech</article-title>
<alt-title alt-title-type="left-running-head">Mental Illness Disorder Diagnosis Using Emotion Variation Detection from Continuous English Speech</alt-title>
<alt-title alt-title-type="right-running-head">Mental Illness Disorder Diagnosis Using Emotion Variation Detection from Continuous English Speech</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western">
<surname>Lalitha</surname>
<given-names>S.</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western">
<surname>Gupta</surname>
<given-names>Deepa</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
<email>g_deepa@blr.amrita.edu</email>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western">
<surname>Zakariah</surname>
<given-names>Mohammed</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western">
<surname>Alotaibi</surname>
<given-names>Yousef Ajami</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<aff id="aff-1"><label>1</label><institution>Department of Electronics &#x0026; Communication Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham</institution>, <addr-line>Bengaluru</addr-line>, <country>India</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Science &#x0026; Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham</institution>, <addr-line>Bengaluru</addr-line>, <country>India</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Engineering, College of Computer and Information Sciences, King Saud University</institution>, <country>Saudi Arabia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1">&#x002A;Corresponding Author: Deepa Gupta. Email: <email>g_deepa@blr.amrita.edu</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2021-08-23">
<day>23</day>
<month>08</month>
<year>2021</year>
</pub-date>
<volume>69</volume>
<issue>3</issue>
<fpage>3217</fpage>
<lpage>3238</lpage>
<history>
<date date-type="received">
<day>07</day>
<month>3</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>4</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2021 Lalitha et al.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Lalitha et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_18406.pdf"></self-uri>
<abstract>
<p>Automatic recognition of human emotions in a continuous dialog model remains challenging where a speaker&#x2019;s utterance includes several sentences that may not always carry a single emotion. Limited work with standalone speech emotion recognition (SER) systems proposed for continuous speech only has been reported. In the recent decade, various effective SER systems have been proposed for discrete speech, <italic>i.e</italic>., short speech phrases. It would be more helpful if these systems could also recognize emotions from continuous speech. However, if these systems are applied directly to test emotions from continuous speech, emotion recognition performance would not be similar to that achieved for discrete speech due to the mismatch between training data (from training speech) and testing data (from continuous speech). The problem may possibly be resolved if an existing SER system for discrete speech is enhanced. Thus, in this work the author&#x2019;s existing effective SER system for multilingual and mixed-lingual discrete speech is enhanced by enriching the cepstral speech feature set with bi-spectral speech features and a unique functional set of Mel frequency cepstral coefficient features derived from a sine filter bank. Data augmentation is applied to combat skewness of the SER system toward certain emotions. Classification using random forest is performed. This enhanced SER system is used to predict emotions from continuous speech with a uniform segmentation method. Due to data scarcity, several audio samples of discrete speech from the SAVEE database that has recordings in a universal language, <italic>i.e</italic>., English, are concatenated resulting in multi-emotional speech samples. Anger, fear, sad, and neutral emotions, which are vital during the initial investigation of mentally disordered individuals, are selected to build six categories of multi-emotional samples. Experimental results demonstrate the suitability of the proposed method for recognizing emotions from continuous speech as well as from discrete speech.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Continuous speech</kwd>
<kwd>cepstral</kwd>
<kwd>bi-spectral</kwd>
<kwd>multi-emotional</kwd>
<kwd>discrete</kwd>
<kwd>emotion</kwd>
<kwd>filter bank</kwd>
<kwd>mental illness</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>A mental disorder, also called mental illness or psychiatric disorder [<xref ref-type="bibr" rid="ref-1">1</xref>], is a mental or behavioral pattern that causes significant impairment or distress in terms of personal functioning [<xref ref-type="bibr" rid="ref-2">2</xref>]. Mental disorders affect emotion, behavioral control, and cognition, and cause substantial interference in the learning ability of children as well as the functioning capability of adults at work and with their families. Mental disorders tend to originate at an early age and if not diagnosed and treated the individual suffers in a chronic recurrent manner [<xref ref-type="bibr" rid="ref-3">3</xref>]. The recent decade has witnessed a significant increase in the number of people suffering from mental illness {[<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>]}. Further, the COVID-19 pandemic has had an adverse effect on the mental health of people directly affected by the corona virus but also on their family members and friends as well as the general public {[<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]}. Thus, there exists an urgency to advance human mental health globally, resulting in a great demand for health care professionals for diagnosis and treatment [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>].</p>
<p>Patients suffering from different mental disorders typically experience certain specific emotions. Anxiety and fear are associated with individuals undergoing stress [<xref ref-type="bibr" rid="ref-14">14</xref>] and seasonal affective disorder [<xref ref-type="bibr" rid="ref-15">15</xref>]. Major depressive disordered [<xref ref-type="bibr" rid="ref-16">16</xref>] and mood disordered individuals [<xref ref-type="bibr" rid="ref-17">17</xref>] are prone to sadness and, in some cases, such individuals are always emotionally neutral and do not respond to situations that would typically cause an emotional response. Anger and fear are usually experienced by COVID-19 affected patients [<xref ref-type="bibr" rid="ref-18">18</xref>]. Borderline Personality Disorder (BPD) is a prevalent mental disorder that has an identifiable emotional component. It is reported that approximately 1.6% of the general population and 20% of the psychiatric population suffer from BPD [<xref ref-type="bibr" rid="ref-19">19</xref>]. Typically, BPD patients have rapid mood swings, tend to be emotionally unstable, and experience intense negative emotions (also referred to as affective dysregulation). People suffering with BPD do not feel the same emotion at all the times [<xref ref-type="bibr" rid="ref-20">20</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]. Apart from mental illness, individuals with medical issues, for example hormonal and heart related issues, experience fear, anger, or sad emotions [<xref ref-type="bibr" rid="ref-24">24</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>]. Thus, anger, fear, sad, and neutral emotions are indicators of mental disorders and other medical conditions. If these emotions could be predicted, then this would greatly help mental healthcare professionals during an initial investigation to diagnose the ailment.</p>
<p>Human emotions can be detected through speech, facial expressions, gestures, electroencephalography signals and autonomic nervous system signals. Amongst these modalities, recognition of emotions using speech is more popular for data collection and speech sample processing is more convenient. During the primary investigation of a mental illness, doctors spend time counseling patients [<xref ref-type="bibr" rid="ref-28">28</xref>]. In the process of continuous conversation, the sequence of emotions experienced by the patients are vital to understand the symptoms and the associated disorder. This situation would benefit from a speech-based automated system that can continuously detect the sequence of a patient&#x2019;s emotions during counseling. Such a system would help doctors identify the mental illness.</p>
<p>Speech-based automated systems have been developed for health care [<xref ref-type="bibr" rid="ref-29">29</xref>&#x2013;<xref ref-type="bibr" rid="ref-31">31</xref>]. These systems are equipped with emotional intelligence, causing mental health services to be further strengthened. Various automated systems that recognize emotions using text or multimodal analysis, <italic>i.e</italic>., a combination of text, images, and the linguistics of speech, have been designed [<xref ref-type="bibr" rid="ref-32">32</xref>&#x2013;<xref ref-type="bibr" rid="ref-34">34</xref>]. However, most existing automated speech emotion recognition (SER) systems are monolingual and can recognize emotions only from discrete speech. If these systems can be further enhanced to recognize emotions from continuous speech they could be more beneficial for doctors to diagnose patients with mental illness. Such a continuous SER system is proposed in this research work.</p>
<p>This remainder of this paper is organized as follows. Section 2 briefly review state of the art SER systems. Section 3 outlines the proposed approach and performance measures. Experiments are described and the results are discussed in Section 4. Conclusions and suggestions for future work are presented in Section 5.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>State-of-the-Art Models</title>
<p>A typical SER system processes and classifies various speech signals to recognize the embedded emotions. There exist several approaches to model emotions; however, categorical and dimensional models are most common [<xref ref-type="bibr" rid="ref-35">35</xref>&#x2013;<xref ref-type="bibr" rid="ref-38">38</xref>]. Categorical models deal with discrete human emotions experienced most commonly in day-to-day life. For example, Ekman proposed six basic human emotions, <italic>i.e</italic>., anger, disgust, fear, surprise, happiness, and sadness [<xref ref-type="bibr" rid="ref-39">39</xref>]. A dimensional model interprets discrete emotion in terms of valence and arousal dimension [<xref ref-type="bibr" rid="ref-40">40</xref>]. In the literature, SER based on dimensional models is referred to as continuous emotion recognition [<xref ref-type="bibr" rid="ref-41">41</xref>&#x2013;<xref ref-type="bibr" rid="ref-43">43</xref>]. In both categorical and dimensional SER models, emotion is recognized from a short duration phrase (2&#x2013;4 s) for monolingual, multilingual, cross-lingual, and mixed-lingual contexts [<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-45">45</xref>].</p>
<p>However, with conversational/continuous speech, speech data lasts for a longer duration, and the same emotion might not exist throughout the spoken utterance. Therefore, to deal with such situations, an SER system for continuous speech is essential. Few studies have investigated SER systems for continuous speech, and emotion databases with continuous speech are not available. Yeh et al. [<xref ref-type="bibr" rid="ref-46">46</xref>] investigated a continuous SER system using a segmentation-based approach to recognize emotions on continuous Mandarin emotion speech. Their study involved discrete emotion samples with categories of angry, happy, neutral, sad, and boredom. Multi-emotional samples with variable lengths were created by combining any two discrete emotion samples belonging to different categories, such as angry&#x2013;happy, neutral&#x2013;sad, and boredom&#x2013;happy, resulting in a total of 10 categories. Frame-based and voiced segmentation techniques were designed to evaluate the two emotions in each voice sample multi-emotional sample. A 128-feature set included jitter, shimmer, formants, linear predictive coefficients, linear prediction cepstral coefficients, Mel Frequency Cepstral Coefficients (MFCC) and MFCC derivatives, log frequency power coefficients, Perceptual Linear Prediction (PLP), and Rasta-PLP served as the speech features. Relevant features were extracted using sequential forward and sequential backward selection methods. A weighted discrete k-nearest neighbor classifier was considered that was trained using variable length utterances created from the database [<xref ref-type="bibr" rid="ref-46">46</xref>]. Fan et al. [<xref ref-type="bibr" rid="ref-47">47</xref>] investigated a multi-scaled time window for continuous SER. Their work involved recognizing two emotions from two classes of voice samples, <italic>i.e</italic>., angry&#x2013;neutral or happy&#x2013;neutral samples from the Emo-dB database and a Chinese database [<xref ref-type="bibr" rid="ref-47">47</xref>]. Various MFCC features, modulation spectral features, and global statistical features were employed in experiments. The LIBSVM library was applied for classification. The training data was combined and segmented uniformly to train the classifier. System performance was compared with a baseline Hidden Markov Model (HMM) system [<xref ref-type="bibr" rid="ref-48">48</xref>]. The best results were obtained using global statistical features.</p>
<p><bold><italic>Summary and Limitation of State of the Art Approaches:</italic></bold></p>
<p>From the survey conducted it is evident that various SER systems for monolingual, multilingual, cross-lingual and mixed-lingual discrete speech have been proposed in the past decade. However, few studies have considered continuous SER. In addition, the discrete speech SER systems used segmented continuous speech to train the classifier. Dedicated segmentation methods were incorporated to detect emotion variation boundaries in continuous speech using German and Chinese language voice samples.</p>
<p>It would be more practical and useful if a well-established SER system that works for discrete speech could also be applied for continuous speech, To a large extent, existing discrete SER systems may not be able to capture the sequence of emotions in continuous speech due to variation in emotion boundaries of training samples (derived from discrete speech) and test samples (derived from continuous speech). To address this, if some enhancements are incorporated in the prevailing SER systems for discrete speech, then continuous emotions could be better detected. Further, the SER systems should be robust for detecting emotions from a universal language, such as English, so that it can be versatile across the globe. Such an SER system is proposed in this article.</p>
<p>The primary contributions of this study are as follows.</p>
<p>a. Unique sine filter bank-based Mel-coefficient functionals are explored to recognize speech emotion.</p>
<p>b. A distinctive compact cepstral and bi-spectral feature combination is proposed for effective SER.</p>
<p>c. The proposed SER system efficiently recognizes emotions in continuous speech as well as discrete speech using a simple uniform segmentation technique.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Approach</title>
<p>The workflow of the implemented methodology is shown in <?A3B2 "fig1",5,"anchor"?><xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The principal constituent modules include database preparation, preprocessing, speech feature extraction, classification, and post-processing.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Database Preparation</title>
<p>Globally, the majority of people communicate in English. Hence, the SAVEE database, which contains recordings of utterances from four male native British English speakers was selected. The focus of this work is toward recognition of emotions from continuous speech of mentally disordered individuals during counseling. Thus, angry, neutral, sad, and fear emotions are considered. In the database used, the recordings comprised fifteen phonetically balanced sentences per emotion from the standard TIMIT corpus, with an additional 30 sentences for neutral emotion [<xref ref-type="bibr" rid="ref-49">49</xref>].</p>
<p><italic>Creation of Multi-Emotional Voice Samples</italic></p>
<p>Here, the focus is on continuous emotion detection. Due to the lack of available continuous speech emotion samples, a database needed to be created from available discrete emotion samples. In the database under consideration each sample includes a discrete emotion of 2&#x2013;4 s. In a practical situation, human emotions exist for a certain period. Thus, 3&#x2013;4 samples of the same emotion class are concatenated to form a voice sample of a single emotion category, as shown in <?A3B2 "fig2",5,"anchor"?><xref ref-type="fig" rid="fig-2">Figs. 2</xref> and <?A3B2 "fig3",5,"anchor"?><xref ref-type="fig" rid="fig-3">3</xref>. Two such voice samples from different emotion categories are concatenated to create a continuous multi-emotional speech sample, as depicted in <?A3B2 "fig4",5,"anchor"?><xref ref-type="fig" rid="fig-4">Fig. 4</xref>. Thus, in this work, continuous speech samples are multi-emotional with a duration of 7&#x2013;12 s. Five different categories of multi-emotional voice samples, <italic>i.e</italic>., angry&#x2013;neutral, sad&#x2013;angry, angry&#x2013;fear, sad&#x2013;neutral, and fear&#x2013;neutral are created using Audacity [<xref ref-type="bibr" rid="ref-50">50</xref>], which is an open-source audio editor and recording application software. Any two emotions from angry, neutral, sad, and fear are considered in multi-emotional speech creation as identification of these emotions are significant in any clinical investigation of an individual thought to be suffering from a mental disorder.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed continuous SER work overflow</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-1.png"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Preprocessing</title>
<p>This phase involves segmentation of continuous speech. As shown in <?A3B2 "fig5",5,"anchor"?><xref ref-type="fig" rid="fig-5">Fig. 5</xref>, the speech signal is segmented uniformly into segments of constant lengths (<italic>e.g</italic>., 2 s) and two consecutive frames make an independent speech sample. Framing is performed without overlapping. Then, the emotion of each segment can be recognized.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Creation of an angry utterance</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-2.png"/>
</fig>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Creation of a neutral utterance</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-3.png"/>
</fig>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Combining utterances</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-4.png"/>
</fig>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Uniform segmentation of continuous speech</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-5.png"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Speech Feature Extraction</title>
<p>In this study, the speech feature set includes cepstral features, bispectral features, and modified sine-based MFCC coefficients.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Cepstral Features</title>
<p>Unique cepstral speech feature functionals derived from Mel, Bark, and inverted Mel filter banks along with modified H-coefficients and additional parameters are found to be quite robust for multilingual and mixed-lingual SER for discrete samples from Indian and western language backgrounds [<xref ref-type="bibr" rid="ref-51">51</xref>]. The feature set in a previous study [<xref ref-type="bibr" rid="ref-51">51</xref>] form a size of 151 coefficients, as shown in <?A3B2 "tbl1",5,"anchor"?><xref ref-type="table" rid="table-1">Tab. 1</xref>, which are part of the speech feature set in this work.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Cepstral feature set [<xref ref-type="bibr" rid="ref-51">51</xref>]</title>
</caption>
<table>
<colgroup>
<col/>
<col charoff="140pt"></col>
<col/>
<col/>
<col charoff="140pt"></col>
<col/>
</colgroup>
<thead>
<tr>
<th>Sl. No.</th>
<th>Speech feature</th>
<th>{ Feature size}</th>
<th>Sl. No.</th>
<th>Speech feature</th>
<th>{ Feature size}</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>MFCC functionals</td>
<td>6</td>
<td>11</td>
<td>Extended IMFCC un-voiced functionals</td>
<td>6</td>
</tr>
<tr>
<td>2</td>
<td>MFCC voiced functionals</td>
<td>6</td>
<td>12</td>
<td>LPC functionals</td>
<td>6</td>
</tr>
<tr>
<td>3</td>
<td>MFCC unvoiced functionals</td>
<td>6</td>
<td>13</td>
<td>Functionals of MEDC, LFPC, LPCC</td>
<td>18</td>
</tr>
<tr>
<td>4</td>
<td>Extended MFCC functionals</td>
<td>6</td>
<td>14</td>
<td>PLPC functionals</td>
<td>6</td>
</tr>
<tr>
<td>5</td>
<td>Extended MFCC voiced functionals</td>
<td>6</td>
<td>15</td>
<td>MFPLPC functionals</td>
<td>6</td>
</tr>
<tr>
<td>6</td>
<td>Extended MFCC un-voiced functionals</td>
<td>6</td>
<td>16</td>
<td>BFCC functionals</td>
<td>6</td>
</tr>
<tr>
<td>7</td>
<td>IMFCC functionals</td>
<td>6</td>
<td>17</td>
<td>RPCC functionals</td>
<td>6</td>
</tr>
<tr>
<td>8</td>
<td>IMFCC voiced functionals</td>
<td>6</td>
<td>18</td>
<td>H-Coefficients</td>
<td>8</td>
</tr>
<tr>
<td>9</td>
<td>IMFCC un-voided functionals</td>
<td>6</td>
<td>19</td>
<td>Functionals of cceps and rceps, rceps_ph</td>
<td>18</td>
</tr>
<tr>
<td>10</td>
<td>Extended IMFCC voiced functionals</td>
<td>6</td>
<td>20</td>
<td>Skewness, kurtosis, variance, frequency, phase, average amplitude, max amplitude, maximum at pitch, pitch, entropy</td>
<td>11</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Bispectral Features</title>
<p>Fundamentally, a bispectrum is a Fourier transform of dimension two from the cumulant function of the third order, as shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.</p>
<p><disp-formula id="eqn-17">
<label>(1)</label>
<mml:math id="mml-eqn-17" display="block"><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>f</mml:mi><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>.</mml:mo></mml:math>
</disp-formula></p>
<p>Here, <italic>P(f</italic><sub>x</sub>, <italic>f</italic><sub>y</sub><italic>)</italic> denotes a bispectrum with frequencies (<italic>f</italic><sub>x</sub><italic>, f</italic><sub>y</sub>). X(f) represents Fourier transform, &#x002A; signifies complex conjugate, and E[.] means expectation of operation [<xref ref-type="bibr" rid="ref-52">52</xref>]. The bispectrum of a speech signal includes redundant data. Thus, bispectral features are selected from the non-redundant area <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, as shown in <?A3B2 "fig6",5,"anchor"?><xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Non-redundant area</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-6.png"/>
</fig>
<p>Frequencies represented in <xref ref-type="fig" rid="fig-6">Fig. 6</xref> are normalized by Nyquist frequency. <xref ref-type="disp-formula" rid="eqn-2">Eqs. (2)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-11">(11)</xref> illustrate the procedural steps to derive bispectral speech features. The mean magnitude of the bispectrum is expressed as follows:</p>
<p><disp-formula id="eqn-1">
<label>(2)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mtext>Mean Amp&#xA0;</mml:mtext><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2217;</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where p denotes the number of points prevailing in that region [<xref ref-type="bibr" rid="ref-53">53</xref>]. The weighted center of bispectrum (WCOB) is derived using <xref ref-type="disp-formula" rid="eqn-5">Eqs. (5)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-8">(8)</xref>.</p>
<p><disp-formula id="eqn-2">
<label>(3)</label>
<mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>c</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-3">
<label>(4)</label>
<mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>d</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-4">
<label>(5)</label>
<mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>3</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>c</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-5">
<label>(6)</label>
<mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>4</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>d</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math>
</disp-formula></p>
<p>Here, c and d provide the bin index of the frequency existing in the region, where <italic>g</italic><sub><italic>1d</italic></sub>, <italic>g</italic><sub><italic>2d</italic></sub> represents WCOB and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>3</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mn>4</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are WCOB absolute values [<xref ref-type="bibr" rid="ref-54">54</xref>].</p>
<p>The log amplitude summation (T<sub>a</sub>) of the bispectrum is derived as follows.</p>
<p><disp-formula id="eqn-6">
<label>(7)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="normal">T</mml:mi><mml:mi mathvariant="normal">a</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>y</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math>
</disp-formula></p>
<p>Similarly, the log amplitude summation from diagonal elements (<italic>T</italic><sub><italic>b</italic></sub>) in the bispectrum derived as follows</p>
<p><disp-formula id="eqn-7">
<label>(8)</label>
<mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>T</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math>
</disp-formula></p>
<p>The amplitude of diagonal elements (<italic>T</italic><sub><italic>c</italic></sub>, <italic>T</italic><sub><italic>d</italic></sub>, <italic>T</italic><sub><italic>e</italic></sub>) with first and second order spectral moments is derived by:</p>
<p><disp-formula id="eqn-8">
<label>(9)</label>
<mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>T</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mi>d</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-9">
<label>(10)</label>
<mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>T</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>T</mml:mi><mml:mi>c</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2217;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>f</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-10">
<label>(11)</label>
<mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>T</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:msqrt><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math>
</disp-formula></p>
<p>A total of six features comprising the bispectrum mean amplitude and five features of bispectrum log amplitudes are derived and form the part of the proposed speech feature set.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Modified Sine-Based MFCC Coefficients</title>
<p>The process flow for extraction of sine-based Mel coefficients is shown in <?A3B2 "fig7",5,"anchor"?><xref ref-type="fig" rid="fig-7">Fig. 7</xref>. Initially, the power spectra of the preprocessed speech signal is derived. Differing from the conventional triangular shaped filter bank for MFCC feature extraction as discussed in an earlier SER study [<xref ref-type="bibr" rid="ref-55">55</xref>], here sinusoidal filter banks, as shown in <?A3B2 "fig8",5,"anchor"?><xref ref-type="fig" rid="fig-8">Fig. 8</xref>, are applied to the power spectra. The center frequencies of the filter banks are given as <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>.</p>
<p><disp-formula id="eqn-11">
<label>(12)</label>
<mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mi mathvariant="normal">f</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mi>N</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mfrac><mml:mrow><mml:mi>B</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mn>2</mml:mn></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>;</mml:mo><mml:mspace width="1em" /><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>F</mml:mi></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is defined in <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref> as follows:</p>
<p><disp-formula id="eqn-12">
<label>(13)</label>
<mml:math id="mml-eqn-12" display="block"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>700</mml:mn><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mi>m</mml:mi><mml:mn>2595</mml:mn></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p>Here, f(p) denotes center frequency, f<sub>s</sub> represents sampling frequency, and N is the window length.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Sine-based MFCC feature extraction process</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-7.png"/>
</fig>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Sine-based filter bank</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-8.png"/>
</fig>
<p>Lastly, the successive application of log and discrete cosine transform to the output of the sine-based Mel filter bank results in deriving the modified MFCC coefficients, <italic>i.e</italic>., sine-based MFCC coefficients. Functionals of the modified MFCC, <italic>i.e</italic>., maximum, minimum, mean, standard deviation, variance, and median are considered.</p>
<p>In this study, for each speech signal, 151 cepstral features, six bispectral features, and six sine-based MFCC functionals (<italic>i.e</italic>., 163 coefficients) are extracted.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Classification and Post Processing</title>
<p>For the proposed SER work, various classifiers from Python [<xref ref-type="bibr" rid="ref-56">56</xref>] were chosen. However, compared with other classifiers, superior performance was achieved with the random forest (RF) classifier. Therefore, the RF classifier was hence considered in this work [<xref ref-type="bibr" rid="ref-57">57</xref>]. With the knowledge acquired by the classifier during training from feature vectors of discrete samples referred as learning from discrete SER Model, emotion is predicted for each continuous speech segment. The feature vector is comprised of cepstral, bi-spectral, and sine filter bank-based MFCC functionals. In the post processing phase, a decision rule is deployed to determine the sequence of emotions. For every consecutive three speech segments, the emotion predicted the maximum number of times is the emotion determined. These predicted emotions are sequences of emotions in the continuous speech.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Evaluation Metrics</title>
<p>In this study, performance measures of recall, precision, F-measure, and accuracy are considered to evaluate the system [<xref ref-type="bibr" rid="ref-58">58</xref>].</p>
<sec id="s3_5_1">
<label>3.5.1</label>
<title>Recall</title>
<p>Recall <italic>is</italic> the number of instances that are relevant among the total number of relevant instances. Recall is also known as sensitivity.</p>
<p><disp-formula id="eqn-13">
<label>(14)</label>
<mml:math id="mml-eqn-13" display="block">
 <mml:mrow>
  <mml:mtext>Recall&#x00A0;</mml:mtext><mml:mo>&#x00A0;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x0025;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow>
  </mml:mfrac>
  <mml:mo>*</mml:mo><mml:mn>100</mml:mn><mml:mo>,</mml:mo></mml:mrow>
</mml:math>
</disp-formula></p>
<p>where, True Positive is number of samples predicted positive that are actually positive, and False Negative is the number of examples predicted negative that are actually negative.</p>
</sec>
<sec id="s3_5_2">
<label>3.5.2</label>
<title>Precision</title>
<p>Precision gives the number of instances that are relevant among the instances retrieved. Precision is also known as the positive predictive value.</p>
<p><disp-formula id="eqn-14">
<label>(15)</label>
<mml:math id="mml-eqn-14" display="block">
 <mml:mrow>
  <mml:mtext>Precision&#x00A0;</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x0025;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow>
  </mml:mfrac>
  <mml:mo>*</mml:mo><mml:mn>100.</mml:mn></mml:mrow>
</mml:math>
</disp-formula></p>
<p>Precision quantifies the number of correct positive predictions.</p>
</sec>
<sec id="s3_5_3">
<label>3.5.3</label>
<title>F-Measure</title>
<p>F-measure is the harmonic mean of recall and precision.</p>
<p><disp-formula id="eqn-15">
<label>(16)</label>
<mml:math id="mml-eqn-15" display="block">
 <mml:mrow>
  <mml:mtext>F-measure&#x00A0;</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x0025;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>*</mml:mo><mml:mfrac>
   <mml:mrow>
    <mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>*</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow>
   <mml:mrow>
    <mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow>
  </mml:mfrac>
  <mml:mo>*</mml:mo><mml:mn>100.</mml:mn></mml:mrow>
</mml:math>
</disp-formula></p>
</sec>
<sec id="s3_5_4">
<label>3.5.4</label>
<title>Accuracy</title>
<p>Accuracy is the number of test samples of a particular emotion classified accurately with respect to the total number of test samples of the emotion under consideration.</p>
<p><disp-formula id="eqn-16">
<label>(17)</label>
<mml:math id="mml-eqn-16" display="block">
 <mml:mrow>
  <mml:mtext>Accuracy&#x00A0;</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x0025;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac>
   <mml:mrow>
    <mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>l</mml:mi><mml:mi>y</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mi>i</mml:mi><mml:mi>z</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow>
   <mml:mrow>
    <mml:mi>T</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x00A0;</mml:mo><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow>
  </mml:mfrac>
  <mml:mo>*</mml:mo><mml:mn>100.</mml:mn></mml:mrow>
</mml:math>
</disp-formula></p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Work, Results, and Discussion</title>
<p>The experimental work is performed in two successive modules. The focus of the first module involves enhancing the author&#x2019;s previously proposed SER system on discrete speech [<xref ref-type="bibr" rid="ref-51">51</xref>]. This is required because, although the existing SER system is suitable for recognizing emotions from various discrete speech languages, when continuous speech is tested, the performance is not similar. Therefore, the initial experimental work involves increasing the robustness of this existing SER system for discrete speech with the addition of a few more important speech features so that emotions could also be detected from continuous speech. The enhanced SER system is referred to as the proposed SER system. The second module involves experimentation on continuous speech using the proposed SER system. Both modules involve extraction of cepstral, bi-spectral and sine-based MFCC functionals speech features. An RF classifier is used. Fivefold cross-validation is applied to analyze system performance.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Module 1: Experimentation and Analysis for Proposed SER System</title>
<p>The previously proposed SER system for multilingual and mixed-lingual discrete speech [<xref ref-type="bibr" rid="ref-51">51</xref>] is considered in this work. The previous system comprised cepstral speech feature functionals of size 151 coefficients for each speech sample and used a simple RF classifier. Data augmentation was applied to avoid system bias toward any specific set of emotion categories. The current study focuses on recognizing emotions that are indicators of mental illness, <italic>i.e</italic>., angry, sad, fear, and neutral emotions. Thus, the initial phase of the work involved investigating the performance of the previous SER system [<xref ref-type="bibr" rid="ref-51">51</xref>] in recognizing these four emotions from discrete samples from the SAVEE database. The results obtained by this investigation are shown in <?A3B2 "tbl2",5,"anchor"?><xref ref-type="table" rid="table-2">Tab. 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Previous SER system performance using cepstral features</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Emotion</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>F-Measure (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Angry</td>
<td>90.6</td>
<td>96.7</td>
<td>93.5</td>
</tr>
<tr>
<td>Fear</td>
<td>77.8</td>
<td>70.0</td>
<td>73.7</td>
</tr>
<tr>
<td>Neutral</td>
<td>80.3</td>
<td>95.0</td>
<td>87.0</td>
</tr>
<tr>
<td>Sad</td>
<td>80.0</td>
<td>82.7</td>
<td style="background:#FFFFFF;">81.3</td>
</tr>
<tr>
<td>Weighted average</td>
<td>82.2</td>
<td>82.1</td>
<td>83.9</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="table" rid="table-2">Tab. 2</xref>, it can be observed that, among the four indicative emotions, the system best recognizes angry and neutral emotions with recall rates of 96.7% and 95.0%, respectively. Precision and F-Scores rates are also reported to be above 80.0% for the aforementioned emotions. The min-max rates achieved are recall 70.0%&#x2013;96.7%, precision 77.8%&#x2013;90.6%, and F-Score 73.7%&#x2013;93.5%. In addition, weighted averages of approximately 82.0% are obtained across all performance measures. Samples from sad emotions are misclassified as neutral while fear is primarily classified as angry. Thus, considerably lower rates are reported for sad and fear emotions. The previously proposed system has to be made more robust in recognizing fear and sad emotions along with angry and neutral emotions, such that emotions in continuous speech can be well detected.</p>
<p>One probable solution the authors considered to overcome this limitation was to expand the existing speech feature set and enhance system performance for fear and sad emotions. Thus, in this work, the cepstral feature set used in the previous system [<xref ref-type="bibr" rid="ref-51">51</xref>] is enhanced using bi-spectral features that capture the higher order statistics of the signal spectra. The experimental work now involves extracting cepstral&#x2013;bi-spectral feature combinations and analyzing the SER system. Thus, a speech feature set of 157 coefficients (151 cepstral features and 6 bi-spectral features) for each speech sample was extracted from all the audio samples of the SAVEE database. The feature set derived from the speech samples were subjected to an emotion recognition task. The SER system performance is shown in <?A3B2 "tbl3",5,"anchor"?><xref ref-type="table" rid="table-3">Tab. 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Performance of the discrete SER system using cepstral and bi-cepstral features</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Emotion</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>F-Measure (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Angry</td>
<td>95.2</td>
<td>100.0</td>
<td>97.6</td>
</tr>
<tr>
<td>Fear</td>
<td>100.0</td>
<td>86.7</td>
<td>97.5</td>
</tr>
<tr>
<td>Neutral</td>
<td>86.4</td>
<td>95.0</td>
<td>90.5</td>
</tr>
<tr>
<td>Sad</td>
<td>87.2</td>
<td>88.3</td>
<td>87.7</td>
</tr>
<tr>
<td>Weighted average</td>
<td>91.0</td>
<td>92.5</td>
<td>93.3</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From the results shown in <xref ref-type="table" rid="table-3">Tab. 3</xref>, it is evident that the higher-order statistics of the bi-spectral features along with the cepstral features are significant for emotion recognition, and all emotions show performance measures greater than 85.0%. Fear and sad emotions show an increased recall rate of approximately 16% and 5%, respectively, compared with the results shown in <xref ref-type="table" rid="table-2">Tab. 2</xref>. The min-max rates achieved are recall 86.7%&#x2013;100.0%, precision 86.4%&#x2013;100%, and {F-Score} 87.7%&#x2013;97.6%. The min rates across the three measures, which were previously less than 80%, have improved and remained above 85%. In addition, with the inclusion of bi-spectral features, weighted averages were approximately 92.0%, which is 10% higher than those reported in <xref ref-type="table" rid="table-2">Tab. 2</xref>, where only cepstral features were considered, across all performance measures. Note that, although SER performance has improved, some errors persist, <italic>i.e</italic>., sad is recognized as neutral and fear is recognized as angry. Thus, the recall rates for sad and fear emotions were between 80.0%&#x2013;90.0%.</p>
<p>To overcome this and further enhance the emotion prediction of the SER system, the speech feature set is further expanded. For this purpose, the authors focused on altering the filter bank shape used to derive cepstral features. With an initial work in this direction, the authors considered altering one of the cepstral feature filter bank shapes proposed in <xref ref-type="table" rid="table-1">Tab. 1</xref>. Among this set, MFCC has been a popular feature for various speech applications, including emotion recognition [<xref ref-type="bibr" rid="ref-59">59</xref>]. Thus, in this study, the filter bank shape of MFCC is altered. Traditionally to date, triangular filter banks have been used for MFCC feature extraction. In this work, sine-shaped filter banks have been considered, and MFCC features are derived. The extraction procedure is discussed in Section 3.</p>
<p>Six functionals of the Mel coefficients are derived from the sine filter bank and appended to the feature vector of the cepstral and bispectral feature combination. This resulted in a size of 163 coefficients for each speech sample. Classification was performed and the robustness of this feature combination is analyzed. The results obtained are shown in <?A3B2 "tbl5",5,"anchor"?><xref ref-type="table" rid="table-5">Tab. 5</xref>. With the incorporation of the new speech feature, all four emotions are optimally recognized with performance rates greater than 95% for all measures. This indicates that the shape of the filter bank has a considerable effect on the extracted Mel coefficients and hence on the emotion discriminating capability. The average accuracy of all three performance measures was 97.9%. The min-max band for recall was 95.8%&#x2013;100%, precision was 96.6%&#x2013;99.2%, and the F-Score was 96.2%&#x2013;99.6%.</p>
<p>From an analysis of the results presented in <?A3B2 "tbl4",5,"anchor"?><xref ref-type="table" rid="table-4">Tab. 4</xref>, the previous SER system [<xref ref-type="bibr" rid="ref-51">51</xref>] is enhanced with the inclusion of bi-spectral and sine filter bank-based MFCC coefficients. This enhanced system has proven to be robust in recognizing all the four emotions of discrete speech and henceforth is referred to as the proposed SER system.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance of the SER model using cepstral, bi-cepstral, and modified sine-based MFCC features</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Emotion</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>F-Measure (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Angry</td>
<td>99.2</td>
<td>100.0</td>
<td>99.6</td>
</tr>
<tr>
<td>Fear</td>
<td>99.2</td>
<td>99.2</td>
<td>99.2</td>
</tr>
<tr>
<td>Neutral</td>
<td>96.6</td>
<td>95.8</td>
<td>96.2</td>
</tr>
<tr>
<td>Sad</td>
<td>96.7</td>
<td>96.7</td>
<td>96.7</td>
</tr>
<tr>
<td>Weighted average</td>
<td>97.9</td>
<td>97.9</td>
<td>97.9</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Module 2: Experimentation and Analysis of Proposed SER System for Continuous Speech</title>
<p>In this module, experiments are conducted with regard to the recognition of emotions from continuous multi-emotional speech samples using the proposed SER system. The detailed workflow is explained in Section 3. Emotions of either angry&#x2013;neutral, fear&#x2013;neutral, sad&#x2013;neutral, angry&#x2013;fear, angry&#x2013;sad, or fear&#x2013;sad are included in each continuous speech sample during data creation. The feature vector of the discrete speech samples were input to train the classifier that was subsequently applied to recognize emotions in continuous speech. In this experimental procedure, for example, consider the context with recognizing emotions of an angry&#x2013;neutral speech, all the samples created of this category is divided into five different folds. When continuous speech samples of a particular fold of angry&#x2013;neutral is tested, all those discrete samples of angry and neutral used for data creation in that fold are removed from the discrete input to the training phase to avoid the bias during testing. The same is repeated during testing the remaining five categories of continuous speech samples.</p>
<p>In this context of experimentation, a multi-emotional sample of angry&#x2013;neutral, as shown in <?A3B2 "fig9",5,"anchor"?><xref ref-type="fig" rid="fig-9">Fig. 9</xref>, was tested using the proposed SER system. For each segment, the system recognizes the associated emotion. The angry emotion is denoted A, and the neutral emotion is denoted N. The decision rule was as follows: for every three consecutive segments the maximally recognized emotion is considered to be emotion. Finally, all these emotions are emotions in the continuous speech. As observed in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, Angry-Angry-Angry&#x2013;Neutral are the emotions of the speech sample tested.</p>
<p>All continuous emotion samples were tested, and the obtained results were analyzed. First, the performance of the proposed SER system across each fold during the fivefold cross-validation was investigated. The bar charts in <?A3B2 "fig10",5,"anchor"?><xref ref-type="fig" rid="fig-10">Figs. 10a</xref>&#x2013;<xref ref-type="fig" rid="fig-10">10f</xref> depict how each emotion paired with another emotion in a speech sample of each fold is recognized in continuous speech using the proposed method. Every fold consists of eight test samples across any continuous emotion category. From the plots, it is observed that both emotions in Angry&#x2013;Neutral and Angry&#x2013;Sad are consistently recognized from the continuous speech across all folds. However, recognition of Sad in the Sad&#x2013;Neutral combination shows a large variation across the folds. Fear in the neutral or angry combination and sad combined with neutral show large variations in recognized emotions across the folds. In addition, both emotions in the Fear&#x2013;Sad combination remained consistent across the folds; however, fear is confused with angry, and sad is confused with neutral, resulting in a lower recognition performance. With the investigation of emotions recognized across folds, the next step involved overall performance analysis, as illustrated in <xref ref-type="table" rid="table-5">Tab. 5</xref>.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Recognized emotions in the continuous speech sample tested</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-9.png"/>
</fig>
<p>The performance of the proposed SER system across each of these six different multi-emotional pairs in continuous speech is shown in <xref ref-type="table" rid="table-5">Tab. 5</xref>. Emotions in continuous speech are better recognized from the Angry&#x2013;Neutral emotion pair, with an accuracy of 85.0% and 97.0% for angry and neutral, respectively. Angry is better recognized with accuracy of at least 80.0% and higher with any of the continuous speech sample. Fear emotion is often confused with angry and is moderately recognized with Fear&#x2013;Neutral and Angry&#x2013;Fear scenarios. Sad is primarily classified as neutral, resulting in lower recognition performance, as observed in the Sad&#x2013;Neutral combination. Thus, recognition of sad remains challenging when associated with neutral emotion. The emotions from the fear&#x2013;sad continuous emotion category was found to be considerably lower, <italic>i.e</italic>., 54.6% for fear and 45.8% for sad emotion.</p>
<p>An analysis of the accuracy performance across the six multi-emotional categories considered in this work is shown in <?A3B2 "fig11",5,"anchor"?><xref ref-type="fig" rid="fig-11">Fig. 11</xref>. Considerable accuracy recognition rates higher than 75.0% are guaranteed for any continuous emotion category. The Angry&#x2013;Neutral emotion pair demonstrated superior recognition rates, with an average accuracy of 91.0%. With the exception of the fear&#x2013;sad combination, the min-max average accuracy band was 71.3%&#x2013;91.0%.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Accuracy (%) performance of proposed SER system for continuous speech across each fold during fivefold cross-validation (a) Angry&#x2013;Neutral continuous speech (b) Angry&#x2013;Fear continuous speech (c) Angry&#x2013;Sad continuous speech (d) Fear&#x2013;Neutral continuous speech (e) Sad&#x2013;Neutral continuous speech (f) Fear&#x2013;Sad continuous speech</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-10.png"/>
</fig>
<p>As shown in <xref ref-type="table" rid="table-5">Tab. 5</xref>, similar performance results were obtained when the first and second emotion samples of the multi-emotional combination were interchanged. For example, Angry&#x2013;Neutral and Neutral&#x2013;Angry continuous samples were recognized with the same accuracy by the proposed SER system.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Accuracy (%) of the proposed SER system for continuous speech</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Continuous speech emotion category</th>
<th>Emotion</th>
<th>Accuracy (%) across each emotion</th>
</tr>
</thead>
<tbody>
<tr>
<td>Angry-Neutral</td>
<td>Angry</td>
<td>85.0</td>
</tr>
<tr>
<td></td>
<td>Neutral</td>
<td>97.0</td>
</tr>
<tr>
<td>Fear-Neutral</td>
<td>Fear</td>
<td>61.0</td>
</tr>
<tr>
<td></td>
<td>Neutral</td>
<td>92.5</td>
</tr>
<tr>
<td>Sad-Neutral</td>
<td>Sad</td>
<td>47.5</td>
</tr>
<tr>
<td></td>
<td>Neutral</td>
<td>95.0</td>
</tr>
<tr>
<td>Angry-Fear</td>
<td>Angry</td>
<td>100.0</td>
</tr>
<tr>
<td></td>
<td>Fear</td>
<td>65.0</td>
</tr>
<tr>
<td>Angry-Sad</td>
<td>Angry</td>
<td>80.0</td>
</tr>
<tr>
<td></td>
<td>Sad</td>
<td>67.5</td>
</tr>
<tr>
<td>Fear-Sad</td>
<td>Fear</td>
<td>54.6</td>
</tr>
<tr>
<td></td>
<td>Sad</td>
<td>45.8</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Average accuracy (%) of SER system for various continuous emotion combinations</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_18406-fig-11.png"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Comparative Analysis of the Proposed System with Existing Continuous SER Studies</title>
<p>In this section, the proposed SER system is compared with existing works for recognition of emotions in continuous speech. As shown in <?A3B2 "tbl6",5,"anchor"?><xref ref-type="table" rid="table-6">Tab. 6</xref>, two similar studies have been identified. Though each work is validated on a different database, the comparison is performed to compare the robustness of the proposed methodology to existing techniques. The proposed work involved uniform segmentation independent of the detection of emotion variation boundaries, as performed in existing studies [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>]. The proposed SER system showed considerable average recognition accuracy of 74.2% using a unique cepstral feature functional set of size 163. Four discrete emotions with six multi-emotional categories were involved. However, in the work of Yeh et al. [<xref ref-type="bibr" rid="ref-46">46</xref>], although five discrete emotions with 10 different multi-emotional categories using a Mandarin database are involved, with uniform segmentation, the model could achieve only 40% accuracy. Applying end point detection in the segmentation method and a feature selection method, 89.0% was achieved. Similarly, in the study conducted by Fan et al. [<xref ref-type="bibr" rid="ref-47">47</xref>], three discrete emotions with only two multi-emotional categories were involved. Though only small feature set of size 85 was applied, a multi-time scale window was applied during the segmentation stage, with an additional task of training and testing samples were chosen to be of the same length. Although both studies [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>] demonstrated accuracy of approximately 89.0%, they only considered continuous speech, and validation was performed on Mandarin (Chinese language voice samples) and Emo-dB (German language voice samples) databases where the speakers emotion voice recordings are highly expressive. Note that Mandarin and German are not universal languages. In contrast, the SAVEE database is considered in this work. The SAVEE database contained voice samples from male speakers in English, which is a universal language. In addition, the emotions are very flat and not expressive. The proposed SER system exhibits considerable emotion recognition for both discrete and continuous speech, proving the robustness of the emotion carrying capability of the chosen speech feature combination, which avoids detection of emotion variation boundaries, feature selection techniques, and the use of segmented continuous speech during training for continuous emotion recognition.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparative analysis of the proposed continuous SER with previous studies</title>
</caption>
<table>
<colgroup>
<col charoff="55pt"></col>
<col charoff="115pt"></col>
<col/>
<col charoff="150pt"></col>
<col charoff="68pt"></col>
<col/>
</colgroup>
<thead>
<tr>
<th>{ Author &#x0026; year [Ref.]}</th>
<th>{ Database/number of discrete emotions in multi-emotional voice sample/discrete emotions}</th>
<th>{ Multi-emotion sample categories}</th>
<th>{ Methodology [Segmentation/ features/feature vector size/ feature selection/classifier]}</th>
<th>{ Average accuracy (%)}</th>
<th></th>
</tr>
</thead>
<tbody>
<tr>
<td>Yeh et al. 2011 [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>Mandarin/5/Angry, neutral, sad, happy, boredom</td>
<td>10</td>
<td>Uniform segmentation or segmentation using end point detection/Cepstral and voice quality features/128/Sequential backward selection/Weighted discrete k-nearest neighbor</td>
<td>40.0 (uniform segmentation), 89.0 (end point detection)</td>
<td></td>
</tr>
<tr>
<td>Fan et al. 2014 [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>EMO-dB/3/Angry, neutral, happy</td>
<td>2</td>
<td>Multi-time scaled window/global statistical features/85/-/Neural network</td>
<td>89.0</td>
<td></td>
</tr>
<tr>
<td>Proposed method</td>
<td>SAVEE/4/Angry, neutral, sad, fear</td>
<td>6</td>
<td>Uniform Segmentation/Cepstral&#x2013;bispectral feature set with data augmentation/163/-/Random forest</td>
<td>74.2</td>
<td></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion and Future Research</title>
<p>This study focused on the recognition of human emotions in continuous speech in a mental health context. In this study, an existing SER system for discrete speech that is quite robust for multilingual and mixed-lingual contexts is enhanced to capture emotion variations in continuous speech. It was demonstrated that altering the filter bank shape during MFCC extraction was effective in improving SER. Sine filter bank-based Mel cepstral coefficients and a cepstral&#x2013;bi-spectral feature set proved to be capable of recognizing emotions from continuous speech. In addition, uniform segmentation is considered. The proposed system is independent of any dedicated segmentation techniques and feature selection algorithms. Differing from existing SER systems, the proposed system is well suited for recognizing continuous emotions in continuous speech besides discrete speech. Thus, the proposed SER system is suitable to be deployed in bots for effective mental disorder investigations.</p>
<p>The proposed SER system recognizes emotions from continuous English speech. Since this system is an enhanced version an existing system that is suitable for multilingual and mixed-lingual contexts, emotions from continuous speech of other languages should also be better recognized. Therefore, in future, the performance of the proposed system for continuous speech of other languages could be tested. This study is intended toward recognizing mental illness based on emotional content of speech. Therefore, real time audio recorded during counseling sessions with mental illness patients could be used to test the proposed system. In this study, two emotions are included in each multi-emotional voice sample. Future research could include more emotion categories. In addition, features could be added to the existing feature set so that the sad emotion could be better recognized in the presence of neutral or fear emotion in continuous speech. More significantly, cepstral features could be derived from different filter bank shapes.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> This work was partially supported by the Research Groups Program (Research Group Number RG-1439-033), under the Deanship of Scientific Research, King Saud University, Riyadh, Saudi Arabia.</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Oliver</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Thayer</surname></string-name></person-group>, &#x201C;<article-title>Mental health disorders</article-title>,&#x201D; <source>British Dental Journal</source>, vol. <volume>227</volume>, no. <issue>7</issue>, pp. <fpage>539</fpage>&#x2013;<lpage>540</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-2"><label>2</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Risal</surname></string-name></person-group>, &#x201C;<article-title>Common mental disorders</article-title>,&#x201D; <source>Kathmandu University Medical Journal</source>, vol. <volume>9</volume>, no. <issue>3</issue>, pp. <fpage>213</fpage>&#x2013;<lpage>217</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-3"><label>3</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>O. F.</given-names> <surname>Norheim</surname></string-name></person-group>, &#x201C;<article-title>Disease control priorities third edition is published: A theory of change is needed for translating evidence to health policy</article-title>,&#x201D; <source>International Journal of Health Policy and Management</source>, vol. <volume>7</volume>, no. <issue>9</issue>, pp. <fpage>771</fpage>&#x2013;<lpage>777</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-4"><label>4</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Chancellor</surname></string-name> and <string-name><given-names>M. D.</given-names> <surname>Choudhury</surname></string-name></person-group>, &#x201C;<article-title>Methods in predictive techniques for mental health status on social media: A critical review</article-title>,&#x201D; <source>NPJ Digital Medicine</source>, vol. <volume>3</volume>, no. <issue>43</issue>, pp. <fpage>1</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-5"><label>5</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>C. H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Su</surname></string-name> and <string-name><given-names>Y. T.</given-names> <surname>Kuo</surname></string-name></person-group>, &#x201C;<article-title>Detecting unipolar and bipolar depressive disorders from elicited speech responses using latent affective structure model</article-title>,&#x201D; <source>IEEE Transactions on Affective Computing</source>, vol. <volume>11</volume>, no. <issue>3</issue>, pp. <fpage>393</fpage>&#x2013;<lpage>404</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-6"><label>6</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Stefanidou</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Greenlaw</surname></string-name> and <string-name><given-names>L. M.</given-names> <surname>Douglass</surname></string-name></person-group>, &#x201C;<article-title>Mental health issues in transition-age adolescents and young adults with epilepsy</article-title>,&#x201D; <source>Seminars in Pediatric Neurology</source>, vol. <volume>36</volume>, pp. <fpage>100856</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-7"><label>7</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Morris</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Edjoc</surname></string-name></person-group>, &#x201C;<article-title>Health fact sheet: Paraben concentrations in canadians, 2014 and 2015</article-title>,&#x201D; <source>Daily (Statistics Canada)</source>, vol. <volume>114</volume>, no. <issue>12</issue>, pp. <fpage>82</fpage>&#x2013;<lpage>625</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-8"><label>8</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I. I.</given-names> <surname>Haider</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Tiwana</surname></string-name> and <string-name><given-names>S. M.</given-names> <surname>Tahir</surname></string-name></person-group>, &#x201C;<article-title>Impact of the COVID-19 pandemic on adult mental health</article-title>,&#x201D; <source>Pakistan Journal of Medical Sciences</source>, vol. <volume>36</volume>, no. <issue>COVID19-S4</issue>, pp. <fpage>90</fpage>&#x2013;<lpage>94</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-9"><label>9</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. L.</given-names> <surname>Hagerty</surname></string-name> and <string-name><given-names>L. M.</given-names> <surname>Williams</surname></string-name></person-group>, &#x201C;<article-title>The impact of COVID-19 on mental health: The interactive roles of brain biotypes and human connection</article-title>,&#x201D; <source>Brain Behavior &#x0026; Immunity-Health</source>, vol. <volume>5</volume>, no. <issue>9</issue>, pp. <fpage>100078</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-10"><label>10</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Ouyang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Huo</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Xia</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Shan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Dual-sampling attention network for diagnosis of COVID-19 from community acquired pneumonia</article-title>,&#x201D; <source>IEEE Transactions on Medical Imaging</source>, vol. <volume>39</volume>, no. <issue>8</issue>, pp. <fpage>2595</fpage>&#x2013;<lpage>2605</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. C.</given-names> <surname>Das</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Roy</surname></string-name> and <string-name><given-names>M. S. I.</given-names> <surname>Salam</surname></string-name></person-group>, &#x201C;<article-title>Potential factors of mental health challenges during COVID-19 on the young people in Dhaka, Bangladesh</article-title>,&#x201D; <source>Advanced Journal of Social Science</source>, vol. <volume>7</volume>, no. <issue>1</issue>, pp. <fpage>109</fpage>&#x2013;<lpage>117</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-12"><label>12</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. B.</given-names> <surname>Turkozer</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Ongur</surname></string-name></person-group>, &#x201C;<article-title>A projection for psychiatry in the post-COVID-19 era: Potential trends, challenges, and directions</article-title>,&#x201D; <source>Molecular Psychiatry</source>, vol. <volume>25</volume>, no. <issue>10</issue>, pp. <fpage>2214</fpage>&#x2013;<lpage>2219</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-13"><label>13</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Samartzis</surname></string-name> and <string-name><given-names>M. A.</given-names> <surname>Talias</surname></string-name></person-group>, &#x201C;<article-title>Assessing and improving the quality in mental health services</article-title>,&#x201D; <source>International Journal of Environmental Research and Public Health</source>, vol. <volume>17</volume>, no. <issue>1</issue>, pp. <fpage>249</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-14"><label>14</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. R.</given-names> <surname>Bandela</surname></string-name> and <string-name><given-names>T. K.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Stressed speech emotion recognition using feature fusion of teager energy operator and MFCC</article-title>,&#x201D; in <conf-name>8th Int. Conf. on Computing, Communication and Networking Technologies</conf-name>, <publisher-loc>Delhi, India</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-15"><label>15</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Melrose</surname></string-name></person-group>, &#x201C;<article-title>Seasonal affective disorder: An overview of assessment and treatment approaches</article-title>,&#x201D; <source>Depression Research and Treatment</source>, vol. <volume>2015</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-16"><label>16</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Mahendran</surname></string-name>, <string-name><given-names>P. M. D. R.</given-names> <surname>Vincent</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Srinivasan</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Sharma</surname></string-name> and <string-name><given-names>D. K.</given-names> <surname>Jayakody</surname></string-name></person-group>, &#x201C;<article-title>Realizing a stacking generalization model to improve the prediction accuracy of major depressive disorder in adults</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>49509</fpage>&#x2013;<lpage>49522</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-17"><label>17</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Deligianni</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Guo</surname></string-name> and <string-name><given-names>G. Z.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>From emotions to mood disorders: A survey on gait analysis methodology</article-title>,&#x201D; <source>IEEE Journal of Biomedical and Health Informatics</source>, vol. <volume>23</volume>, no. <issue>6</issue>, pp. <fpage>2302</fpage>&#x2013;<lpage>2316</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-18"><label>18</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Kene</surname></string-name></person-group>, &#x201C;<article-title>Mental health implications of the COVID-19 pandemic in India</article-title>,&#x201D; <source>Psychological Trauma: Theory, Research, Practice and Policy</source>, vol. <volume>12</volume>, pp. <fpage>585</fpage>&#x2013;<lpage>587</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-19"><label>19</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W. D.</given-names> <surname>Ellison</surname></string-name>, <string-name><given-names>L. K.</given-names> <surname>Rosenstein</surname></string-name>, <string-name><given-names>T. A.</given-names> <surname>Morgan</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Zimmerman</surname></string-name></person-group>, &#x201C;<article-title>Community and clinical epidemiology of borderline personality disorder</article-title>,&#x201D; <source>Psychiatric Clinics of North America</source>, vol. <volume>41</volume>, no. <issue>4</issue>, pp. <fpage>561</fpage>&#x2013;<lpage>573</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-20"><label>20</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Foxhall</surname></string-name>, <string-name><given-names>C. H.</given-names> <surname>Giachritsis</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Button</surname></string-name></person-group>, &#x201C;<article-title>The link between rejection sensitivity and borderline personality disorder: A systematic review and meta-analysis</article-title>,&#x201D; <source>British Journal of Clinical Psychology</source>, vol. <volume>58</volume>, no. <issue>3</issue>, pp. <fpage>289</fpage>&#x2013;<lpage>326</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-21"><label>21</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Driessen</surname></string-name> and <string-name><given-names>S. D.</given-names> <surname>Hollon</surname></string-name></person-group>, &#x201C;<article-title>Cognitive behavioral therapy for mood disorders: Efficacy, moderators and mediators</article-title>,&#x201D; <source>The Psychiatric Clinics of North America</source>, vol. <volume>33</volume>, no. <issue>3</issue>, pp. <fpage>537</fpage>&#x2013;<lpage>555</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-22"><label>22</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Soler</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Vega</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Elices</surname></string-name>, <string-name><given-names>A. F.</given-names> <surname>Soler</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Soto</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Testing the reinforcement sensitivity theory in borderline personality disorder compared with major depression and healthy controls</article-title>,&#x201D; <source>Personality and Individual Differences</source>, vol. <volume>61</volume>, pp. <fpage>43</fpage>&#x2013;<lpage>46</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-23"><label>23</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Kulacaoglu</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Kose</surname></string-name></person-group>, &#x201C;<article-title>Borderline personality disorder (BPD): In the midst of vulnerability, chaos, and awe</article-title>,&#x201D; <source>Brain Sciences</source>, vol. <volume>8</volume>, no. <issue>11</issue>, pp. <fpage>201</fpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-24"><label>24</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. S.</given-names> <surname>Paul</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Sher</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Tamietto</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Mendl</surname></string-name></person-group>, &#x201C;<article-title>Towards a comparative science of emotion: Affect and consciousness in humans and animals</article-title>,&#x201D; <source>Neuroscience and Biobehavioral Reviews</source>, vol. <volume>108</volume>, no. <issue>205</issue>, pp. <fpage>749</fpage>&#x2013;<lpage>770</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-25"><label>25</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. D.</given-names> <surname>Spielberger</surname></string-name> and <string-name><given-names>E. C.</given-names> <surname>Reheiser</surname></string-name></person-group>, &#x201C;<article-title>Assessment of emotions: Anxiety, anger, depression, and curiosity</article-title>,&#x201D; <source>Applied Psychology: Health and Well-Being</source>, vol. <volume>1</volume>, no. <issue>3</issue>, pp. <fpage>271</fpage>&#x2013;<lpage>302</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-26"><label>26</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. A.</given-names> <surname>Arias</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Williams</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Raghvani</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Aghajani</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Baez</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>The neuroscience of sadness: A multidisciplinary synthesis and collaborative review</article-title>,&#x201D; <source>Neuroscience and Biobehavioral Reviews</source>, vol. <volume>111</volume>, no. <issue>Suppl 2</issue>, pp. <fpage>199</fpage>&#x2013;<lpage>228</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-27"><label>27</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Lalitha</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Tripathi</surname></string-name></person-group>, &#x201C;<article-title>Emotion detection using perceptual based speech features</article-title>,&#x201D; in <conf-name>Proc. 2016 IEEE Annual India Conf.</conf-name>, <publisher-loc>Bangalore, India</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2016</year>. </mixed-citation></ref>
<ref id="ref-28"><label>28</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Hassan</surname></string-name>, <string-name><given-names>R.</given-names> <surname>McCabe</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Priebe</surname></string-name></person-group>, &#x201C;<article-title>Professional-patient communication in the treatment of mental illness: A review</article-title>,&#x201D; <source>Communication and Medicine</source>, vol. <volume>4</volume>, no. <issue>2</issue>, pp. <fpage>141</fpage>&#x2013;<lpage>152</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-29"><label>29</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Bates</surname></string-name></person-group>, &#x201C;<article-title>Health care chatbots are here to help</article-title>,&#x201D; <source>IEEE Pulse</source>, vol. <volume>10</volume>, no. <issue>3</issue>, pp. <fpage>12</fpage>&#x2013;<lpage>14</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-30"><label>30</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Latif</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Qadir</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Qayyum</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Usama</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Younis</surname></string-name></person-group>, &#x201C;<article-title>Speech technology for healthcare: Opportunities, challenges, and state of the art</article-title>,&#x201D; <source>IEEE Reviews in Biomedical Engineering</source>, vol. <volume>14</volume>, pp. <fpage>342</fpage>&#x2013;<lpage>356</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-31"><label>31</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. N.</given-names> <surname>Vaidyam</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wisniewski</surname></string-name>, <string-name><given-names>J. D.</given-names> <surname>Halamka</surname></string-name>, <string-name><given-names>M. S.</given-names> <surname>Kashavan</surname></string-name> and <string-name><given-names>J. B.</given-names> <surname>Torous</surname></string-name></person-group>, &#x201C;<article-title>Chatbots and conversational agents in mental health: A review of the psychiatric landscape</article-title>,&#x201D; <source>Canadian Journal of Psychiatry. Revue Canadienne de Psychiatrie</source>, vol. <volume>64</volume>, no. <issue>7</issue>, pp. <fpage>456</fpage>&#x2013;<lpage>464</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-32"><label>32</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Poria</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Majumder</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Mihalcea</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Hovy</surname></string-name></person-group>, &#x201C;<article-title>Emotion recognition in conversation: Research challenges, datasets, and recent advances</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>100943</fpage>&#x2013;<lpage>100953</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-33"><label>33</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Bhargava</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Varshney</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Anita</surname></string-name></person-group>, &#x201C;<article-title>Emotionally intelligent chatBot for mental healthcare and suicide prevention</article-title>,&#x201D; <source>International Journal of Advanced Science and Technology</source>, vol. <volume>29</volume>, no. <issue>6</issue>, pp. <fpage>2597</fpage>&#x2013;<lpage>2605</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-34"><label>34</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Oh</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Ko</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Choi</surname></string-name></person-group>, &#x201C;<article-title>A chatbot for psychiatric counseling in mental healthcare service based on emotional dialogue analysis and sentence generation</article-title>,&#x201D; in <conf-name>Proc. 18th IEEE Int. Conf. on Mobile Data Management</conf-name>, <publisher-loc>Daejeon</publisher-loc>, pp. <fpage>371</fpage>&#x2013;<lpage>375</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-35"><label>35</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Bao</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Cummins</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Attention-enhanced connectionist temporal classification for discrete speech emotion recognition</article-title>,&#x201D; in <conf-name>Interspeech 2019</conf-name>, <conf-loc>Graz, Austria</conf-loc>, pp. <fpage>206</fpage>&#x2013;<lpage>210</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-36"><label>36</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Pan</surname></string-name></person-group>, &#x201C;<article-title>Speech emotion recognition using fusion of three multi-task learning-based classifiers: HSF-DNN, MS-CNN and LLD-RNN</article-title>,&#x201D; <source>Speech Communication</source>, vol. <volume>120</volume>, no. <issue>3</issue>, pp. <fpage>11</fpage>&#x2013;<lpage>19</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-37"><label>37</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. S.</given-names> <surname>Poorna</surname></string-name> and <string-name><given-names>G. J.</given-names> <surname>Nair</surname></string-name></person-group>, &#x201C;<article-title>Multistage classification scheme to enhance speech emotion recognition</article-title>,&#x201D; <source>International Journal of Speech Technology</source>, vol. <volume>22</volume>, no. <issue>2</issue>, pp. <fpage>327</fpage>&#x2013;<lpage>340</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-38"><label>38</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B. T.</given-names> <surname>Atmaja</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Akagi</surname></string-name></person-group>, &#x201C;<article-title>Dimensional speech emotion recognition from speech features and word embeddings by using multitask learning</article-title>,&#x201D; <source>APSIPA Transactions on Signal and Information Processing</source>, vol. <volume>9</volume>, pp. <fpage>2825</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-39"><label>39</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Wolf</surname></string-name></person-group>, &#x201C;<article-title>Measuring facial expression of emotion</article-title>,&#x201D; <source>Dialogues in Clinical Neuroscience</source>, vol. <volume>17</volume>, no. <issue>4</issue>, pp. <fpage>457</fpage>&#x2013;<lpage>462</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-40"><label>40</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Yazdani</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Skodras</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Fakotakis</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Ebrahimi</surname></string-name></person-group>, &#x201C;<article-title>Multimedia content analysis for emotional characterization of music video clips</article-title>,&#x201D; <source>EURASIP Journal on Image and Video Processing</source>, vol. <volume>26</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-41"><label>41</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Huang</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Epps</surname></string-name></person-group>, &#x201C;<article-title>An investigation of partition-based and phonetically-aware acoustic features for continuous emotion prediction from speech</article-title>,&#x201D; <source>IEEE Transactions on Affective Computing</source>, vol. <volume>11</volume>, no. <issue>4</issue>, pp. <fpage>653</fpage>&#x2013;<lpage>668</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-42"><label>42</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Tao</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Lian</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Niu</surname></string-name></person-group>, &#x201C;<article-title>Multimodal transformer fusion for continuous emotion recognition</article-title>,&#x201D; in <conf-name>Proc. Int. Conf. on Acoustics, Speech and Signal Processing</conf-name>, <publisher-loc>Barcelona, Spain</publisher-loc>, pp. <fpage>3507</fpage>&#x2013;<lpage>3511</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-43"><label>43</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Coutinho</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Schuller</surname></string-name></person-group>, &#x201C;<article-title>Dynamic difficulty awareness training for continuous emotion prediction</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>21</volume>, no. <issue>5</issue>, pp. <fpage>1289</fpage>&#x2013;<lpage>1301</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-44"><label>44</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Meng</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Yuan</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Wei</surname></string-name></person-group>, &#x201C;<article-title>Speech emotion recognition from 3D log-mel spectrograms with deep learning network</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>125868</fpage>&#x2013;<lpage>125881</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-45"><label>45</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Praveena</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Govind</surname></string-name></person-group>, &#x201C;<article-title>Significance of incorporating excitation source parameters for improved emotion recognition from speech and electroglottographic signals</article-title>,&#x201D; <source>International Journal of Speech Technology</source>, vol. <volume>20</volume>, no. <issue>4</issue>, pp. <fpage>787</fpage>&#x2013;<lpage>797</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-46"><label>46</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. H.</given-names> <surname>Yeh</surname></string-name>, <string-name><given-names>T. L.</given-names> <surname>Pao</surname></string-name>, <string-name><given-names>C. Y.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Y. W.</given-names> <surname>Tsai</surname></string-name> and <string-name><given-names>Y. T.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Segment-based emotion recognition from continuous Mandarin Chinese speech</article-title>,&#x201D; <source>Computers in Human Behavior</source>, vol. <volume>27</volume>, no. <issue>5</issue>, pp. <fpage>1545</fpage>&#x2013;<lpage>1552</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-47"><label>47</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Fan</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wu</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Cai</surname></string-name></person-group>, &#x201C;<article-title>Automatic emotion variation detection in continuous speech</article-title>,&#x201D; in <conf-name>Proc. Signal and Information Processing Association Annual Summit and Conf.</conf-name>, <publisher-loc>Siem Reap, Cambodia</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-48"><label>48</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H. B.</given-names> <surname>Kang</surname></string-name></person-group>, &#x201C;<article-title>Affective content detection using HMMs</article-title>,&#x201D; in <conf-name>Eleventh ACM Int. Conf. on Multimedia</conf-name>, <publisher-loc>New York, NY, USA</publisher-loc>, pp. <fpage>259</fpage>&#x2013;<lpage>262</lpage>, <year>2003</year>. </mixed-citation></ref>
<ref id="ref-49"><label>49</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Alshamsi</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Kepuska</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Alshamsi</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Meng</surname></string-name></person-group>, &#x201C;<article-title>Automated speech emotion recognition on smart phones</article-title>,&#x201D; in <conf-name>9th IEEE Annual Ubiquitous Computing, Electronics &#x0026; Mobile Communication Conf.</conf-name>, <publisher-loc>New York, USA</publisher-loc>, pp. <fpage>44</fpage>&#x2013;<lpage>50</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-50"><label>50</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Franklin</surname></string-name></person-group>, &#x201C;<article-title>The sheer audacity: How to get more, in less time, from the audacity digital audio editing software</article-title>,&#x201D; in <conf-name>IEEE Int. Professional Communication Conf.</conf-name>, <publisher-loc>Saragota Springs, NY, USA</publisher-loc>, pp. <fpage>92</fpage>&#x2013;<lpage>105</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-51"><label>51</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Lalitha</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zakariah</surname></string-name> and <string-name><given-names>Y. A.</given-names> <surname>Alotaibi</surname></string-name></person-group>, &#x201C;<article-title>Investigation of multilingual and mixed-lingual emotion recognition using enhanced cues with data augmentation</article-title>,&#x201D; <source>Applied Acoustics</source>, vol. <volume>170</volume>, no. <issue>1</issue>, pp. <fpage>107519</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-52"><label>52</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Muthuswamy</surname></string-name>, <string-name><given-names>D. L.</given-names> <surname>Sherman</surname></string-name> and <string-name><given-names>N. V.</given-names> <surname>Thakor</surname></string-name></person-group>, &#x201C;<article-title>Higher-order spectral analysis of burst patterns in EEG</article-title>,&#x201D; <source>IEEE Transactions on Bio-Medical Engineering</source>, vol. <volume>46</volume>, no. <issue>1</issue>, pp. <fpage>92</fpage>&#x2013;<lpage>99</lpage>, <year>1999</year>.</mixed-citation></ref>
<ref id="ref-53"><label>53</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T. T.</given-names> <surname>Ng</surname></string-name>, <string-name><given-names>S. F.</given-names> <surname>Chang</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Blind detection of photomontage using higher order statistics</article-title>,&#x201D; in <conf-name>Int. Symp. on Circuits and Systems</conf-name>, <publisher-loc>Vancouver, BC, Canada</publisher-loc>, vol. <volume>5</volume>, pp. <fpage>688</fpage>&#x2013;<lpage>691</lpage>, <year>2004</year>. </mixed-citation></ref>
<ref id="ref-54"><label>54</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Du</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Dua</surname></string-name>, <string-name><given-names>R. U.</given-names> <surname>Acharya</surname></string-name> and <string-name><given-names>C. K.</given-names> <surname>Chua</surname></string-name></person-group>, &#x201C;<article-title>Classification of epilepsy using high-order spectra features and principle component analysis</article-title>,&#x201D; <source>Journal of Medical Systems</source>, vol. <volume>36</volume>, no. <issue>3</issue>, pp. <fpage>1731</fpage>&#x2013;<lpage>1743</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-55"><label>55</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Lalitha</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Tripathi</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Gupta</surname></string-name></person-group>, &#x201C;<article-title>Enhanced speech emotion detection using deep neural networks</article-title>,&#x201D; <source>International Journal of Speech Technology</source>, vol. <volume>22</volume>, no. <issue>3</issue>, pp. <fpage>497</fpage>&#x2013;<lpage>510</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-56"><label>56</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Dobesova</surname></string-name></person-group>, &#x201C;<article-title>Programming language Python for data processing</article-title>,&#x201D; in <conf-name>Int. Conf. on Electrical and Control Engineering</conf-name>, <publisher-loc>Yichang, China</publisher-loc>, pp. <fpage>4866</fpage>&#x2013;<lpage>4869</lpage>, <year>2011</year>. </mixed-citation></ref>
<ref id="ref-57"><label>57</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Noroozi</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Sapinski</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kaminska</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Anbarjafari</surname></string-name></person-group>, &#x201C;<article-title>Vocal-based emotion recognition using random forests and decision tree</article-title>,&#x201D; <source>International Journal of Speech Technology</source>, vol. <volume>20</volume>, no. <issue>2</issue>, pp. <fpage>239</fpage>&#x2013;<lpage>246</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-58"><label>58</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Goutte</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Gaussier</surname></string-name></person-group>, &#x201C;<article-title>A probabilistic interpretation of precision, recall and f-ccore, with implication for evaluation</article-title>,&#x201D; <source>Lecture Notes in Computer Science</source>, vol. <volume>3408</volume>, pp. <fpage>345</fpage>&#x2013;<lpage>359</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-59"><label>59</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Lalitha</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Gupta</surname></string-name></person-group>, &#x201C;<article-title>An encapsulation of vital non-linear frequency features for speech applications</article-title>,&#x201D; <source>Journal of Computational and Theoretical Nanoscience</source>, vol. <volume>17</volume>, no. <issue>1</issue>, pp. <fpage>303</fpage>&#x2013;<lpage>307</lpage>, <year>2018</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>