<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CSSE</journal-id>
<journal-id journal-id-type="nlm-ta">CSSE</journal-id>
<journal-id journal-id-type="publisher-id">CSSE</journal-id>
<journal-title-group>
<journal-title>Computer Systems Science &#x0026; Engineering</journal-title>
</journal-title-group>
<issn pub-type="ppub">0267-6192</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">24967</article-id>
<article-id pub-id-type="doi">10.32604/csse.2022.024967</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Speak-Correct: A Computerized Interface for the Analysis of Mispronounced Errors</article-title><alt-title alt-title-type="left-running-head">Speak-Correct: A Computerized Interface for the Analysis of Mispronounced Errors</alt-title><alt-title alt-title-type="right-running-head">Speak-Correct: A Computerized Interface for the Analysis of Mispronounced Errors</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Jambi</surname><given-names>Kamal</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref><email>kjambi@kau.edu.sa</email>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Al-Barhamtoshy</surname><given-names>Hassanin</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Al-Jedaibi</surname><given-names>Wajdi</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Rashwan</surname><given-names>Mohsen</given-names></name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Abdou</surname><given-names>Sherif</given-names></name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<aff id="aff-1"><label>1</label><institution>Faculty of Computing &#x0026; Information Technology, King Abdulaziz University</institution>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Faculty of Engineering, Cairo University</institution>, <country>Egypt</country></aff>
<aff id="aff-3"><label>3</label><institution>Faculty of Computers, Cairo University</institution>, <country>Egypt</country></aff>
</contrib-group><author-notes><corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Kamal Jambi. Email: <email>kjambi@kau.edu.sa</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-05-06"><day>06</day>
<month>05</month>
<year>2022</year></pub-date>
<volume>43</volume>
<issue>3</issue>
<fpage>1155</fpage>
<lpage>1173</lpage>
<history>
<date date-type="received"><day>06</day><month>11</month><year>2021</year></date>
<date date-type="accepted"><day>09</day><month>12</month><year>2021</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Jambi et al.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Jambi et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CSSE_24967.pdf"></self-uri>
<abstract>
<p>Any natural language may have dozens of accents. Even though the equivalent phonemic formation of the word, if it is properly called in different accents, humans do have audio signals that are distinct from one another. Among the most common issues with speech, the processing is discrepancies in pronunciation, accent, and enunciation. This research study examines the issues of detecting, fixing, and summarising accent defects of average Arabic individuals in English-speaking speech. The article then discusses the key approaches and structure that will be utilized to address both accent flaws and pronunciation issues. The proposed SpeakCorrect computerized interface employs a cutting-edge speech recognition system and analyses pronunciation errors with a speech decoder. As a result, some of the most essential types of changes in pronunciation that are significant for speech recognition are performed, and accent defects defining such differences are presented. Consequently, the suggested technique increases the Speaker&#x2019;s accuracy. SpeakCorrect uses 100 h of phonetically prepared individuals to construct a pronunciation instruction repository. These prerecorded sets are used to train Hidden Markov Models (HMM) as well as weighted graph systems. Their speeches are quite clear and might be considered natural. The proposed interface is optimized for use with an integrated phonetic pronounced dataset, as well as for analyzing and identifying speech faults in Saudi and Egyptian dialects. The proposed interface detects, analyses, and assists English learners in correcting utterance faults, overcoming problems, and improving their pronunciations.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Speech recognition</kwd>
<kwd>computerized interface</kwd>
<kwd>arabic dialects</kwd>
<kwd>accent defects</kwd>
<kwd>acoustic error</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>As english has become a global language, several people are learning and speaking it. Among the most prevalent, but extremely complicated, activities to address when learning English as an additional language is the use of English in verbal communication. This is because English proficiency has become a need, particularly for those wishing to develop in specific domains of human effort. Speaking English as a foreign language is a significant issue for several people throughout the world. This occurs for a variety of reasons. When learning a foreign language, you must grasp that it employs a wide variety of sounds as well as orthographic norms than your native speech. Learners frequently try to mimic the sounds by using ones they are already acquainted with and pronouncing texts as if they were composed in their home languages.</p>
<p>Any speech processor is made up of two or three components. The first is a word recognizer, which turns any activities with respect into fundamental word sequence. The next is a phoneme recognition system based on any paradigm, such as HTK, that turns the provided utterance into a phoneme series. This series is then evaluated with a matching algorithm to identify the most relevant keywords.</p>
<p>The acoustic input O is treated as a sequence of individual &#x201C;symbols&#x201D; or &#x201C;observations&#x201D;, represented by symbols: O &#x003D; o1, o2, o3, &#x2026;, ot. Likewise, a sentence/word will be preserved as a string of words/phonemes: W &#x003D; w1, w2, w3, &#x2026;, wn. Various general terminologies used throughout this document, are explained in this section [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>This research paper is written specifically for practising non-specialist learners/students who really need to communicate In english appropriately. The suggested approach uses the fewest linguistic terms possible and seeks to deliver straightforward evaluation results in visual form. This is especially true in the case of speech, where unnecessary technical information can be perplexing to non-specialists. Furthermore, we consider that the explanations and technical aspects provided here are accurate and sufficiently detailed within the context of this technique.</p>
<p>The objectives of this study include designing, building, and assessing a prototype system which can detect and analyse mispronounced mistakes in adult learners&#x2019; teaching-based operations. This system is not designed to be a total replacement for the actual class teacher, but rather to serve as an additional tool to assist him/her in teaching the fundamental skills of learning and practising in the various fields of advancement.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<p>To comprehend signs communicated by a voice signal, &#x201C;Spoken Language Understanding (SLU)&#x201D; as well as &#x201C;Natural Language Understanding (NLU)&#x201D; are utilised. The derived conceptual description of natural language sentences is contributed by SLU and NLU [<xref ref-type="bibr" rid="ref-5">5</xref>]. Signs can be used for understanding and can be encoded into signals that provide extra information. Moreover, the suggested two systems feature an Automatic Speech Recognition (ASR) component and should be sound sensitive, depending on the pattern of spoken language as well as ASR failures.</p>
<p>Dialog categorization and automated segmentation are critical for interpreting SLU. This research proposed a system for segmenting and classifying multiparty meetings based on contextual speech as well as prosodic traits. &#x201C;Contextual elements are better for recognising, but prosodic features are better for identifying base processes and backchannels,&#x201D; they discover [<xref ref-type="bibr" rid="ref-6">6</xref>]. In data screening, voice is employed. A phonetic matching method has been provided, and the given application is used in the music sector; search errors on text and spoken queries have been reduced [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Heracleous et al. (2009) offer consonant and vowel identification in French using HMM in their work. Speech hand-shapes with lip-patterns (as a graphic communication medium) build all oral language sounds clearly, particularly for deaf as well as hearing-impaired individuals. The goal of this study is to overcome the difficulty of lip reading and so enable deaf children and adults to completely perceive spoken language [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Two ways are used to improve comprehensibility for those who are deaf or deafeningly. The first technique is utilised in the setting of discussion for hearing-impaired users; the second strategy tries to improve speaking-impaired people&#x2019;s intelligibility. The results revealed that an improvement in intelligibility was not attained, and listeners preferred the altered speech of an alternative approach.</p>
<p>Linguistic knowledge is commonly employed in ASR to improve mistake prediction. Tsubota et al. modelled 79 different types of pronunciation error patterns to distinguish Japanese students&#x2019; English, and the paper proposes a straightforward way to following the pitch of two active presenters. It used HMM to track frequency over time. To illustrate experimental findings, a statistical approach might be utilised. The research demonstrated that the suggested technique outperformed the multi-pitch tracking algorithm [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>To enable students to acquire Japanese, a Computer-Assisted Language Learning (CALL) approach has been designed [<xref ref-type="bibr" rid="ref-11">11</xref>]. The study provides a model for detecting lexico-grammatical problems in pronunciation, as well as inputting sentence faults. The researchers of [<xref ref-type="bibr" rid="ref-12">12</xref>] presents a technique for generating acoustic sub-word entities from a spoken term recognition system, which can be used to replace standard phone models. The system uses an unsupervised method to construct a series of speaker models that are independent of the training data. Another work [<xref ref-type="bibr" rid="ref-13">13</xref>] presents a decision tree that may be used for error categorization in automatic voice recognition to discover critical and redundant errors.</p>
<p>The study [<xref ref-type="bibr" rid="ref-14">14</xref>] concentrates on non-native accent issues in uninterrupted speech recognition. It attempts to investigate the transformation principles of non-native speech expressed in Mandarin by local speakers. As a result, a corpus in Mandarin is used to train the HMM algorithms and check the effect of voice recognition. As a consequence, the results show that the collected information is useful for adjusting a native speaker ASR method to model nonnative accented content.</p>
<p>Another article, titled &#x201C;Vowel Effects on Dental Arabic Consonants Based on Spectrogram&#x201D; [<xref ref-type="bibr" rid="ref-15">15</xref>], investigated the influence of Arabic vowels on Arabic consonants with three easy diacritics by Malaysian youngsters. Malaysian children add these vowels to the fundamental consonants using three simple diacritics. The location of articulation is essential in dental consonants and formant frequencies, according to the report.</p>
<p>Kensaku Fujii et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] cancelled acoustic echo by replacing the differential between adaptable filter coefficients for the prediction error in the former. Nevertheless, Juraj Kac and Gregor Rozinaj et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] investigated the effect of replacing a number of speech parameters with voiced/unvoiced data in approximated pitch value. Investigations were carried out on the data for smartphone platforms, employing context dependent as well as independent phonemes from HMM models.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>General Terminology</title>
<p>This general terminology is intended to help readers comprehend commonly used terminology and phrases when researching, interpreting, and assessing scholarly method of investigation. Basic words or phrases are also discussed in the perspective of how they relate to commence research for the SpeakCorrect computerised interface.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Phoneme</title>
<p>A phoneme is the smallest element of speech; it is used to differentiate meaning. The phoneme is the most essential element in the word; every word is made up of phonemes, therefore changing them changes the essence of the word. When the sound [b] in the word &#x201C;pin&#x201D; is substituted by the sound [p], the word becomes &#x201C;bin.&#x201D; As a result, /b/ is a phoneme [<xref ref-type="bibr" rid="ref-18">18</xref>]. The phoneme is characterized purely in terms of changes in its immediate phonetic context that constrain allophonic variations, with no regard for higher-level language structures. To describe a phoneme, just remark that two essentially identical utterances (in practise, words because they are the shortest utterance length in a naming task) vary due to the presence of two separate sound components.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Phone</title>
<p>It is the lowest physical sound part. Phones are thus the physical manifestation of phonemes. Allophones are sometimes defined as phonic variations of phonemes [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Phonetics</title>
<p>Phonetics is the analysis of human voice, specifically the qualities of spoken tones [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Phonology</title>
<p>Phonology is the study of audio systems and the abstraction of sound units, such as phonemes as well as phonological principles. As a result, phonetics concepts are universal, whereas phonology is language specific. [<xref ref-type="bibr" rid="ref-21">21</xref>] represents the phonetic of a sound, while // represents the phoneme.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Syllable</title>
<p>The term &#x201C;syllable&#x201D; refers to a component of pronunciation. It is greater than a single sound but smaller than a word. The syllable begins and finishes with consonants and also is made up of vowels [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Common Pronunciation Errors</title>
<p>One of the major issues with spoken English is the preference for informal over formal communication. In informal communication, ellipsis, contractions, and relative clauses lacking relative pronouns are increasingly common. Formal speech follows traditional grammatical standards and is typically employed for strangers or in writing. &#x201C;He is my first brother,&#x201D; for instance, sounds more formal than &#x201C;He&#x2019;s my first brother,&#x201D; which sounds more informal. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> depicts almost all of the pronunciation problems that must be considered in effective training and evaluation models [<xref ref-type="bibr" rid="ref-23">23</xref>]. It can be divided into two sorts of errors, as illustrated in the <xref ref-type="fig" rid="fig-1">Fig. 1</xref>: phonemic and prosodic.<list list-type="bullet"><list-item>
<p>In this work, phonemic errors are classified as substituted, removed, or introduced. There are even minor errors &#x201C;when the proper phoneme is greater or fewer being pronounced&#x201D; [<xref ref-type="bibr" rid="ref-24">24</xref>].</p></list-item><list-item>
<p>Stress, rhythm, and intonation are the three types of prosodic faults.</p></list-item></list></p>
<p>As a result of these two types of faults, pronunciation becomes a multi-dimensional issue. As a result, a wide variety of measures are used to quantify these characteristics [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Classification of pronunciation errors</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-1.png"/>
</fig>
<p>We identified a considerable body of research outlining the usual patterns of error committed by Korean Learner Segmental Errors (KLEs) while developing SpeakCorrect. A pilot corpus of suggested English speech data from a variety of various forms of material is collected and phonetically labeled. Short sections of text, phrase prompts, and words with terribly challenging consonant clusters are included in the <italic>corpus</italic> (e.g., refrigerator). In total, 25,000 speech specimens were taken from 111 Korean learners for the pilot <italic>corpus</italic>. The <italic>corpus</italic> allows for direct comparisons of realised phone sequences to anticipated canonical patterns from native speakers. A description of Korean for English speakers likewise does not distinguish between fricative /f/ and /v/, instead substituting /p/ and /b/. Other commonly reported errors involve aspirated /t/ for / as well as un-aspirated /t/ for /. In addition, the report demonstrates the most common segmentation errors encountered. The current work focuses on mistake substitution, removal, and addition in phonemics.</p>
</sec>
<sec id="s5">
<label>5</label>
<title>SpeakCorrect Error Detection Processing Engine</title>
<p>The proposed computerised interface is divided into three major steps. The first stage is utilised for pre-processing or feature training collection, development of pronunciation hypotheses, and HMM adaptation. The second stage employs a decoder that detects the user&#x2019;s supplied speech, as well as a confidence measurement and a pronunciation fault detector. This step&#x2019;s output is the third stage, which comprises analysis of the respond to changes and generates recognised errors, advice to the repaired errors, and an evaluation method. As a result, phone estimation use statistical techniques (such as neural networks or Gaussian models) to detect particular speech sounds such as f or s. This step creates a vector of possibilities over phones for every frames. The last stage contains numerous sub-modules, the most significant of which is the decoding module, which is used to discover the sequence of words with the greatest chance given the acoustic occurrences.</p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> depicts the Data Flow Diagram (DFD) of the suggested SpeakCorrect interface. Each word entails the students being depicted as in ATN or Lattice, which is a visual showing several phonetic levels and accompanying training (Saudi or Egyptian accent defects). As a result, Speak Correct gives users the option of choosing their own degrees and instances. After entering the syllables, computerized speech recognition is used, which is backed by trained instances in the manner of a grammar network (Lattice graph) for the specific word. Errors will be discovered and assessed, and feedback statistics will be provided to users, resulting in the evolution of the model.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>SpeakCorrect error detection overview</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-2.png"/>
</fig>
<p>To explain the feature extraction procedure, beginning with sound waves as well as finishing with a feature vector. Firstly, an input audio wave is digitised (analog-to-digital conversion), which is accomplished in two steps: Quantification and sampling.</p>
<p>The number of observations taken per second is referred to as the sampling rate. To precisely measure a waveform, at least two samples are required in each cycle: one for the positive half of the wave and one for the negative part, and more than two samples per cycle improve amplitude precision.</p>
<p>There are volume measurements for each second of voice at each sample rate. Quantization refers to the technique of expressing such quantities as integers. The waveform is transformed after digitization to set the spectral characteristics. Any prominent feature set (Linear Predictive Coding (LPC) or Perceptual Linear Predictive (PLP)) can be used directly to view HMM symbols [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-30">30</xref>]. The phones assessment is a fast method for calculating the likelihood of an ordered set given weighted automata. HMM enables us to add together numerous routes that individually account for the very same observation sequence.</p>
<p>The decoding stage is concerned with determining the right &#x201C;underlying&#x201D; succession of symbols/patterns. As a result, the Veterbi algorithm provides an efficient method of addressing the decoding problem by evaluating all potential strings and employing addition rules (such as the Bays rule) to determine their probability of obtaining the observed sequence.</p>
<p>The features are frequently subjected to additional processing in order to match the standard speech models to the presenter&#x2019;s speech attributes. As a result, the speaker adjustment module of Maximum Likelihood Linear Regression (MLLR) is employed to refine the adapted module.</p>
<sec id="s5_1">
<label>5.1</label>
<title>SpeakCorrect Background Architecture</title>
<p>Several techniques used in voice recognition have been developed by different academics. As a result, the terms phone and syllable are formed [<xref ref-type="bibr" rid="ref-31">31</xref>&#x2013;<xref ref-type="bibr" rid="ref-33">33</xref>]. Furthermore, the N-gram language concept as well as the HMM are described in greater depth.</p>
<p>Initially, HMM were presented as stochastic approaches for modelling temporal pattern classification and analysis systems. As a result, the HMMs can be shown employing finite state machines, with an assessment from a given state at each transition and output symbol emissions for each state. In another words, to select the most likely word given the insight, the single word with the highest P (word | observation). If w is the predicted right terminology and O is the recorded sequence (individual observation), the <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> for deciding the proper word is:</p>
<p><disp-formula id="eqn-1"><label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mi mathvariant="normal">W</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">x</mml:mi><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">o</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">w</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">w</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where P(o|w) denotes likelihood, P(w) denotes prior, w denotes vocabulary, w denotes right word, and o denotes observation. Once the probability computations and decoding difficulties for a reduced input comprising of strings of phones have been addressed, feature extraction would be swiftly implicated.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Acoustic Probabilities Counting</title>
<p>As previously stated, the speech intake can be transformed into a sequence of feature vectors by passing it through signal conditioning transformations, with every vector reflecting one time-slice of the spoken input signal. One frequent method for computing probability on extracted features is to first cluster them into discrete numbered symbols. As a result, the likelihood of a specific cluster can be computed (number of times it occurs in some training set).</p>
<p>This technique is known as vector quantization, which is used to compute observation possibilities or probability density functions (pdf). There really are two typical methods: Gaussian pdfs, which convert the observation vector Ot to a likelihood, and neural networks or multi-layer perspectives, which may be trained to attribute a possibility to a real-valued feature space in audio. A neural network is a collection of small computer units linked together by weighted connections. When given vector variables, the network calculates a vector of target value.</p>
<p>Mishra et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] proposed a standard model which is founded on a probabilistic neural network that is suited for testing as well as pattern categorization. The number of input voice variables M, the number of recognition patterns required N, and the training instances for each pattern are denoted by S1, S2,&#x2026; SN comprise the architecture of such probabilistic neural network model. There are four layers: input layer, model layer, summation layer, as well as output layer; the weights among accumulation layer with output layer are calculated as follows (<xref ref-type="disp-formula" rid="eqn-2">Eq. 2</xref>):</p>
<p><disp-formula id="eqn-2"><label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mi mathvariant="normal">W</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:math>
</disp-formula></p>
<p>As a result, when speaker adaptation occurs, specific qualities such as acoustic gender, dialects, and age would be modelled in speech processing. As a result, the speaker&#x2019;s unique accent should be unaffected.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>SpeakCorrect Principles Modules</title>
<p>Our purpose is to create a model that can find out how it changed this &#x201C;true&#x201D; term and therefore recover it. The essential recognition procedures for the complete talk right tasks are as follows.</p>
<sec id="s5_3_1">
<label>5.3.1</label>
<title>Main Module</title>
<p>Step 1: Gathering and collecting the speech input samples.</p>
<p>Step 2: Dividing such samples into two parts, one part is for training samples and the second part is for testing.</p>
</sec>
<sec id="s5_3_2">
<label>5.3.2</label>
<title>Training Module</title>
<p>Step 3: Do the following:</p>
<p>3.1 Speech Adaption.</p>
<p>3.2 Confidence measuring.</p>
<p>3.3 Tuning the native Arabic speaker accent.</p>
<p>a. Tuning Saudi accent.</p>
<p>b. Tuning Egyptian accent.</p>
<p>3.4 Intonation training and teaching the pronunciation effects.</p>
<p>Step 4: Using the feature vector of training samples to train the SpeakCorrect model.</p>
</sec>
<sec id="s5_3_3">
<label>5.3.3</label>
<title>Testing Module</title>
<p>Step 5: Do the following steps:</p>
<p>Step 5.1: Establishing the system with the associated acoustic and language model.</p>
<p>Step 5.2: The feature vectors are used to input test samples into network which has been trained.</p>
<p>Step 5.3: Judging the equivalent speech signal class and the speaker characteristics according to the output values.</p>
</sec>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Weighted Finite State and Weighted ATN/Lattice</title>
<p>Computing languages and automata concept were utilised to forecast letter sequences, characterise natural language, apply Context-Free Grammars (CFG), present tree transducer ideas, and parse automatic natural language writing. In the 1970s, voice processing investigators recorded NLP grammar utilizing weighted Finite State Acceptors (FSAs), which could be trained on machine-readable dictionary, <italic>corpus</italic>, and corpora [<xref ref-type="bibr" rid="ref-35">35</xref>&#x2013;<xref ref-type="bibr" rid="ref-41">41</xref>].</p>
<p>In the nineties, finite state machines and massive training corpora has become the dominant models in speech processing, prompting the development of software development tools for Weighted Finite State Machines (WFSM) [<xref ref-type="bibr" rid="ref-42">42</xref>]. In the twenty-first century, common tree automata toolkits [<xref ref-type="bibr" rid="ref-42">42</xref>&#x2013;<xref ref-type="bibr" rid="ref-47">47</xref>] have now been designed to aid research.</p>
<p>Although the single WFST or Augmented Transition Network (ATN) that depicts P(S|E) remains complex, model conversion can be converted into a sequence of transducers as shown in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.</p>
<p><disp-formula id="eqn-3"><label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">F</mml:mi><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">M</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">E</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">x</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo mathvariant="bold" rspace="0pt" stretchy="false">&#x2190;</mml:mo><mml:mo lspace="0pt" stretchy="false">&#x2192;</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">F</mml:mi><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">M</mml:mi><mml:mi mathvariant="normal">b</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">E</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>As a result, a simple model can be utilised to compute 1-, 2-, and n-gram language models of characters. If a <italic>corpus</italic> has 1,000,000 characters, the letter e appears 127,000 times, with a probability P(e) of 0.127.</p>
<p>In the context of a 2-gram system, it can be computed by recalling the preceding letter context- its WFSA condition. The probability P(e|s) can be used to compute the transformation between states s and e, which results in the letter e. The n-gram model produces more word-like elements than the (n-1)-gram approach. The weighted or lattice automaton is a simple automaton whereby each arc is connected with a transition, that can be expressed by a probability level suggesting how that path should be pursued. All arcs escaping a node must have a probability of one. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> depicts a weighted ATN trained on actual pronunciation examples for the English word &#x201C;around.&#x201D; This is an example of a HMM. The behaviour transitioning in the weighted ATN is depicted graphically in this picture. The principle of the transition is as follows:<list list-type="bullet"><list-item>
<p>Starts in some initial state (start: s1 ) with probability p(si).</p></list-item><list-item>
<p>On each move, goes from state si to state sj according to transition probability P (si, sj).</p></list-item><list-item>
<p>At each state si, it emits a symbol w<sub>k</sub> according to the emit probability P&#x2019;(s<sub>i</sub>, w<sub>k</sub>).</p></list-item></list></p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>A pronunciation network (weighted atn) for word &#x201C;about&#x201D;</title>
<p>Legends: P(w | ax ) &#x003D; .68 P(w | ix ) &#x003D; .20</p></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-3.png"/>
</fig>
<p>The proposed SpeakCorrect computerised interface is referred to as a hybrid approach since it employs components of the HMM or weighted state-graph model of a word&#x2019;s pronunciation, as well as observation-probability computing through multilayer perception. This network contains a single output unit for each phone; by adding all output unit numbers to one, the SpeakCorrect may be used to calculate the probability of a condition j given an observation vector Ot, P(qj | ot), or P(ot | qj). As a result, when given the series of spoken words which created a specific aural speech, a standard model - such as the one depicted in <xref ref-type="fig" rid="fig-3">Fig. 3</xref> is utilised. The model produces P(E|S) given a received speech signal S, and it is described as follows.<list list-type="bullet"><list-item>
<p>A series of phonemes is detected with different probability for each phonetic in S and can thus be construed as the word.</p></list-item><list-item>
<p>A word-to-phone correspondence is created for every phonetic.</p></list-item><list-item>
<p>Every phone can be represented by a different set of audio signals.</p></list-item></list></p>
<p>Once constructed, the input audio sequence and the final communicative approach are weighted using the likelihood technique and detecting probabilities from the training dataset.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Training the SpeakCorrect</title>
<p>In most ASR systems, a concise summary of the integrated training process is provided. A few of the algorithm&#x2019;s features are discussed in [<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>]. To train the SpeakCorrect system, four probabilistic models are required:<list list-type="bullet"><list-item>
<p>Language model probabilities: P(wi|wi-1 wi-2).</p></list-item><list-item>
<p>Likelihood observation: bj(ot).</p></list-item><list-item>
<p>Transition probabilities: aij.</p></list-item><list-item>
<p>Pronunciation Lexicon: Lattice or Weighted ATN of HMM state graph structure.</p></list-item></list></p>
<p>The SpeakCorrect features the following corpora for training the past probabilities element:<list list-type="bullet"><list-item>
<p>Speech wave document training <italic>corpus</italic>: that is compiled from news websites on the internet, specific individuals and so on. Such speech wave documents are grouped with word-transactions.</p></list-item><list-item>
<p>Text <italic>corpus</italic> containing a large number of comparable texts, such as the word-transaction out from speech database.</p></list-item><list-item>
<p>A smaller standard training <italic>corpus</italic> of audio that has been phonetically labelled, i.e., frames have been hand-annotated containing phonemes.</p></list-item></list></p>
<p>An off-the-shelf pronunciation vocabulary is used to build the HMM lexicon architecture. As a result, training begins with running the system on the observations and determining which transitions as well as observations were utilized. Any state can produce one observation symbol, and all observation possibilities are 1.0 [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>]. The probabilities pij of a specific transition from condition I to condition j can be calculated by computing the number of transitions undertaken; c I j), then normalising such result using the <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>.</p>
<p><disp-formula id="eqn-4"><label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">i</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mi mathvariant="normal">j</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>Two strategies are utilised for lattice or weighted ATN and HMM. The first is to iterate calculate the numbers and observation possibilities, and then use similar estimated probabilities to create progressively better probability values. The second step is to compute forward probability along all possible paths to obtain predicted probabilities. Considering the automaton A, calculate the forward probabilities in state I after witnessing the first t occurrences (<xref ref-type="disp-formula" rid="eqn-5">Eqs. (5)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-8">(8)</xref>).</p>
<p><disp-formula id="eqn-5"><label>(5)</label>
<mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="normal">P</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">i</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mi>A</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:math>
</disp-formula></p>
<p>Formally, describe the following iteration:</p>
<p>1. Initialization:</p>
<p><disp-formula id="eqn-6"><label>(6)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>&#x2217;</mml:mo><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">j</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">N</mml:mi></mml:mrow></mml:math>
</disp-formula></p>
<p>2. Iteration:</p>
<p><disp-formula id="eqn-7"><label>(7)</label>
<mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">t</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>&#x2217;</mml:mo><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">j</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">N</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">t</mml:mi><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:math>
</disp-formula></p>
<p>3. Termination:</p>
<p><disp-formula id="eqn-8"><label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>&#x2217;</mml:mo><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:math>
</disp-formula></p>
<p>To calculate the alternative phonemes which were most likely given the observational sequence [axe b], with a forward algorithm is performed, and the combination P(o | w) P(w) is produced for each potential word. So, for every word, the probability of ordered set o provided the word w times the previous likelihood of the word is determined, and the word with the greatest value is chosen.</p>
<p>A forward method is an edit distance method, and an intermediary table is utilized to hold the observation sequence&#x2019;s probability numbers. The data can be presented in the table by rows that are orientated; the rows are marked by a state-graph that contains multiple paths from one state to another. With determining the number of each cell from the three cells surrounding it, the table is completed as a matrix. In addition, the forward technique calculates the total of probability for all alternative paths that could yield the observation series. Formally, each cell represents the probability (<xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>):</p>
<p><disp-formula id="eqn-9"><label>(9)</label>
<mml:math id="mml-eqn-9" display="block"><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>j</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>o</mml:mi><mml:mrow><mml:mn mathvariant="italic">1</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>o</mml:mi><mml:mrow><mml:mn mathvariant="italic">2</mml:mn></mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>q</mml:mi><mml:mi>t</mml:mi><mml:mspace width="thickmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>j</mml:mi><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mspace width="thickmathspace" /><mml:mi>A</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>The forward algorithm is implemented to any word is described in the pseudo code below.</p>
<p>forwardAlgorithm ( observation, state-graph )</p>
<p>begin</p>
<p>ns &#x003D; numOfStates(state-graph);</p>
<p>no &#x003D; length(observation);</p>
<p>/&#x002A; create probability matrix &#x002A;/</p>
<p>forward [ns &#x002B; 2 , no &#x002B; 2];</p>
<p>forward [0,0] &#x003D; 1.0;</p>
<p>foreach time step t from 0 to no do</p>
<p>foreach states from 0 to ns do</p>
<p>foreach transition s&#x2019; from s specified by state-graph</p>
<p>forward [s&#x2019; , t &#x002B;1] &#x003D; forward [s, t] &#x002A; a[s, s&#x2019;] &#x002A; b [s&#x2019;, ot];</p>
<p>return sum of the probabilities in the final column of forward;</p>
<p>end.</p>
<p>where:</p>
<p>a [s , s&#x2019;] signifies transition possibility from present state s to subsequent state s&#x2019;.</p>
<p>b [s&#x2019;, ot] is the observation probability of s&#x2019; given ot.</p>
<p>b [s&#x2019;, ot] is equal 1 if the observation symbol matches the state, and is equal 0 otherwise.</p>
<p>The part of the forward-backward algorithm is the backward probability. This backward algorithm is almost the mirroring of the forward probability. It computes the probability of the observations from t&#x002B;1 to the end. Suppose that we are in state j at time t in given automaton A; then (<xref ref-type="disp-formula" rid="eqn-10">Eqs. (10)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-13">(13)</xref>):</p>
<p><disp-formula id="eqn-10"><label>(10)</label>
<mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi><mml:mspace width="thickmathspace" /><mml:mo stretchy="false">(</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">j</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mi>A</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:math>
</disp-formula></p>
<p>The backward computation is defined as the following:</p>
<p>Initialization:</p>
<p><disp-formula id="eqn-11"><label>(11)</label>
<mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">i</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">N</mml:mi></mml:mrow></mml:math>
</disp-formula></p>
<p>Iteration:</p>
<p><disp-formula id="eqn-12"><label>(12)</label>
<mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mn>1</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">j</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>&#x27E8;</mml:mo><mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">N</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x27E9;</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">t</mml:mi><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&lt;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math>
</disp-formula></p>
<p>Termination:</p>
<p><disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x007C;</mml:mo><mml:mi>A</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x2009;</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mtext>&#x2009;</mml:mtext></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mstyle><mml:mtext>&#x2009;</mml:mtext><mml:msub><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x2009;</mml:mtext><mml:mo>&#x002A;</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msub><mml:mrow></mml:mrow><mml:mtext>&#x2009;</mml:mtext></mml:msub><mml:mtext>&#x2009;</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>Therefore, the transition probability, aij, and observation probability, bi(ot), will be computed from an observation sequence.</p>
</sec>
<sec id="s5_6">
<label>5.6</label>
<title>SpeakCorrect Editor and Phonetic-based Approach</title>
<p>SpeakCorrect&#x2019;s conversion principles are a phonetic-based technique that converts text words into pronunciation terms. As a result, an intermediary form among source and the destination words is utilised, in addition to the phonetic-translation principle that captures the pronunciation of the specific words. There would be three phonetic-based principles used: identification, mapping, and generating rules. The identification rule analyses phonemes in the input word(s), and the mapping converts those phonemes into characters in the destination word(s) (orthographic representation). The targeted word is generated as a spoken word by the generating rule (letter-to-sound rule).</p>
<p>The conversion rule principles are founded on a number model and are written as <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>:</p>
<p><disp-formula id="eqn-14"><label>(14)</label>
<mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">x</mml:mi><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p>where P (Ws | Is ) is the probability of pronouncing the source word; P (Is | Wt ) is the probability of generating the written Wt from the pronunciation in Is ; P (Wt) represents probability of sequence Wt occurring in the target language.</p>
<p>HMM or ATN can be thought of as a conversion rule that takes the source input (Ws) and maps it to the desired response (Wt) utilizing weight for every transition between stages, indicating which output patterns have a higher probability than the others. P(Wt) is a unigram phrase model that may be developed using any <italic>corpus</italic>. Depending on frequency components, P (Ws | Is) can be approximated.</p>
</sec>
<sec id="s5_7">
<label>5.7</label>
<title>Dataset of SpeakCorrect</title>
<p>The first sample we used is made up of two distinct categories. The first category we gathered from Aljazeera&#x2019;s internet news website. This dataset contains around 140 collected h, of which 100 h were used to develop the SpeakCorrect linguistic model and 40 h were utilized to test the SpeakCorrect method. The second domain is separated into two regions: Saudi Arabia and Egypt. <xref ref-type="table" rid="table-1">Tabs. 1</xref> and <xref ref-type="table" rid="table-2">2</xref> illustrate the layout of our dataset after it was recorded.</p>
<table-wrap id="table-1"><label>Table 1</label>
<caption>
<title>Structure of the dataset 1</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dataset 1</th>
<th colspan="2">Al-Jazeera&#x2019;s news</th>
</tr>
<tr>
<td>Type</td>
<td>Training</td>
<td>Testing</td>
</tr>
</thead>
<tbody>
<tr>
<td>No of hours</td>
<td>100</td>
<td>40</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-2"><label>Table 2</label>
<caption>
<title>Structure of the dataset 2</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dataset 2</th>
<th colspan="2">Saudi</th>
<th colspan="2">Egypt</th>
</tr>
<tr>
<td>Type</td>
<td>Male</td>
<td>Female</td>
<td>Male</td>
<td>Female</td>
</tr>
</thead>
<tbody>
<tr>
<td>No of students</td>
<td>39</td>
<td>&#x2013;</td>
<td>40</td>
<td>30</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Both of the training as well as testing datasets are obtained from native speakers, as evidenced by dataset 1. We discovered that during annotating dataset 2, there are only few sounds.</p>
</sec>
<sec id="s5_8">
<label>5.8</label>
<title>Similarity between Two English Words</title>
<p>Given two word phonemes, W<sub>1</sub>(p<sup>1</sup> p<sup>2</sup> p<sup>3</sup> &#x2026; p<sup>n</sup>) and W<sub>2</sub> (p<sup>1</sup> p<sup>2</sup> p<sup>3</sup> &#x2026; p<sup>n</sup>), three factors are used to describe and evaluate similarity:<list list-type="bullet"><list-item>
<p>The similarity of pronunciation in each phoneme pair ( Pi(W1) , Pj(W2)) between W1 and W2.</p></list-item><list-item>
<p>The similarity between the length of W1 and the length of W2, where 0 &#x003C;&#x003D; i &#x003C;&#x003D; n, 0 &#x003C;&#x003D; j &#x003C;&#x003D; m.</p></list-item><list-item>
<p>The similarity between each character pair (recognized sound character) between W<sub>1</sub> and W<sub>2</sub>.</p></list-item></list></p>
<p>The three factors having different effect in calculating the similarity between two words, we get <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>:</p>
<p><disp-formula id="eqn-15"><label>(15)</label>
<mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="normal">S</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">w</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">W</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>3</mml:mn></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mi>S</mml:mi><mml:msub><mml:mrow><mml:msub><mml:mi mathvariant="normal">w</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">S</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">w</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">W</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>In our case, w<sub>1</sub> is selected to be 0.5, w<sub>2</sub> equals 1 and w<sub>3</sub> is selected to be 1.</p>
</sec>
<sec id="s5_9">
<label>5.9</label>
<title>SpeakCorrect Error Correction</title>
<p>As a result, a kernel characteristic model matrix is applied to compute the correlation between two words in order to correct the speech sequence (Syllable words). The confusion rating from Wi to Wj may be derived as the averaged confusion of every speech segment Si labeled as Wi to trained HMM model of Wj, Awj, given two words Wi and Wj.</p>
<p>Consequently, the confusion similarity score from Wi to Wj is estimated as the following equation <xref ref-type="disp-formula" rid="eqn-16">Eqs. (16)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-18">(18)</xref>:</p>
<p><disp-formula id="eqn-16"><label>(16)</label>
<mml:math id="mml-eqn-16" display="block"><mml:mrow><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">O</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">O</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:math>
</disp-formula></p>
<p>The Guassian kernel function can be applied to calculate the confusion score between Wi and Wj.</p>
<p><disp-formula id="eqn-17"><label>(17)</label>
<mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">f</mml:mi></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">x</mml:mi><mml:mi mathvariant="normal">p</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">W</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi>&#x03C3;</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where &#x03C3; represents the variance calculated over the distribution Sim (Wi , Wj).</p>
<p>The best correction result can be calculated according to the following equation:</p>
<p><disp-formula id="eqn-18"><label>(18)</label>
<mml:math id="mml-eqn-18" display="block"><mml:mrow><mml:mi mathvariant="normal">C</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">x</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">P</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">W</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">E</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">W</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi mathvariant="normal">P</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">W</mml:mi></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="normal">S</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p>where P(W) represents the word language model for the corrected word sequence W. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates the proposed SpeakCorrect correction system.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>SpeakCorrect errors&#x2019; correction</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-4.png"/>
</fig>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Implementation and Testing</title>
<p>The following developing model is based on components to detect error for pronunciation analysis and pronunciation adaption.</p>
<sec id="s6_1">
<label>6.1</label>
<title>The SpeakCorrect Corpus Architecture</title>
<p>The <italic>SpeakCorrect</italic> corpus is based on annotated speech; it will be designed to provide acoustic for supporting the development and evaluation of automatic speech recognition systems.</p>
<sec id="s6_1_1">
<label>6.1.1</label>
<title>The SpeakCorrect Structure</title>
<p>SpeakCorrect, such as the Brown <italic>Corpus</italic>, contains a broad range of dialects, speakers, and texts. It has two primary accent zones, each with two dialect localities, 150 men and women presenters ranging in age (18&#x2013;21 years old) and academic background, and every reads 390 specially selected words. The terms were selected to be phonetically rich as well as to address all of the Arabic participants&#x2019; pronunciation flaws (substitution, deletion, and insertion) (Saudi and Egypt regions). Furthermore, the design employs numerous speakers saying the same words in order to allow comparison among speakers, as well as having a vast variety of terms to ensure maximum coverage of flaws. As a result of the speakers reading by region (and two locales), 150 hundred captured statements are saved in the <italic>corpus</italic>, and each data file has inner structure, as illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Structure of speakcorrect structure</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-5.png"/>
</fig>
<p>With the help of <xref ref-type="table" rid="table-1">Tabs. 1</xref> and <xref ref-type="table" rid="table-2">2</xref>, and <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-18">(18)</xref>, Every item has a phonetic transcription, as well as the associated word tokens, which can be retrieved.</p>
</sec>
<sec id="s6_1_2">
<label>6.1.2</label>
<title>The SpeakCorrect Design Features</title>
<p>As illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, SpeakCorrect incorporates <italic>corpus</italic> design elements. Firstly, such a <italic>corpus</italic> incorporates description at the phonetic as well as orthographic layers, with various labelling techniques at each level. A second property of SpeakCorrect is its balancing across various dimensions of variance, to encompass accent and accent areas and places, which aids later use of the <italic>corpus</italic> for applications such as sociolinguistics, which were not intended when the <italic>corpus</italic> was established.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Structure of the implemented speakcorrect <italic>corpus</italic></title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-6.png"/>
</fig>
</sec>
<sec id="s6_1_3">
<label>6.1.3</label>
<title>The SpeakCorrect Data Acquisition</title>
<p>The internet is a great repository of information for so many natural language processing applications. Consequently, in our scenario, a huge number of data samples are required to obtain. As a result, one of these ways is to get available data from the internet. The benefit of employing such well-defined online data is that it allows for documented, consistent, and repeatable investigation.</p>
</sec>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>The SpeakCorrect User Interface</title>
<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> depicts the computerized user interface layout of SpeakCorrect, which is separated into three tiers. The presenting tier is at the top; the logical or commercial tier is in the centre; it begins with registration, in which the login occurs, microphone configuration and adaption, language and speech lessons, and ultimately assessment. The third tier is internal, and it houses all of the attributes, records, files, and so on.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>The different tiers of the speakcorrect system</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-7.png"/>
</fig>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>The Speak Correct Testing</title>
<p>Because of historical issues with English pronunciation in Arabian accents, the accompanying testing methodology is built on components to evaluate as well as assist students for pronunciation assessment, pronunciation adaptability, and pronunciation detect error. <xref ref-type="fig" rid="fig-8">Fig. 8</xref> depicts the login information for students.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>The login interface of the speakcorrect system</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-8.png"/>
</fig>
<p>As a result, the user interface is built using Silverlight technology. This type of user interface has several visual qualities that allow it to execute basic activities such as going to prior and subsequent demos, streaming a sample (predefined example), human voice, and collecting the user&#x2019;s voice. <xref ref-type="fig" rid="fig-9">Fig. 9</xref> depicts device configuration and microphone modification.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>The device setting and microphone adjustment of the speakcorrect system</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-9.png"/>
</fig>
<p>The operational code includes a cooperation module that connects C# code (.Net Client/Server), HMM code, and HTK element code. The second element, as illustrated in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>, aims at comparing the input voice to the predetermined trained voices and, as a result, provides a mistake-if any.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>The levels and associated lessons testing of the speakcorrect system</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24967-fig-10.png"/>
</fig>
<p>As a result, the SpeakCorrect user experience is built using Silverlight technology. As seen in earlier drawings, such a user interface comprises tabs for basic operations such as going to previous and next demos, playing a sample (predefined example), user voice, and recording the user&#x2019;s voice. As a result, the MVVM design is employed to facilitate the tabs&#x2019; connectivity &#x201C;Click Event.&#x201D; The attributes of such visual elements are constrained in the underlying ViewModel class.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Conclusions</title>
<p>The purpose of this paper is to discuss the development of the SpeakCorrect computerised interface, which could be used for speech correction for non-native English speakers. The data set includes information for two major countries: Saudi Arabia as well as Egypt. Every region is divided by two locality regions. An interactive suggestion system is included in the planned system to encourage individuals to enhance their linguistic skills. SpeakCorrect would be utilised in the future to train and assess pronouncing individuals. As a result, an analytical component should be put in place to examine the phonetic characteristics of the unidentified word and the missing phonemes. As a result, post-testing as well as a comparative study would be conducted. As a result, independent of the speaker&#x2019;s accent or sexual identity, the experiment is constructed and performed to assess the proposed resolution of the SpeakCorrect interface based on phonetically trained database. The performance of the SpeakCorrect interface indicates that the recognition system performed satisfactorily.</p>
</sec>
</body>
<back>
<ack>
<p>The authors thank Science and Technology Unit, King Abdulaziz University for the technical support.</p>
</ack><fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> This project was funded by the National Plan for Science, Technology and Innovation (MAARIFAH)&#x2013;King Abdulaziz City for Science and Technology (KACST)- Kingdom of Saudi Arabia &#x2013; Project Number (10-INF-1406-03).</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I. P.</given-names> <surname>Wiggers</surname></string-name> and <string-name><given-names>L. J. M.</given-names> <surname>Rothkrantz</surname></string-name></person-group>, &#x201C;<article-title>Automatic speech recognition using hidden Markov models</article-title>,&#x201D; <source>Course IN4012TU-Real-time AI &#x0026; Automatische Spraakherkenning</source>, vol. <volume>70</volume>, no. <issue>8</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>20</lpage>, <year>2003</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Noormamode</surname></string-name>, <string-name><given-names>B. G.</given-names> <surname>Rahimbux</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Peerboccus</surname></string-name></person-group>, &#x201C;<article-title>A speech engine for mauritian creole</article-title>,&#x201D; <source>Information Systems Design and Intelligent Application</source>, vol. <volume>36</volume>, no. <issue>7</issue>, pp. <fpage>389</fpage>&#x2013;<lpage>398</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. A.</given-names> <surname>Kemble</surname></string-name></person-group>, &#x201C;<article-title>An introduction to speech recognition</article-title>,&#x201D; <source>Voice Systems Middleware Education-IBM Corporation</source>, vol. <volume>16</volume>, no. <issue>8</issue>, pp. <fpage>154</fpage>&#x2013;<lpage>163</lpage>, <year>2001</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>O.</given-names> <surname>Bracha</surname></string-name></person-group>, &#x201C;<article-title>The folklore of informationalism: The case of search engine speech</article-title>,&#x201D; <source>Fordham Law Review</source>, vol. <volume>82</volume>, no. <issue>5</issue>, pp. <fpage>1629</fpage>&#x2013;<lpage>1633</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Tur</surname></string-name> and <string-name><given-names>R. De</given-names> <surname>Mori</surname></string-name></person-group>, &#x201C;<article-title>Spoken language understanding: Systems for extracting semantic information from speech</article-title>,&#x201D; <source>Fordham Law Review</source>, vol. <volume>82</volume>, no. <issue>5</issue>, pp. <fpage>1629</fpage>&#x2013;<lpage>1633</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Laskowski</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Shriberg</surname></string-name></person-group>, &#x201C;<article-title>Comparing the contributions of context and prosody in text-independent dialog act recognition</article-title>,&#x201D; in <conf-name>Proc. of the 2010 IEEE Int. Conf. on Acoustics, Speech and Signal Processing Newyork</conf-name>, <publisher-loc>USA</publisher-loc>, pp. <fpage>5374</fpage>&#x2013;<lpage>5377</lpage>, <year>2010</year>. </mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y. I.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Y. Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y. C.</given-names> <surname>Ju</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Seltzer</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Tashev</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Voice search of structured media data</article-title>,&#x201D; in <conf-name>Proc. of the 2009 IEEE Int. Conf. on Acoustics, Speech and Signal Processing</conf-name>, <publisher-loc>Delhi, India</publisher-loc>, pp. <fpage>3941</fpage>&#x2013;<lpage>3944</lpage>, <year>2009</year>. </mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Heracleous</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Aboutabit</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Beautemps</surname></string-name></person-group>, &#x201C;<article-title>Lip shape and hand position fusion for automatic vowel recognition in cued speech for French</article-title>,&#x201D; <source>IEEE Signal Processing Letters</source>, vol. <volume>16</volume>, no. <issue>5</issue>, pp. <fpage>339</fpage>&#x2013;<lpage>342</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. J.</given-names> <surname>Osberger</surname></string-name></person-group>, &#x201C;<article-title>Speech intelligibility in the hearing impaired: Research and clinical implications</article-title>,&#x201D; <source>Intelligibility in Speech Disorders</source>, vol. <volume>74</volume>, no. <issue>6</issue>, pp. <fpage>233</fpage>&#x2013;<lpage>265</lpage>, <year>1992</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Tsubota</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Dantsuji</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Kawahara</surname></string-name></person-group>, &#x201C;<article-title>An English pronunciation learning system for Japanese students based on diagnosis of critical pronunciation errors</article-title>,&#x201D; <source>ReCALL</source>, vol. <volume>16</volume>, no. <issue>1</issue>, pp. <fpage>173</fpage>&#x2013;<lpage>188</lpage>, <year>2004</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. G.</given-names> <surname>Almekhlafi</surname></string-name></person-group>, &#x201C;<article-title>The effect of computer assisted language learning on United Arab Emirates English as a foreign language school students&#x2019; achievement and attitude</article-title>,&#x201D; <source>Journal of Interactive Learning Research</source>, vol. <volume>17</volume>, no. <issue>2</issue>, pp. <fpage>121</fpage>&#x2013;<lpage>142</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Huijbregts</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Mclaren</surname></string-name> and <string-name><given-names>D. V.</given-names> <surname>Leeuwen</surname></string-name></person-group>, &#x201C;<article-title>Unsupervised acoustic sub-word unit detection for query-by-example spoken term detection</article-title>,&#x201D; in <conf-name>Proc. of the 2011 IEEE Int. Conf. on Acoustics, Speech and Signal Processing</conf-name>, <publisher-loc>Biging, China</publisher-loc>, pp. <fpage>4436</fpage>&#x2013;<lpage>4439</lpage>, <year>2011</year>. </mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>C. J.</given-names> <surname>Waple</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Kawahara</surname></string-name></person-group>, &#x201C;<article-title>Computer assisted language learning system based on dynamic question generation and error prediction for automatic speech recognition</article-title>,&#x201D; <source>Speech Communication</source>, vol. <volume>51</volume>, no. <issue>10</issue>, pp. <fpage>995</fpage>&#x2013;<lpage>1005</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N. T.</given-names> <surname>Vu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Klose</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Mihaylova</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Schultz</surname></string-name></person-group>, &#x201C;<article-title>Improving asr performance on non-native speech using multilingual and crosslingual information</article-title>,&#x201D; in <conf-name>Proc. of the Fifteenth Annual Conf. of the Int. Speech Communication Association</conf-name>, <publisher-loc>LA, USA</publisher-loc>, pp. <fpage>200</fpage>&#x2013;<lpage>203</lpage>, <year>2014</year>. </mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N. A. A.</given-names> <surname>Kadir</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Sudirman</surname></string-name></person-group>, &#x201C;<article-title>Vowel effects towards dental arabic consonants based on spectrogram</article-title>,&#x201D; in <conf-name>Proc. of the 2011 Second Int. Conf. on Intelligent Systems, Modelling and Simulation</conf-name>, <publisher-loc>NY, USA</publisher-loc>, pp. <fpage>183</fpage>&#x2013;<lpage>188</lpage>, <year>2011</year>. </mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Fujii</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Saitoh</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Oka</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Muneyasu</surname></string-name></person-group>, &#x201C;<article-title>Acoustic echo cancellation algorithm tolerable for double talk</article-title>,&#x201D; in <conf-name>Proc. of the 2008 Hands-Free Speech Communication and Microphone Arrays</conf-name>, <publisher-loc>NY, USA</publisher-loc>, pp. <fpage>200</fpage>&#x2013;<lpage>203</lpage>, <year>2008</year>. </mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Kacur</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Rozinaj</surname></string-name></person-group>, &#x201C;<article-title>Adding voicing features into speech recognition based on HMM in Slovak</article-title>,&#x201D; in <conf-name>Proc. of the 2009 16th Int. Conf. on Systems, Signals and Image Processing</conf-name>, <publisher-loc>Bhopal, India</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>, <year>2009</year>. </mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. K.</given-names> <surname>Yopp</surname></string-name> and <string-name><given-names>R. H.</given-names> <surname>Yopp</surname></string-name></person-group>, &#x201C;<article-title>Supporting phonemic awareness development in the classroom</article-title>,&#x201D; <source>Reading Teacher</source>, vol. <volume>54</volume>, no. <issue>2</issue>, pp. <fpage>130</fpage>&#x2013;<lpage>143</lpage>, <year>2000</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Treiman</surname></string-name></person-group>, &#x201C;<article-title>Onsets and rimes as units of spoken syllables: Evidence from children</article-title>,&#x201D; <source>Journal of Experimental Child Psychology</source>, vol. <volume>39</volume>, no. <issue>1</issue>, pp. <fpage>161</fpage>&#x2013;<lpage>181</lpage>, <year>1985</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. J.</given-names> <surname>Hazen</surname></string-name>, <string-name><given-names>I. L.</given-names> <surname>Hetherington</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Shu</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Livescu</surname></string-name></person-group>, &#x201C;<article-title>Pronunciation modeling using a finite-state transducer representation</article-title>,&#x201D; <source>Speech Communication</source>, vol. <volume>46</volume>, no. <issue>2</issue>, pp. <fpage>189</fpage>&#x2013;<lpage>203</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G. Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Z. H.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>G. J.</given-names> <surname>Hwang</surname></string-name></person-group>, &#x201C;<article-title>Developing multi-dimensional evaluation criteria for English learning websites with university students and professors</article-title>,&#x201D; <source>Computers &#x0026; Education</source>, vol. <volume>56</volume>, no. <issue>1</issue>, pp. <fpage>65</fpage>&#x2013;<lpage>79</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Young</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Evermann</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Gales</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Hain</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kershaw</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>The HTK book</article-title>,&#x201D; <source>Cambridge University Engineering Department</source>, vol. <volume>3</volume>, no. <issue>175</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2002</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Kurian</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Balakriahnan</surname></string-name></person-group>, &#x201C;<article-title>Continuous speech recognition system for Malayalam language using PLP cepstral coefficient</article-title>,&#x201D; <source>Journal of Computing and Business Research</source>, vol. <volume>3</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>24</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. T. J.</given-names> <surname>Ansari</surname></string-name> and <string-name><given-names>N. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Worldwide Covid-19 vaccines sentiment analysis through twitter content</article-title>,&#x201D; <source>Electronic Journal of General Medicine</source>, vol. <volume>18</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>15</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Mishra</surname></string-name>, <string-name><given-names>C. N.</given-names> <surname>Bhende</surname></string-name> and <string-name><given-names>B. K.</given-names> <surname>Panigrahi</surname></string-name></person-group>, &#x201C;<article-title>Detection and classification of power quality disturbances using S-transform and probabilistic neural network</article-title>,&#x201D; <source>IEEE Transactions on Power Delivery</source>, vol. <volume>23</volume>, no. <issue>1</issue>, pp. <fpage>280</fpage>&#x2013;<lpage>287</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. T. J.</given-names> <surname>Ansari</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Pandey</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Alenezi</surname></string-name></person-group>, &#x201C;<article-title>STORE: Security threat oriented requirements engineering methodology</article-title>,&#x201D; <source>Journal of King Saud University-Computer and Information Sciences</source>, vol. <volume>54</volume>, no. <issue>5</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>18</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Mogran</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Bourlard</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Hermansky</surname></string-name></person-group>, &#x201C;<article-title>Automatic speech recognition: An auditory perspective</article-title>,&#x201D; <source>Speech Processing in the Auditory System Journal</source>, vol. <volume>68</volume>, no. <issue>12</issue>, pp. <fpage>309</fpage>&#x2013;<lpage>338</lpage>, <year>2004</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Radha</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Vimala</surname></string-name></person-group>, &#x201C;<article-title>A review on speech recognition challenges and approaches</article-title>,&#x201D; <source>Encyclopedias of Journals</source>, vol. <volume>2</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. H.</given-names> <surname>Lee</surname></string-name> and <string-name><given-names>S. M.</given-names> <surname>Siniscalchi</surname></string-name></person-group>, &#x201C;<article-title>An information-extraction approach to speech processing: Analysis, detection, verification, and recognition</article-title>,&#x201D; <source>Proceedings of the IEEE</source>, vol. <volume>101</volume>, no. <issue>5</issue>, pp. <fpage>1089</fpage>&#x2013;<lpage>1115</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sahu</surname></string-name>, <string-name><given-names>F. A.</given-names> <surname>Alzahrani</surname></string-name>, <string-name><given-names>R. K.</given-names> <surname>Srivastava</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Hesitant fuzzy sets based symmetrical model of decision-making for estimating the durability of web application</article-title>,&#x201D; <source>Symmetry</source>, vol. <volume>12</volume>, no. <issue>11</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>20</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. T. J.</given-names> <surname>Ansari</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Baz</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Alhakami</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Alhakami</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>P-STORE: Extension of store methodology to elicit privacy requirements</article-title>,&#x201D; <source>Arabian Journal for Science and Engineering</source>, vol. <volume>64</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>24</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Durability challenges in software engineering</article-title>,&#x201D; <source>CrossTalk</source>, vol. <volume>42</volume>, no. <issue>4</issue>, pp. <fpage>29</fpage>&#x2013;<lpage>31</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sahu</surname></string-name>, <string-name><given-names>F. A.</given-names> <surname>Alzahrani</surname></string-name>, <string-name><given-names>R. K.</given-names> <surname>Srivastava</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Evaluating the impact of prediction techniques: Software reliability perspective</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>67</volume>, no. <issue>2</issue>, pp. <fpage>1471</fpage>&#x2013;<lpage>1488</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Attaallah</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ahmad</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Tarique</surname></string-name>, <string-name><given-names>A. K.</given-names> <surname>Pandey</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Device security assessment of internet of healthcare things</article-title>,&#x201D; <source>Intelligent Automation &#x0026; Soft Computing</source>, vol. <volume>27</volume>, no. <issue>2</issue>, pp. <fpage>593</fpage>&#x2013;<lpage>603</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sahu</surname></string-name> and <string-name><given-names>R. K.</given-names> <surname>Srivastava</surname></string-name></person-group>, &#x201C;<article-title>Needs and importance of reliability prediction: An industrial perspective</article-title>,&#x201D; <source>Information Sciences Letters</source>, vol. <volume>9</volume>, no. <issue>1</issue>, pp. <fpage>33</fpage>&#x2013;<lpage>37</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sahu</surname></string-name> and <string-name><given-names>R. K.</given-names> <surname>Srivastava</surname></string-name></person-group>, &#x201C;<article-title>Revisiting software reliability</article-title>,&#x201D; <source>Advances in Intelligent Systems and Computing</source>, vol. <volume>802</volume>, pp. <fpage>221</fpage>&#x2013;<lpage>235</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Secure serviceability of software: Durability perspective</article-title>,&#x201D; <source>Communications in Computer and Information Science</source>, vol. <volume>628</volume>, pp. <fpage>104</fpage>&#x2013;<lpage>110</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Fuzzy analytic hierarchy process for software durability: Security risks perspective</article-title>,&#x201D; <source>Advances in Intelligent Systems and Computing</source>, vol. <volume>508</volume>, pp. <fpage>469</fpage>&#x2013;<lpage>478</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Revisiting software security: Durability perspective</article-title>,&#x201D; <source>International Journal of Hybrid Information Technology</source>, vol. <volume>8</volume>, no. <issue>2</issue>, pp. <fpage>311</fpage>&#x2013;<lpage>322</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Alosaimi</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alharbi</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Alyami</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ahmad</surname></string-name>, <string-name><given-names>A. K.</given-names> <surname>Pandey</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Impact of tools and techniques for securing consultancy services</article-title>,&#x201D; <source>Computer Systems Science and Engineering</source>, vol. <volume>37</volume>, no. <issue>3</issue>, pp. <fpage>347</fpage>&#x2013;<lpage>360</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sahu</surname></string-name> and <string-name><given-names>R. K.</given-names> <surname>Srivastava</surname></string-name></person-group>, &#x201C;<article-title>&#x2018;Predicting software bugs of newly and large datasets through a unified neuro-fuzzy approach: Reliability perspective</article-title>,&#x201D; <source>Advances in Mathematics: Scientific Journal</source>, vol. <volume>10</volume>, no. <issue>1</issue>, pp. <fpage>543</fpage>&#x2013;<lpage>555</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Revisiting software security risks</article-title>,&#x201D; <source>Journal of Advances in Mathematics and Computer Science</source>, vol. <volume>11</volume>, no. <issue>6</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Analytical network process for software security: A design perspective</article-title>,&#x201D; <source>CSI Transactions on ICT</source>, vol. <volume>4</volume>, no. <issue>2</issue>, pp. <fpage>255</fpage>&#x2013;<lpage>258</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sahu</surname></string-name> and <string-name><given-names>R. K.</given-names> <surname>Srivastava</surname></string-name></person-group>, &#x201C;<article-title>Soft computing approach for prediction of software reliability</article-title>,&#x201D; <source>ICIC Express Letters</source>, vol. <volume>12</volume>, no. <issue>12</issue>, pp. <fpage>1213</fpage>&#x2013;<lpage>1222</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Software security testing: A pertinent framework</article-title>,&#x201D; <source>Journal of Global Research in Computer Science</source>, vol. <volume>5</volume>, no. <issue>3</issue>, pp. <fpage>23</fpage>&#x2013;<lpage>27</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F. A.</given-names> <surname>Alzahrani</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ahmad</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Nadeem</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Integrity assessment of medical devices for improving hospital services</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>67</volume>, no. <issue>3</issue>, pp. <fpage>3619</fpage>&#x2013;<lpage>3633</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Durable security in software development: Needs and importance</article-title>,&#x201D; <source>CSI Communications</source>, vol. <volume>10</volume>, no. <issue>10</issue>, pp. <fpage>34</fpage>&#x2013;<lpage>36</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Agrawal</surname></string-name> and <string-name><given-names>R. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Measuring the security attributes through fuzzy analytic hierarchy process: Durability perspective</article-title>,&#x201D; <source>ICIC Express Letters</source>, vol. <volume>12</volume>, no. <issue>6</issue>, pp. <fpage>615</fpage>&#x2013;<lpage>620</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Alyami</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Nadeem</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Alosaimi</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alharbi</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Analyzing the data of software security life-span: Quantum computing era</article-title>,&#x201D; <source>Intelligent Automation &#x0026; Soft Computing</source>, vol. <volume>31</volume>, no. <issue>2</issue>, pp. <fpage>707</fpage>&#x2013;<lpage>716</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. H.</given-names> <surname>Seh</surname></string-name>, <string-name><given-names>J. F.</given-names> <surname>Alamri</surname></string-name>, <string-name><given-names>A. F.</given-names> <surname>Subahi</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Agrawal</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Machine learning based framework for maintaining privacy of healthcare data</article-title>,&#x201D; <source>Intelligent Automation &#x0026; Soft Computing</source>, vol. <volume>29</volume>, no. <issue>3</issue>, pp. <fpage>697</fpage>&#x2013;<lpage>712</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. H.</given-names> <surname>Seh</surname></string-name>, <string-name><given-names>J. F.</given-names> <surname>Alamri</surname></string-name>, <string-name><given-names>A. F.</given-names> <surname>Subahi</surname></string-name>, <string-name><given-names>M. T. J.</given-names> <surname>Ansari</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Hybrid computational modeling for web application security assessment</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>70</volume>, no. <issue>1</issue>, pp. <fpage>469</fpage>&#x2013;<lpage>489</lpage>, <year>2022</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>