<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CSSE</journal-id>
<journal-id journal-id-type="nlm-ta">CSSE</journal-id>
<journal-id journal-id-type="publisher-id">CSSE</journal-id>
<journal-title-group>
<journal-title>Computer Systems Science &#x0026; Engineering</journal-title>
</journal-title-group>
<issn pub-type="ppub">0267-6192</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">37113</article-id>
<article-id pub-id-type="doi">10.32604/csse.2023.037113</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Visual Lip-Reading for Quranic Arabic Alphabets and Words Using Deep Learning</article-title>
<alt-title alt-title-type="left-running-head">Visual Lip-reading for Quranic Arabic Alphabets and Words using Deep Learning</alt-title>
<alt-title alt-title-type="right-running-head">Visual Lip-reading for Quranic Arabic Alphabets and Words using Deep Learning</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Aljohani</surname><given-names>Nada Faisal</given-names></name><email>naljohani0084@stu.kau.edu.sa</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Jaha</surname><given-names>Emad Sami</given-names></name></contrib>
<aff><institution>Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University</institution>, <addr-line>Jeddah, 21589</addr-line>, <country>Saudi Arabia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Nada Faisal Aljohani. Email: <email>naljohani0084@stu.kau.edu.sa</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>31</day>
<month>3</month>
<year>2023</year>
</pub-date>
<volume>46</volume>
<issue>3</issue>
<fpage>3037</fpage>
<lpage>3058</lpage>
<history>
<date date-type="received">
<day>24</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>12</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Aljohani and Jaha</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Aljohani and Jaha</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CSSE_37113.pdf"></self-uri>
<abstract>
<p>The continuing advances in deep learning have paved the way for several challenging ideas. One such idea is visual lip-reading, which has recently drawn many research interests. Lip-reading, often referred to as visual speech recognition, is the ability to understand and predict spoken speech based solely on lip movements without using sounds. Due to the lack of research studies on visual speech recognition for the Arabic language in general, and its absence in the Quranic research, this research aims to fill this gap. This paper introduces a new publicly available Arabic lip-reading dataset containing 10490 videos captured from multiple viewpoints and comprising data samples at the letter level (i.e., single letters (single alphabets) and Quranic disjoined letters) and in the word level based on the content and context of the book <italic>Al-Qaida Al-Noorania</italic>. This research uses visual speech recognition to recognize spoken Arabic letters (Arabic alphabets), Quranic disjoined letters, and Quranic words, mainly phonetic as they are recited in the Holy Quran according to Quranic study aid entitled Al-Qaida Al-Noorania. This study could further validate the correctness of pronunciation and, subsequently, assist people in correctly reciting Quran. Furthermore, a detailed description of the created dataset and its construction methodology is provided. This new dataset is used to train an effective pre-trained deep learning CNN model throughout transfer learning for lip-reading, achieving the accuracies of 83.3%, 80.5%, and 77.5% on words, disjoined letters, and single letters, respectively, where an extended analysis of the results is provided. Finally, the experimental outcomes, different research aspects, and dataset collection consistency and challenges are discussed and concluded with several new promising trends for future work.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Visual speech recognition</kwd>
<kwd>lip-reading</kwd>
<kwd>deep learning</kwd>
<kwd>quranic Arabic dataset</kwd>
<kwd>Tajwid</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Human language is the most fundamental means of communication. Therefore, artificial intelligence algorithms have been developed to use spoken speech in different languages for automatic speech recognition. Speech recognition transforms spoken language into data bits that can be utilized for various applications. The field of automatic speech recognition has been gradually and increasingly a hot research area at different levels. It began with only audio speech recognition (ASR), then audio-visual speech recognition (AVSR), followed by continuous speech recognition (CSR), and finally, as the most challenging task among all its counterparts, performing only visual speech recognition (VSR). Lip-reading is a desired capability that can play an important role in human-computer interaction. In computer vision, recognizing of visual speech is a common complex research topic, as lip-reading can be affected by many factors such as lighting, age, makeup, viewpoint angle, and variation of lip shape.</p>
<p>On the other hand, a visual speech recognition system is not affected by surrounding noise because it does not require sounds for processing, as it is limited only to visual information processing. Lip-reading, or visual speech recognition, can be described as extracting spoken words based on visual information only through lips, tongue, and teeth movements made by the pronunciation of these spoken words. The significance of this research field appears in the possibility of applying the concepts of visual lip-reading to different languages for several applications, such as criminal scrutiny by surveillance cameras, educational evaluation, biometric authentication, and connecting with smart cars. Therefore, many research efforts are increasingly devoted to improving performance and increasing the accuracy of visual speech recognition.</p>
<p>The Arabic language is the language of the Holy Quran, meant to preserve and spread it in service to Islam, facilitating the recitation and understanding of the Quran. The Arabic language has elaborate and important origins, derivations, rules, amplitude, flexibility, and formation, which are not often found in other languages. Its letters are distinguished by their articulation (exits) and sound by harmony. Lip-reading may be one of the most important topics that can be effectively employed in human-computer interaction in multiple languages. However, the Arabic language still needs to be sufficiently studied by researchers, in contrast to other languages.</p>
<p>Recently, most research explorations related to the Holy Quran focused on only audio speech recognition; there has been almost no interest in employing the approaches of visual speech recognition or considerations of the potential improvement it would bring. The lack of Arabic visual speech recognition studies may affect progress and development in research related to the Holy Quran in visual speech recognition. The lack or unavailability of Arabic visual speech datasets is likely to decrease the number of studies in deep learning for the Arabic language. Addressing this problem will have practical benefits for studies based on Arabic and contribute to progress in Quranic research areas.</p>
<p>Deep learning is an emanative part of artificial intelligence and is designed in such a way as to simulate the strategies of the human brain. Its architecture contains several layers used for processing, similar to neurons. Deep learning algorithms have been found superior at most recognition tasks and are considered a vast improvement compared to traditional methods in various automatic recognition applications. Recently, some research studies have applied deep learning techniques to lip-reading and proved that they could be used successfully in VSR. Due to the tremendous success achieved by deep learning methods in many fields, we aim is to visually recognize Arabic speech using deep learning techniques based on the rules of the <italic>Al</italic>-Qaida <italic>Al-Noorania book</italic> [<xref ref-type="bibr" rid="ref-1">1</xref>] (<bold><italic>QNbook</italic></bold>), which is used in learning Quran recitation alongside <bold>Tajwid</bold>. <bold>Tajwid</bold> is a science that contains a set of rules that helps reciters of the Holy Quran pronounce the letters with the correct articulation (<bold>Makharij</bold>) and gives each letter its distinctive adjectives (<bold>Sefat</bold>). <italic>QNbook</italic> is considered one of the easiest and most useful means of teaching the pronunciation of the Quranic words as the Prophet Muhammad; peace be upon him, recited them. By learning the smallest building block of the Quran, the letter, whoever masters it can read the Holy Quran by spelling without difficulty because of the author&#x0027;s precision and care in collecting it. The author started gradually with single letters, compound letters, disjoined letters, vowels, formation, and the provisions of Tajwid. Because of the importance of <italic>QNbook</italic>, the fruit of which is correct and eloquent pronunciation and a distinct ability to read Arabic in general and the Quran in particular, we worked on building a dataset that will benefit researchers and those interested in studies related to the Holy Quran in several areas. The most important of these is artificial intelligence to build promising studies in the future and spread its benefit. In this work, the main contributions are as follows:
<list list-type="simple">
<list-item><label>&#x25A0;</label><p>Building, to the best of our knowledge, the first new public dataset designed for audio-visual speech recognition purposes in the Arabic language according to the <italic>Al-Qaida Al-Noorania book</italic> (<italic>QNbook</italic>) from three different face/mouth angles (viewpoints). We called it Al-Qaida Al-Noorania Dataset (AQAND) for inducing a variety of further novel and effective research.</p></list-item>
<list-item><label>&#x25A0;</label><p>Moving the field of speech recognition from audio to visual in studies related to the Holy Quran by training, validating, and testing the pre-trained convolutional neural network (CNN) model using our generated dataset, achieving an accuracy of 83%. This move may help to improve interactive recitation systems and tools as well as systems and tools for learning the recitation of the Holy Quran in the future.</p></list-item>
<list-item><label>&#x25A0;</label><p>Recognizing, to the best of our knowledge, the single letters (Horof Alheja&#x0101; Almufradah) and disjoined letters (Horof Almuqata&#x0101;) for the first time, using lip-reading and reliance on the classical Arabic language. At the word level, words from the Holy Quran were used rather than from various colloquial dialects used in many Arab nations. As far as we know, this has never been done before in visual speech recognition.</p></list-item>
</list></p>
<p>The paper is organized in the following manner: the related work is presented in Section 2; the AQAND&#x2019;s design, its pre-processing details, and the recommended classification model are all described in Section 3; Section 4 of the paper discusses the experiments and results; finally, Section 5 presents the conclusion and future work.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>In the last few years, many researchers have been working on VSR. Before the advent of deep learning, researchers used various methods and algorithms in machine learning for lip-reading. Where numerous research works in lip-reading were based on hand-engineered features that are usually modeled by hidden Markov model (HMM) based pipelines, as shown in [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>]. There has been a significant advancement in lip-reading techniques during the past ten years. Many studies initially centered on 2D fully convolutional networks [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. But as hardware improved, 3D convolutions over 2D convolutions [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>] or recurrent neural networks [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>] quickly became an option for more effective use of temporal information. This concept has developed into a specific architecture consisting of two parts: the front end and the back end. In this process, the final temporal information is summed using recurrent layers as a backend after the local lip movement information has been extracted as a frontend using a 3D &#x002B; 2D convolutions backbone [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. This proposed architecture [<xref ref-type="bibr" rid="ref-12">12</xref>] has been very effective. It has achieved a 17.5% recognizable improvement in accuracy in the lip-reading datasets LRW [<xref ref-type="bibr" rid="ref-13">13</xref>] and LRW-1000 [<xref ref-type="bibr" rid="ref-14">14</xref>], and even now, many state-of-the-art lip-reading solutions still use it as its meta-architecture [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>]. Following the level of recognized speech, in this section, we present related work in the current research area by distributing it into four sub-sections:</p>
<sec id="s2_1">
<label>2.1</label>
<title>Recognition of Arabic at the Letter Level</title>
<p>Arabic letter-level audio speech recognition researchers have achieved more than 99% accuracy [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. Other than studies on visual speech recognition, most studies focus on numbers, words, and sentences rather than considering the alphabet, which forms the foundation of any language. Nevertheless, it is a must when learning any language to know the proper pronunciation of that language&#x0027;s alphabet using the correct articulations and exits. However, in [<xref ref-type="bibr" rid="ref-21">21</xref>], researchers classified the alphabet of the Arabic language into ten visemes (different letters having similar lips movements during pronunciation) and established its viseme mapping for four speakers. They based their study on an analysis of geometrical features of the face extracted from the front lip movement. As a result, they demonstrated that it is possible to recognize a vowel and determine whether it is a short vowel (i.e., fataha, damma, or kasra) or a long one. Also, the research of [<xref ref-type="bibr" rid="ref-22">22</xref>] presented an Arabic viseme system that classified the language into 11 visemes based on two methods: the statistical parameters for lip image and the geometrical parameters for the internal and external lip contour using multilayer perceptron (MLP) neural networks.</p>
<p>On the other hand, the work described in [<xref ref-type="bibr" rid="ref-23">23</xref>] is closely related to ours. The authors researched the mouth shapes of the alphabet during its pronunciation with Tajwid to find out the differences between its pronunciation from the geometry and sequence of movements of the lip and divided the 28 letters into five groups based on those movements. They introduced a lip tracking system to extract lip movement data from a single professional reciter and compare it with novice users&#x2019; lip movements to verify their pronunciation&#x2019;s correctness through a graphical user interface (GUI) that they designed and programmed. Their study depends on calculating the displacement between the height and width of the lips for each video frame and plotting it in a graph using machine learning algorithms. As noted in the literature, there need to be more visual speech datasets for the Arabic alphabet that will open promising avenues of research in the future. In AQAND, we included Arabic and the 14 Quranic letters as novel content in letter-level recognition.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Recognition of Arabic at the Word Level</title>
<p>Most studies focused on lip-reading at the word level have achieved high results. In reference [<xref ref-type="bibr" rid="ref-3">3</xref>], the researchers proposed a novel approach that aims to detect Arabic consonant-vowel letters in a word using two HMM-based classifiers: one for the recognition of the &#x201C;consonant part&#x201D; and one for the &#x201C;vowel part.&#x201D;. They tested their method on 20 words collected from 4 speakers. Their algorithm scored an accuracy of 81.7%. In addition to that study, [<xref ref-type="bibr" rid="ref-24">24</xref>] collected 1100 videos of 10 Arabic words from 22 speakers, calling it the Arabic visual speech dataset (AVSD). The authors cropped the mouth region manually from video frames, and the support vector machine (SVM) model was used to evaluate the AVSD with a 70% word recognition rate (WRR). Also, [<xref ref-type="bibr" rid="ref-25">25</xref>] presented the read my lips (RML) system, an Arabic word lip-reading system. The authors collected dataset videos of 10 commonly used Arabic words from 73 speakers, with RGB and grayscale versions of each. They trained and tested the dataset using three different deep-learning models, as listed in <xref ref-type="table" rid="table-2">Table 2</xref>. The RGB version of the dataset obtained higher accuracy than the grayscale. They also suggested a voting model for the three models in the RML system to improve the overall accuracy, as they succeeded and achieved an accuracy of 82.84%, which is 3.64% higher than the highest accuracy obtained for one of the three models. This accuracy was the highest Arabic word prediction accuracy in any related work. However, in our Arabic Quranic word dataset, we scored a higher accuracy by 0.5%.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Recognition of Arabic at the Sentence Level</title>
<p>At this level, only a few studies have been conducted on lip-reading. Over the last two decades, researchers have proposed for the first time a novel Arabic lip-reading system by combining the hypercolumn model (HCM) with the HMM [<xref ref-type="bibr" rid="ref-2">2</xref>]. They used HCM to extract the relevant features and HMM for feature sequence recognition. For testing their proposed system, they used nine sentences uttered by 9 Arabic speakers and achieved 62.9% accuracy. Recently, researchers continued to improve the recognition process and performance at this level, as shown in [<xref ref-type="bibr" rid="ref-26">26</xref>]; they presented an Arabic dataset of sentences and numbers for visual speech recognition purposes, which contains 960 sentence videos and 2400 number videos from 24 speakers. They used concatenated frame images (CFIs) as pre-processing for their dataset, resulting in one single image containing utterance sequences, which are fed into their proposed model: visual geometry group (VGG-19) network with batch normalization for feature extraction and classification. They excelled, achieving a competitive accuracy of 94% for numbers prediction and 97% for sentence prediction.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Holy Quran-Related Studies in Speech Recognition</title>
<p>In recent decades, researchers have presented many studies related to the Holy Quran, which are concerned with helping to learn and recite the Holy Quran correctly [<xref ref-type="bibr" rid="ref-27">27</xref>&#x2013;<xref ref-type="bibr" rid="ref-29">29</xref>]. The researchers were also interested in several studies in the field of the seven readings recognition (<bold>Qiraat</bold>) and distinguishing it, such as in [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>]. However, while these studies have shown remarkable performance in audio speech recognition, they have yet to be involved in visual speech recognition (VSR) in Quranic-related work areas. In terms of the <bold>Qiraat</bold>, we take into consideration that there is some confusion between two sciences (<bold>Qiraat science</bold> and <bold>Tajwid science</bold>) which must be clarified. Qiraat science is focused on how some words in verses of the Quran should be pronounced, while Tajwid science is more focused on the letters, their exits, and the attributes they carry. This means that the issue of &#x201C;exits of letters&#x201D; is therefore included under Tajwid science. Additionally, because the issue of exits deals with each letter separately, the variations between Qiraat do not affect this issue. In our work, we left the reciters the freedom of choice in pronouncing the Quranic words during data collection. However, by including the different Qiraat competitive accuracy was achieved.</p>
<p>Eventually, we conclude that deep learning networks have proved to be more efficient compared to other approaches. There is also a need to design and build open-source large-scale Arabic datasets of visual speech recognition, enabling more research efforts on this rich language. So, this research is an invitation to all researchers interested in the Holy Quran and Arabic visual speech recognition to continue and advance our work. Visual speech recognition datasets are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>. <xref ref-type="table" rid="table-2">Table 2</xref> summarizes all Arabic-related work described above.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Arabic and other languages lip-reading datasets statistics</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Language</th>
<th>Dataset</th>
<th>Dataset contents</th>
<th>Number of speakers</th>
<th>Recording environment</th>
<th>Resolution</th>
<th>Open source</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">Arabic</td>
<td>[<xref ref-type="bibr" rid="ref-2">2</xref>] 2004</td>
<td>9 Arabic sentences</td>
<td>9</td>
<td rowspan="5">Lab</td>
<td>160 &#x00D7; 120 pix</td>
<td>No</td>
</tr>
<tr>
<td>AVSD [<xref ref-type="bibr" rid="ref-24">24</xref>] 2019</td>
<td>1100 video samples of 10 daily communication Arabic words</td>
<td>22 (8 males &#x0026; 14 females)</td>
<td>1920 &#x00D7; 1080 pix</td>
<td>No</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>] 2022</td>
<td>1051 video samples of 10 common Arabic words</td>
<td>73 (40 males &#x0026; 33 females)</td>
<td>66 &#x00D7; 100 pix</td>
<td>Yes</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>] 2022</td>
<td>3360 video samples of 10 Arabic digits &#x0026; 4 Arabic sentences</td>
<td>24 (14 males &#x0026; 10 females)</td>
<td>1920 &#x00D7; 1080 pix</td>
<td>Yes</td>
</tr>
<tr>
<td><bold>AQAND (ours)</bold></td>
<td>10490 video samples of 29 Arabic alphabets &#x0026; 14 Quranic letters &#x0026; 10 Quranic words</td>
<td>22 (18 males &#x0026; 4 females)</td>
<td>1920 &#x00D7; 1080 pix</td>
<td>Yes</td>
</tr>
<tr>
<td>English</td>
<td>LRW [<xref ref-type="bibr" rid="ref-13">13</xref>] 2016</td>
<td>500 words</td>
<td>&#x002B;1K</td>
<td>TV</td>
<td>256 &#x00D7; 256 pix</td>
<td>Yes</td>
</tr>
<tr>
<td>Mandarin</td>
<td>LRW-1000 [<xref ref-type="bibr" rid="ref-14">14</xref>] 2019</td>
<td>1000 words</td>
<td>&#x002B;2K</td>
<td>TV</td>
<td>Naturally distributed</td>
<td>Yes</td>
</tr>
<tr>
<td>Russian</td>
<td>LRWR [<xref ref-type="bibr" rid="ref-15">15</xref>] 2021</td>
<td>235 words</td>
<td>135</td>
<td>YouTube</td>
<td>1920 &#x00D7; 1080 pix</td>
<td>Yes</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Recent Arabic VSR-related work</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Reference</th>
<th colspan="2" align="center">Pre-processing Techniques</th>
<th>Features extraction</th>
<th>Classifier</th>
<th>Level recognition</th>
<th>Results and accuracy</th>
</tr>
<tr>
<th/>
<th>Face detection</th>
<th>Lip localization</th>
<th/>
<th/>
<th/>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-2">2</xref>] 2004</td>
<td>N/A</td>
<td>N/A</td>
<td>Hypercolumn model (HCM)</td>
<td>HMM, with five states</td>
<td>9 Arabic sentences</td>
<td>Accuracy of 62.9%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-3">3</xref>] 2011<sup>1</sup></td>
<td>N/A</td>
<td>N/A</td>
<td>Using low-level statistical methods, calculated the pixels of the ROI for extraction of geometrical features in 4 points on lips (W, H, A, D).</td>
<td>HMM with three states</td>
<td>20 Arabic words</td>
<td>Accuracy of 81.7%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-24">24</xref>] 2019</td>
<td>N/A</td>
<td>Manually cropped for ROI</td>
<td>DCT</td>
<td>SVM</td>
<td>10 Arabic words</td>
<td>Word recognition rate of 70%</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>] 2022</td>
<td>Python Dlib library</td>
<td>Dlib facial landmark points and generate two versions: RGB &#x0026; grayscale</td>
<td>Three models: CNN &#x0026; TD &#x002B; LSTM &#x0026; TD &#x002B; BiLSTM</td>
<td>Softmax layer</td>
<td>10 Arabic words</td>
<td>Accuracy of: RGB in CNN (79.2%), grayscale in CNN (76.6%), RGB in TD &#x002B; LSTM (70.1%), grayscale in TD &#x002B; LSTM (67.5%), RGB in TD &#x002B; BiLSTM (74.1%), grayscale in TD &#x002B; BiLSTM (70.1%), and RGB in a voting model (82.8%)</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>] 2022</td>
<td>OpenCV<break/>Library using Dlib toolkit to detect facial landmarks</td>
<td>Use facial landmarks to locate key points of mouth and generate CFI</td>
<td>VGG-19 with batch normalization</td>
<td>Softmax layer</td>
<td>10 Arabic digits and 4 Arabic sentences</td>
<td>Accuracy of digits is 94%, in sentences is 97%, in digits &#x0026; sentences is 93%</td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn><p>Note: <sup>1</sup>Speaker-independent</p></fn></table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>The Al-Qaida Al-Noorania Dataset (AQAND)</title>
<p>Most lip-reading systems involve two phases (analyzing visual information in the input image and transforming this information into corresponding words or sentences) and three parts (lip detection and localization, lip feature extraction, and visual speech recognition). Many studies have given end-to-end models driven by deep neural networks (DNN) to hasten the development of deep learning technologies. These models have shown promising performance outcomes when compared to approaches based on traditional neural networks. In this research, we aim to train and evaluate an effective model for lip-reading through three main phases: AQAND design and collection, AQAND preparation and splitting, and implementation of the classification model.</p>
<sec id="s3_1">
<label>3.1</label>
<title>AQAND Design</title>
<p>In this research investigation, we built our dataset since no datasets were available in our research scope nor suitable to achieve our objectives. Therefore, we introduce a new and publicly available lip-reading dataset, which, to the best of our knowledge, may be the first dataset on <italic>QNbook</italic>. As such, the AQAND is designed and created to be used for training and testing a model&#x2019;s capabilities initially in lip-reading (recognition of the lip movements) and further, in recognizing the correct and accurate exits of the letters during their pronunciation by following the intonation (Tajwid) rules for Quran recitation. The AQAND consists of around 16 h of RGB videos with a resolution of 1920 &#x00D7; 1080 pixels and 30 frames per second, resulting in a total of 10490 RGB video samples, each ranging from two to ten seconds long. The total size of AQAND is approximately 60 GB.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>AQAND Collection</title>
<p>The collected videos of the dataset were sorted into three main categories: Horof Alheja&#x0101; Almufradah (29 single letters), Horof Almuqata&#x0101; (14 disjoined letters), and Quranic words (10 words). Each video was recorded in an indoor environment from three different angles (0&#x00B0;, 30&#x00B0;, and 90&#x00B0;) simultaneously using three digital cameras, with each camera placed within 90 cm of the subject. All videos were obtained from 22 reciters (contributors), 16 men and four women from the ages of 20 to 59. Each reciter repeated the same video recording scenario (for all three categories) in three different sessions. In addition, there were several individual differences between reciters, including geometric features of the lip, beard, mustache, movements, and shape of mouth, teeth, and the alveolar ridge. There were also observable variations that occurred between the same repeated samples over the three recording sessions from the same reciters, which helped in producing unbiased data sampling and can enable data augmentation for deep learning.</p>
<p>The proposed setup and recoding scenario were to allow the following of Tajwid rules based on the <italic>QNbook</italic> method, where the reciters were asked to read letters and words correctly, clearly, and accurately and wait for five seconds (a silence between each letter/word) before moving to the next item. After taking the videos, we entirely reviewed each video to verify their authenticity and to make sure that our instructions were correctly followed. As such, the videos were collected for the three main categories (<italic>QNbook</italic> lessons as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>) and included in the AQAND as follows:</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The three main categories contained in the AQAND as presented in <italic>QNbook</italic>: (a) Single alphabets (Horof Alheja&#x0101; Almufradah), (b) Disjoined letters (Horof Almuqata&#x0101;), and (c) The selected ten Quranic words that are highlighted in red boxes</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-1.tif"/>
</fig>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Face and mouth detection in video frames captured from three viewpoints (0&#x00B0;, 30&#x00B0;, and 90&#x00B0;)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The sequence frames of pronunciation of the word &#x2018;Lawoh&#x2019; from 0&#x00B0; (frontend), 30&#x00B0;, and 90&#x00B0; (profile), up-to-down, respectively, where a voice letter is written under each corresponding frame</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-3.tif"/>
</fig>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Arabic Single Alphabets Dataset</title>
<p>The first and core lesson in <italic>QNbook</italic> is Horof Alheja&#x0101; Almufradah (29 single letters), as in <xref ref-type="fig" rid="fig-1">Fig. 1a</xref>, where their pronunciation is clarified with English letters in <xref ref-type="table" rid="table-3">Table 3</xref>. A reciter reads the single letters as written in red above each letter on the <italic>QNbook</italic> page, shown in <xref ref-type="fig" rid="fig-1">Fig. 1a</xref>. This means they utter it independently, as the letter&#x2019;s conventional full name and not the letter&#x2019;s isolated sound (e.g., for &#x2018;<inline-graphic xlink:href="CSSE_37113-inline-1.tif"/>,&#x2019; saying &#x201C;alif&#x201D; and not &#x201C;a&#x2019;a&#x201D;) as per the <italic>QNbook</italic> rules. This dataset consists of 5742 videos.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The Arabic letters in isolated forms, their corresponding forms in English, the letters in Arabic script, and the pronunciation of single Arabic letters in English. (Note that it is difficult to write the exact spelling of some Arabic letters as there are no similar phonemes to them in the English language)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="center"/>
</colgroup>
<tbody>
<tr>
<td><inline-graphic xlink:href="CSSE_37113-inline-2.tif"/></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Disjoined Letters Dataset</title>
<p>We also shed light on the third lesson from the lessons of <italic>QNbook</italic>, which is Horof Almuqatah (disjoined letters or Quranic letters), as shown in <xref ref-type="fig" rid="fig-1">Fig. 1b</xref>. They are the fourteen combinations of disjoined letters stated in the Quran at the beginning of some surahs, and they are read as individual letters with no silence between them. For example, (<inline-graphic xlink:href="CSSE_37113-inline-3.tif"/>) is read connectedly as &#x2018;alif, lam, meem.&#x2019;. The reciters were asked to correctly perform the extension of letters&#x2019; vowel segments if they are marked above by the (&#x007E;) sign, according to the correct length (or duration, known in Tajwid science as the number of moves; in this case, six moves) for the marked letter. <xref ref-type="table" rid="table-4">Table 4</xref> clarifies the extended letter pronunciation by six moves in English by repeating the extended vowel letter, such as using &#x2018;Qaaaaaaf&#x2019; for the extended letter (<inline-graphic xlink:href="CSSE_37113-inline-4.tif"/>) instead of &#x2018;Qaf&#x2019; as used for the unmarked standard letter (<inline-graphic xlink:href="CSSE_37113-inline-5.tif"/>), shown in <xref ref-type="table" rid="table-3">Table 3</xref>. This dataset consists of 2772 videos.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Extend letter pronunciation by six moves of Arabic disjoined letters written in English. (Note that it is difficult to write the exact spelling of some Arabic letters because there are no similar phonemes to them in the English language)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CSSE_37113-inline-6.tif"/></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Quranic Words Dataset</title>
<p>In this part, we have collected data for ten Quranic words from the sixth and ninth lessons of <italic>the QNbook</italic>. They were randomly chosen from among several Quranic words. A reciter was asked to read the ten Quranic words correctly with diacritics (signs above or under letters, which affect letter pronunciation as short vowels), as shown in the red boxes in <xref ref-type="fig" rid="fig-1">Fig. 1c</xref>. <xref ref-type="table" rid="table-5">Table 5</xref> shows the pronunciation and meaning of the ten chosen Arabic Quranic words as written in English. This dataset consists of 1980 videos.</p>
<table-wrap id="table-5"><label>Table 5</label>
<caption>
<title>Pronunciation and meaning of the ten Quranic words in the dataset</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CSSE_37113-inline-7.tif"/></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Pre-Processing the AQAND</title>
<p>In lip-reading, pre-processing is a necessary step, primarily affecting the validity and accuracy of recognition tasks. In this section, we present the conducted pre-processing steps. First, during pronunciation, each letter or word may have a different length. Therefore, the length of each video sample in the dataset is reformed to exactly 60, 80, or 300 frames, and each video is ensured to accurately contain the complete visual representation of the target letter or word. The number of frames is selected to range from 2 to 10 seconds because most reciters are observed to spend time within this range to complete an utterance of one letter or word. In a few instances, when letter/word utterance is performed faster than this predetermined period, the whole video is concatenated with additional black frames by using the zeros function to compensate for the missing frames. Thus, we get inputs that have a fixed sequence of frames of 60, 80, and 300 and durations of 2, 2.5, and 10 s for single letters, Quranic words, and disjoined letters, respectively.</p>
<p>Second, if the data samples include unneeded details, this may adversely affect the lip-reading task. To combat this, the AQAND is designed to present a region of interest (ROI) in videos, such that the data sample is pre-processed using the Haar-cascade of OpenCV to detect and extract the ROI (in this case, the mouth region) from each video, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. In visual speech recognition, the pre-processing works to accurately look for the position of the lip needed to adjust for any challenges, such as variations in positions of the face and mouth, variation in illumination, and image shadow, all of which can affect the quality of lip video frames. Therefore, a lip-reading system relying on accurate localization of the lip is more likely to achieve higher accuracy in visual speech recognition. A few samples of video frames are presented in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, which display the sequence of visemes of the word &#x2018;Lawoh&#x2019; for the same reciter from three angles (0&#x00B0;, 30&#x00B0;, and 90&#x00B0;).</p>
<p>Third, to decrease the size of the dataset while maintaining reliable data quality, each frame is resized to 88 &#x00D7; 88 pixels, which is the input size used to train our model. Then, the resized frame segments are converted to grayscale, and their pixel values are normalized to a range between 0 and 1.</p>
<p>A cross-validation technique with a ratio of 70:30 had used to split normalized data samples into three sets for training, validation, and testing. Note that this work is conducted based on speaker-independent experiments (i.e., it seeks to identify anyone&#x2019;s lip movements for certain spoken letters/words, regardless of the speaker). All reciters are exclusively distributed into the three groups, everyone in only one of the groups, based on the variation between reciters in the same group in terms of skin color, gender, and age. The validation set is used to fine-tune the model&#x2019;s weights and biases, enhance performance, and prevent overfitting, while the training set is used to train and fit the deep neural network model. Unseen data (i.e., unseen reciters) is utilized in the testing set to assess the model&#x2019;s ability to recognize and classify spoken letters or words. The training set consists of 7630 videos, the validation set constitutes 1430 videos, and the testing set comprises 1430 videos. The size of the training and validation sets is twice doubled using data augmentation techniques, including horizontal flipping and an affine transformation (i.e., training and validation sets end with 22890 and 4290 videos, respectively). Finally, the pre-processed and labeled dataset is fed to the network model as an input with dimensions 88 &#x00D7; 88 &#x00D7; 1, as will be further discussed in Subsection 3.4. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> summarizes all AQAND pre-processing steps described above. The deep learning model in use automatically extracts features from image frames that include lip movements and then uses those features to classify the input video of spoken letters or words. Thus, the model can be utilized to test and classify an unseen input video of a completely unseen reciter. The proposed AQAND specifications are summarized in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>An overview of the AQAND dataset specifications</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Specification aspect</th>
<th>Factor of specification</th>
<th>Specifications</th>
</tr>
</thead>
<tbody>
<tr>
<td>Language</td>
<td>Language</td>
<td>Arabic</td>
</tr>
<tr>
<td/>
<td>Type of utterance</td>
<td>Single letters &#x0026; disjoined letters &#x0026; Quranic words</td>
</tr>
<tr>
<td/>
<td>Number of utterances</td>
<td>53 utterance classes</td>
</tr>
<tr>
<td>Reciters (contributors)</td>
<td>Number of reciters</td>
<td>22 reciters</td>
</tr>
<tr>
<td/>
<td>Gender</td>
<td>18 males and four females</td>
</tr>
<tr>
<td/>
<td>Number of repetition of utterances</td>
<td>Three times</td>
</tr>
<tr>
<td/>
<td>Face view angle of the reciter (viewpoint)</td>
<td>0&#x00B0; (frontend), 30&#x00B0;, and 90&#x00B0; (profile)</td>
</tr>
<tr>
<td/>
<td>Speaker-independent (acquisition/usage)</td>
<td>Yes</td>
</tr>
<tr>
<td>Technicality</td>
<td>Cameras</td>
<td>Three digital cameras</td>
</tr>
<tr>
<td/>
<td>Output data</td>
<td>Videos with MOV format</td>
</tr>
<tr>
<td/>
<td>Resolution</td>
<td>Full HD (1920 &#x00D7; 1080 Pixel)</td>
</tr>
<tr>
<td/>
<td>Frame rate</td>
<td>30 fps</td>
</tr>
<tr>
<td/>
<td>Method for controlling vibration</td>
<td>Using a camera tripod stand</td>
</tr>
<tr>
<td/>
<td>Recording environment</td>
<td>Controlled lab with good illumination and plain background</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Classification Model</title>
<p>The transfer learning technique was used to investigate various efficient lip-reading deep learning models and examine their classification capabilities on our video data from AQAND, as shown in Subsections 3.1 and 3.2, to enforce a reliable deep learning model capable of achieving a high recognition accuracy for visual speech recognition. In this work, based on transfer learning, we apply AQAND, a modified state-of-the-art method [<xref ref-type="bibr" rid="ref-12">12</xref>], represented in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, and evaluate the performance efficacy of the nascent model, then give a thorough analysis of the findings and insights for future research. The classification model is given an input video after pre-processing and applying data augmentation techniques where &#x201C;HF &#x002B; AT&#x201D; in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> means that horizontal flip and affine transformation techniques are applied. Then, we fed it into the residual network (ResNet-18), modifying 2D convolution to 3D. The output features from the global average pooling are then given to the backend network. The total number of letter or word classes is the output dimension of the final fully connected layer.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Flow chart of whole pre-processing. <italic>N</italic> refers to the fixed number of video frames. Section A is applied only one time, while the remaining steps include data pre-processing that is related to classification and feature extraction for the model and can be replaceable</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-4.tif"/>
</fig>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The overview of workflow a classification model architecture in our task</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments and Results</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup and Initialization</title>
<p>The experiments are carried out utilizing the Google Colab environment with a large amount of RAM (51 GB) and a single GPU. The Google Drive is connected and accessed by the Colab environment, where all data and necessary files are stored. The Pytorch library is used for coding the proposed classification model outlined in Section 3. The Cross-Entropy (CE) loss function for optimization and the Adam optimizer with an initial learning rate of 7e-4 is used and combined with cosine learning rate scheduling and weight decay of 1e-4 with a batch size of 32. The cosine learning rate scheduling training trick is used in the same manner as in [<xref ref-type="bibr" rid="ref-12">12</xref>]. Every time the validation error reaches a plateau throughout three consecutive epochs, the learning rate will be reduced by a factor of two when validating the model at the end of each epoch. This is to prevent any abrupt reduction of the learning rate. The learning rate <italic>&#x03B7;</italic> at an epoch <italic>&#x03C4;</italic> in the cosine setting is determined by <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> as follows:</p>
<p><disp-formula id="eqn-1">
<label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mi>&#x03C4;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mi>&#x03C0;</mml:mi></mml:mrow><mml:mi>T</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mi>&#x03B7;</mml:mi></mml:math></disp-formula>where <italic>&#x03B7;</italic> represents the initial learning rate, and <italic>T</italic> stands for the overall number of epochs, which in our experiments is 100.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Dataset Statistics</title>
<p>In this paper, all conducted lip-reading experiments are performed only on the zero-angle video samples (i.e., frontend) of AQAND, with all categories (classes) of the dataset (single letters, disjoined letters, and Quranic words), consisting of a total of 53 letters and words classes, obtained from 22 reciters repeated three times, resulting in 3498 total samples, which we split into 2544 training samples, 477 validation samples, and 477 testing samples. <xref ref-type="table" rid="table-7">Table 7</xref> shows the number of samples distributed in each category per split set. Each category of data is trained individually.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Number of samples per split set of zero-angle video data</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Dataset split</th>
<th rowspan="2">Splitting ratio</th>
<th rowspan="2">Number of reciters</th>
<th colspan="3" align="center">Main categories in the dataset</th>
</tr>
<tr>
<th>Single alphabets</th>
<th>Disjoined letters</th>
<th>Quranic words</th>
</tr>
</thead>
<tbody>
<tr>
<td>Training set</td>
<td>70%</td>
<td>16</td>
<td>1392</td>
<td>672</td>
<td>480</td>
</tr>
<tr>
<td>Validation set</td>
<td>15%</td>
<td>3</td>
<td>261</td>
<td>126</td>
<td>90</td>
</tr>
<tr>
<td>Testing set</td>
<td>15%</td>
<td>3</td>
<td>261</td>
<td>126</td>
<td>90</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experimental Analysis</title>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Arabic Speech Production</title>
<p>Each letter in every language has a specific sound. Each sound is produced by a speaker from a specific place of the mouth using a precise utterance mechanism. Twenty-nine letters make up the Arabic alphabet, as per <italic>QNbook</italic>. The Arabic language can be distinguished from other languages in that each letter has a specific phoneme, and some of these letters are unique and cannot be found in other languages, such as the letter &#x2018;Dhad,&#x2019; which in Arabic is written as &#x2018;<inline-graphic xlink:href="CSSE_37113-inline-8.tif"/>&#x2019;. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> explains all 29 Arabic letters, and from which place in the mouth they are each produced (also described as the phonic production/exit place of a spoken letter). It also shows that Arabic letters are produced from 13 different places with specific phonemes and lip movements for each. All letters are arranged in groups of different colors according to their phonic production place in the mouth, starting with the letters issued from the lips and ending with the letters issued from the back of the throat (glottal).</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>An illustration of Arabic letters production/exit places of the mouth during each letter utterance</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-6.tif"/>
</fig>
<p>The Arabic-based visual speech recognition task may face many challenges regarding the letters in groups 11, 12, and 13, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. These letters are distinguished in the Arabic language and cannot be found in most other languages. Recognizing these letters only visually is considered challenging as they are produced from the back of the throat with no lip movement. This difficulty in recognizing these letters is also considered a major problem for deaf persons because they cannot notice slight changes in the movement of the lips, so they may not be able to imitate and pronounce these letters. As such, in visual speech recognition, we must pay attention to the concept of &#x201C;visemes&#x201D;. Visemes are letters or words that are similar in the visual mouth movement and shape during their pronunciation but are different in sound and meaning. <xref ref-type="table" rid="table-8">Table 8</xref> summarizes all visemes found in the Arabic alphabet. This includes the group of letters &#x2018;Tha&#x2019; (<inline-graphic xlink:href="CSSE_37113-inline-9.tif"/>), &#x2018;Dha&#x2019; (<inline-graphic xlink:href="CSSE_37113-inline-10.tif"/>), and &#x2018;Dha&#x2019; (<inline-graphic xlink:href="CSSE_37113-inline-11.tif"/>), which are three completely different letters in sound, but their visual movements look very similar on the speaker&#x2019;s lips, as shown for viseme number 3 in <xref ref-type="table" rid="table-8">Table 8</xref>.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>The twelve visemes of Arabic isolated letters, illustrating the visual similarity between letters in the same group that have the same exits, resulting in one single inferred viseme for them</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="center"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CSSE_37113-inline-12.tif"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Focusing on the pronunciation style of the Arabic alphabet was the first step in the analysis, and we observed that it relied on two factors:
<list list-type="simple">
<list-item><label>&#x25A0;</label><p><bold>The ARTICULATIONs</bold> are the places of the letter originates (area of letters exit from the mouth), which are identified and summarized per similar shapes of the mouth and are called visemes. Some Arabic letters share the same <italic>articulations</italic> and, thus, mimic the same movements of the lips (i.e., the same viseme), which creates a challenge to recognize and classify them. Note that the entire alphabet is categorized into 12 visemes, as shown in <xref ref-type="table" rid="table-8">Table 8</xref>.</p>
</list-item>
<list-item><label>&#x25A0;</label><p><bold>The HOWs</bold>, which are the styles or the manners that accompany each letter when pronouncing it, such as &#x2018;whispering&#x2019; (&#x2018;Hams&#x2019;), &#x2018;loudness&#x2019; (&#x2018;Jahr&#x2019;), &#x2018;intensity&#x2019; (&#x2018;Shedah&#x2019;), and &#x2018;looseness&#x2019; (&#x2018;Rakhawah&#x2019;), etc. They vary from strong to weak, and the letter&#x2019;s strength or weakness is determined by the number of <italic>hows</italic> it carries and whether strong or weak.</p></list-item>
</list></p>
<p>Since our work is isolated from audio speech recognition and dependent solely on visual speech recognition, we divided the single letters based on their exits, strength, and weakness in pronunciation into three groups, so the letters with similar visemes are separated into two different groups. The <italic>hows</italic> have a key role in distinguishing between similar visemes; this is their primary function. If we had dealt with audio speech recognition along with visual speech recognition, the audio information of <italic>hows</italic> would have had an additional role in discriminating between the letters with visually similar visemes.</p>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experimental Results and Discussions</title>
<p>The confusion matrix of the classification model using the testing set for the Quranic words category is shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. The number of samples in the testing set for each class is represented by the sum of the numbers in each row, which is constant (i.e., nine testing samples from 3 reciters). Since the mouth movements of the letters in words W3 (&#x2018;Hasad,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-13.tif"/>), W6 (&#x2018;Seraja,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-14.tif"/>), and W9 (&#x2018;Ahad,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-15.tif"/>) are so similar, it is apparent that W3 is the most confusable and difficult to predict accurately. In the sequence of frames for the words W3, W6, and W9, respectively, as shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the visemes are clear to see. Another significant observation is that the words W2 (&#x2018;Lawoh,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-16.tif"/>), W8 (&#x2018;Kholeeqa,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-17.tif"/>), and W10 (&#x2018;Kofua,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-18.tif"/>) are very similar in their use of an &#x2018;O&#x2019; mouth shape, which creates a challenge in differentiating and predicting them accurately. The loss and overall classification accuracy of the Quranic words in the testing set resulted in 83.33% accuracy, as shown in <xref ref-type="table" rid="table-9">Table 9</xref>, beside each class&#x2019;s precision and recall statistics.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>The confusion matrix for Quranic words classification</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Pronunciation sequence frames for three words: (a) W3 (&#x2018;Hasad,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-19.tif"/>), (b) W6 (&#x2018;Seraja,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-20.tif"/>), and (c) W9 (&#x2018;Ahad,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-21.tif"/>). Each letter (approximate corresponding voice) is written under each frame</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-8.tif"/>
</fig><table-wrap id="table-9"><label>Table 9</label>
<caption>
<title>The precision, recall, loss, and testing accuracy for Quranic words</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Class</th>
<th>Adhaab<break/>(W1)</th>
<th>Whyl<break/>(W2)</th>
<th>Hasad<break/>(W3)</th>
<th>Enaba<break/>(W4)</th>
<th>Lawoh<break/>(W5)</th>
<th>Seraja<break/>(W6)</th>
<th>Maleke<break/>(W7)</th>
<th>Kholeeqa<break/>(W8)</th>
<th>Ahad<break/>(W9)</th>
<th>Kofua<break/>(W10)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Precision</td>
<td>100</td>
<td>81.8</td>
<td>71.4</td>
<td>80</td>
<td>100</td>
<td>75</td>
<td>80</td>
<td>88.9</td>
<td>64.3</td>
<td>100</td>
</tr>
<tr>
<td>Recall</td>
<td>66.7</td>
<td>100</td>
<td>55.6</td>
<td>88.9</td>
<td>100</td>
<td>66.7</td>
<td>88.9</td>
<td>88.9</td>
<td>100</td>
<td>66.7</td>
</tr>
<tr>
<td colspan="11" align="center"><bold>Overall accuracy 83.33%</bold></td>
</tr>
<tr>
<td colspan="11" align="center"><bold>Loss 1.07</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The confusion matrix of the disjoined letters category for the testing set is shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. We observed that the letters DL8 (&#x2018;TaSieen,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-22.tif"/>) and DL9 (&#x2018;YaSieen,&#x2019; <inline-graphic xlink:href="CSSE_37113-inline-23.tif"/>) caused the most mutual confusion in classification due to visemes used in the second part of the word. The loss and overall classification accuracy of the disjoined letters in the testing set to result in an accuracy of 80.47% with corresponding precision and recall results for each class are shown in <xref ref-type="table" rid="table-10">Table 10</xref>.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>The confusion matrix for disjoined letters classification</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-9.tif"/>
</fig><table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>The precision, recall, loss, and overall accuracy of disjoined letters</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Class</th>
<th>DL1</th>
<th>DL2</th>
<th>DL3</th>
<th>DL4</th>
<th>DL5</th>
<th>DL6</th>
<th>DL7</th>
<th>DL8</th>
<th>DL9</th>
<th>DL10</th>
<th>DL11</th>
<th>DL12</th>
<th>DL13</th>
<th>DL14</th>
</tr>
</thead>
<tbody>
<tr>
<td>Precision</td>
<td>100</td>
<td>100</td>
<td>87.5</td>
<td>81.8</td>
<td>100</td>
<td>100</td>
<td>100</td>
<td>83.3</td>
<td>50</td>
<td>85.7</td>
<td>64.3</td>
<td>64.3</td>
<td>100</td>
<td>70</td>
</tr>
<tr>
<td>Recall</td>
<td>66.7</td>
<td>66.7</td>
<td>77.8</td>
<td>100</td>
<td>100</td>
<td>100</td>
<td>66.7</td>
<td>55.6</td>
<td>77.8</td>
<td>66.7</td>
<td>100</td>
<td>100</td>
<td>66.7</td>
<td>77.8</td>
</tr>
<tr>
<td colspan="15" align="center"><bold>Overall accuracy 80.47%</bold></td>
</tr>
<tr>
<td colspan="15" align="center"><bold>Loss 1.5</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Given the single letters splitting method outlined in Sub-section 4.3, the confusion matrices for testing the three groups (G1, G2, and G3) of single letters are shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>. It appears that single Arabic letters are the most difficult classification task compared with the other two categories. This is due to several letters possessing the same exits, which results in sharing similar visemes. Furthermore, such spoken letters are short, and some of them have no apparent movement of the lips because the spoken letters are produced from the back of the throat (Horof Halaqia). We obtained accuracies of 72.66%, 70.31%, and 77.48% for G1, G2, and G3, respectively, despite the challenging factors affecting the performance of the classification model in this task. The loss and overall classification accuracy for each group of the single letters test set and precision and recall statistics are shown in <xref ref-type="table" rid="table-11">Tables 11</xref>, <xref ref-type="table" rid="table-12">12</xref>, and <xref ref-type="table" rid="table-13">13</xref> for G1, G2, and G3, respectively.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Confusion matrices for single alphabets classification. (a) G1, (b) G2, (c) G3</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-10a.tif"/><graphic mimetype="image" mime-subtype="tif" xlink:href="CSSE_37113-fig-10b.tif"/>
</fig><table-wrap id="table-11"><label>Table 11</label>
<caption>
<title>Precision, recall, overall accuracy, and loss for G1 of single alphabets, shown in <xref ref-type="fig" rid="fig-10">Fig. 10a</xref></title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Group G1</th>
<th>Alif<break/>(SL1)</th>
<th>Ba<break/>(SL2)</th>
<th>Ta<break/>(SL3)</th>
<th>Tha<break/>(SL4)</th>
<th>Jeem<break/>(SL5)</th>
<th>Ha<break/>(SL6)</th>
<th>Dal<break/>(SL7)</th>
<th>Ra<break/>(SL8)</th>
<th>Sien<break/>(SL9)</th>
<th>Ya<break/>(SL10)</th>
<th>Wow<break/>(SL11)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Precision</td>
<td>80</td>
<td>77.8</td>
<td>50</td>
<td>62.5</td>
<td>100</td>
<td>57.1</td>
<td>25</td>
<td>55.6</td>
<td>50</td>
<td>83.3</td>
<td>100</td>
</tr>
<tr>
<td>Recall</td>
<td>88.9</td>
<td>77.8</td>
<td>55.6</td>
<td>55.6</td>
<td>100</td>
<td>44.4</td>
<td>44.4</td>
<td>55.6</td>
<td>33.3</td>
<td>55.6</td>
<td>100</td>
</tr>
<tr>
<td colspan="12" align="center"><bold>Overall accuracy 72.66%</bold></td>
</tr>
<tr>
<td colspan="12" align="center"><bold>Loss 1.6</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-12"><label>Table 12</label>
<caption>
<title>Precision, recall, overall accuracy, and loss for G2 of single alphabets, shown in <xref ref-type="fig" rid="fig-10">Fig. 10b</xref></title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Group G2</th>
<th>Kha<break/>(SL1)</th>
<th>Dhal<break/>(SL2)</th>
<th>Za<break/>(SL3)</th>
<th>Sheen<break/>(SL4)</th>
<th>Dhad<break/>(SL5)</th>
<th>Ta<break/>(SL6)</th>
<th>Ghien<break/>(SL7)</th>
<th>Fa<break/>(SL8)</th>
<th>Kaf<break/>(SL9)</th>
<th>Ba<break/>(SL10)</th>
<th>Wow<break/>(SL11)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Precision</td>
<td>71.4</td>
<td>100</td>
<td>42.9</td>
<td>100</td>
<td>100</td>
<td>100</td>
<td>36</td>
<td>53.8</td>
<td>100</td>
<td>64.3</td>
<td>90</td>
</tr>
<tr>
<td>Recall</td>
<td>55.6</td>
<td>11.1</td>
<td>66.7</td>
<td>55.6</td>
<td>44.4</td>
<td>22.2</td>
<td>100</td>
<td>77.8</td>
<td>44.4</td>
<td>100</td>
<td>100</td>
</tr>
<tr>
<td colspan="12" align="center"><bold>Overall accuracy 70.31%</bold></td>
</tr>
<tr>
<td colspan="12" align="center"><bold>Loss 1.7</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Precision, recall, overall accuracy, and loss for G3 of single alphabets, shown in <xref ref-type="fig" rid="fig-10">Fig. 10c</xref></title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Group G3</th>
<th>Sad<break/>(SL1)</th>
<th>Dhad<break/>(SL2)</th>
<th>Aieen<break/>(SL3)</th>
<th>Qaf<break/>(SL4)</th>
<th>Noon<break/>(SL5)</th>
<th>Hamzah<break/>(SL6)</th>
<th>Meem<break/>(SL7)</th>
<th>Wow<break/>(SL8)</th>
<th>Ha<break/>(SL9)</th>
<th>Lam<break/>(SL10)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Precision</td>
<td>77.8</td>
<td>87.5</td>
<td>50</td>
<td>69.2</td>
<td>100</td>
<td>100</td>
<td>85.7</td>
<td>75</td>
<td>66.7</td>
<td>100</td>
</tr>
<tr>
<td>Recall</td>
<td>77.8</td>
<td>77.8</td>
<td>66.7</td>
<td>100</td>
<td>66.7</td>
<td>100</td>
<td>66.7</td>
<td>100</td>
<td>66.7</td>
<td>55.6</td>
</tr>
<tr>
<td colspan="11" align="center"><bold>Overall accuracy 77.48%</bold></td>
</tr>
<tr>
<td colspan="11" align="center"><bold>Loss 1.3</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-14">Table 14</xref> summarizes of overall accuracy results according to all experiments and observations in the above three paragraphs. Though the classification model&#x2019;s prediction performance for Quranic words and disjoined letters was superior, ranging from 80.47% to 83.33%, it was somewhat lower for single Arabic letters, with a 73.5% average accuracy for all groups. Therefore, it may be important to carry out further investigations to enhance the performance of their prediction, which can be a potential avenue for future research. Our findings are comparable to or even better than those of the mandarin language, as shown in <xref ref-type="table" rid="table-15">Table 15</xref>, which may emphasize the effect of distinctions between the benchmarked languages in the lip-reading challenge. The transfer learning model we employed in our experiments was pre-trained utilizing the LRW dataset.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>All the experimental results of the classification model on the AQAND</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th>Quranic words</th>
<th>Disjoined letters</th>
<th colspan="3" align="center">Single letters</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">Accuracy</td>
<td rowspan="2">83.33%</td>
<td rowspan="2">80.47%</td>
<td>G1</td>
<td>G2</td>
<td>G3</td>
</tr>
<tr>
<td>72.66%</td>
<td>70.31%</td>
<td>77.48%</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Comparison with the existing work</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Dataset</th>
<th rowspan="2">The experiment content type</th>
<th rowspan="2">Language</th>
<th colspan="2" align="center">Model</th>
<th rowspan="2">Accuracy</th>
</tr>
<tr>
<th>Frontend</th>
<th>Backend</th>
</tr>
</thead>
<tbody>
<tr>
<td>The LRW</td>
<td rowspan="3">Word-level</td>
<td>English</td>
<td rowspan="3">ResNet-18</td>
<td rowspan="3">3 Layers GRU</td>
<td>83.7%</td>
</tr>
<tr>
<td>LRW1000</td>
<td>Mandarin</td>
<td>46.5%</td>
</tr>
<tr>
<td>AQAND (ours)</td>
<td>Arabic</td>
<td>83.3%</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions and Future Work</title>
<p>In this paper, to the best of our knowledge, we provide the first dataset of the Arabic language based on the book <italic>Al-Qaida Al-Noorania</italic> (AQAND) comprising single letters, disjoined letters, and Quranic words as video data samples captured from three different viewpoints (0&#x00B0;, 30&#x00B0;, and 90&#x00B0;). We use the new, proposed AQAND to train a CNN-based deep learning model using transfer learning for visual speech recognition (known as lip-reading). For disjoined letters, our model achieved an accuracy of 80.5% while achieving an accuracy of 77.5% for single letters. However, the experimental results showed that the performance of lip-reading for Quranic words in recognizing completely unseen test data samples achieved the highest accuracy of 83.33%.</p>
<p>Three future directions that we see are as follows: first, an additional CNN deep learning model can be trained as an initial step before the classification model to split and sort the single alphabets of each viseme into subgroups instead of the manually implemented splitting, to be then fed into the classification model for prediction. Second, the experimental work on the dataset of AQAND can be expanded by examining the proposed deep learning model on the remaining two replications of video data captured in 30&#x00B0; and 90&#x00B0; viewpoints and comparing their performance results from different aspects and in multiple scenarios. Third, the serving and contributing to the field of continuous speech recognition at a sentence level, as we intend to add complete Quranic verses to the dataset in the future. As such, establishing the basic artificial intelligence capability of Arabic Quranic lip-reading can be considered as a preface for further developments in recognition of continuous recitation of Holy Quran verses, which contributes to effectively employing the field of artificial intelligence in teaching the correct recitation of the Holy Quran.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to thank King Abdulaziz University Scientific Endowment for funding the research reported in this paper. They would also like to thank the Islamic University of Madinah for its extensive help in the data collection phase.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This research was supported and funded by <funding-source>KAU Scientific Endowment, King Abdulaziz University</funding-source>, Jeddah, Saudi Arabia.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The AQAND dataset proposed in this work is available at (<ext-link ext-link-type="uri" xlink:href="https://forms.gle/x5tQcDLeZUwJyG789">https://forms.gle/x5tQcDLeZUwJyG789</ext-link>).</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Alhaqani</surname></string-name></person-group>, &#x201C;<article-title>Al-qaida Al-noorania</article-title>,&#x201D; <source>Al-Furqan Center for Quran Learning</source>, vol. <volume>1</volume>, no. <issue>1</issue>, pp. <fpage>36</fpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sagheer</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Naoyuki</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Taniguchi</surname></string-name></person-group>, &#x201C;<article-title>Arabic lip-reading system: A combination of hypercolumn neural network model with hidden Markov model</article-title>,&#x201D; <conf-name>Proceedings of International Conference  
on Artificial Intelligence and Soft Computing</conf-name>, vol. <volume>2004</volume>, pp. <fpage>311</fpage>&#x2013;<lpage>316</lpage>, <year>2004</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Pascal</surname></string-name></person-group>, &#x201C;<article-title>Visual speech recognition of modern classic Arabic language</article-title>,&#x201D; in <conf-name>2011 Int. Symp. on Humanities, Science and Engineering Research</conf-name>, <publisher-loc>Kuala Lumpur, Malaysia</publisher-loc>, vol. <volume>1</volume>, pp. <fpage>50</fpage>&#x2013;<lpage>55</lpage>, <publisher-name>IEEE</publisher-name>,  <year>2011</year>. </mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. S.</given-names> <surname>Chung</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zisserman</surname></string-name></person-group>, &#x201C;<article-title>Learning to lip read words by watching videos</article-title>,&#x201D; <source>Computer Vision and Image Understanding</source>, vol. <volume>173</volume>, no. <issue>5</issue>, pp. <fpage>76</fpage>&#x2013;<lpage>85</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Tao</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Busso</surname></string-name></person-group>, &#x201C;<article-title>End-to-end audio-visual speech recognition system with multitask learning</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>23</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>11</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Xue</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Hong</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Cao</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>LCSNet: End-to-end lipreading with channel-aware feature selection</article-title>,&#x201D; <source>ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)</source>, vol. <volume>18</volume>, no. <issue>4</issue>, pp. <fpage>27</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Stafylakis</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Tzimiropoulos</surname></string-name></person-group>, &#x201C;<article-title>Pushing the boundaries of audio-visual word recognition using residual networks and LSTMs</article-title>,&#x201D; <source>Computer Vision and Image Understanding</source>, vol. <volume>176</volume>, pp. <fpage>22</fpage>&#x2013;<lpage>32</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Pu</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>A lip reading method based on 3D convolutional vision transformer</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>10</volume>, pp. <fpage>77205</fpage>&#x2013;<lpage>77212</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Tian</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Liu</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Decoding lip language using triboelectric sensors with deep learning</article-title>,&#x201D; <source>Nature communications</source>, vol. <volume>13</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Jeon</surname></string-name> and <string-name><given-names>M. S.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>End-to-end sentence-level multi-view lipreading architecture with spatial attention module integrated multiple CNNs and cascaded local self-attention-CTC</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>22</volume>, no. <issue>9</issue>, pp. <fpage>3597</fpage>, <year>2022</year>; <pub-id pub-id-type="pmid">35591284</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Tsourounis</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kastaniotis</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Fotopoulos</surname></string-name></person-group>, &#x201C;<article-title>Lip reading by alternating between spatiotemporal and spatial convolutions</article-title>,&#x201D; <source>Journal of Imaging</source>, vol. <volume>7</volume>, no. <issue>5</issue>, pp. <fpage>91</fpage>, <year>2021</year>; <pub-id pub-id-type="pmid">34460687</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Shan</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Learn an effective lip reading model without pains</article-title>,&#x201D; <comment>arXiv preprint arXiv: 2011.07557</comment>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. S.</given-names> <surname>Chung</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zisserman</surname></string-name></person-group>, &#x201C;<article-title>Lip reading in the wild</article-title>,&#x201D; <source>Asian Conference on Computer Vision</source>, vol. <volume>13</volume>, no. <issue>2</issue>, pp. <fpage>87</fpage>&#x2013;<lpage>103</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>LRW-1000: A naturally-distributed large-scale benchmark for lip reading in the wild</article-title>,&#x201D; in <conf-name>2019 14th IEEE Int. Conf. on Automatic Face &#x0026; Gesture Recognition (FG, 2019)</conf-name>, <publisher-loc>Lille,
France</publisher-loc>, vol. <volume>1</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Egorov</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Kostyumov</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Konyk</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Kolesnikov</surname></string-name></person-group>, &#x201C;<article-title>LRWR: Large-scale benchmark for lip reading in russian language</article-title>,&#x201D; <comment>arXiv preprint arXiv: 2109.06692</comment>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Jeon</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Elsharkawy</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Sang Kim</surname></string-name></person-group>, &#x201C;<article-title>Lipreading architecture based on multiple convolutional neural networks for sentence-level visual speech recognition</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>22</volume>, no. <issue>1</issue>, pp. <fpage>72</fpage>, <year>2021</year>; <pub-id pub-id-type="pmid">35009612</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>&#x00DC;.</given-names> <surname>Atila</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Sabaz</surname></string-name></person-group>, &#x201C;<article-title>Turkish lip-reading using Bi-LSTM and deep learning models</article-title>,&#x201D; <source>Engineering Science and Technology, An International Journal</source>, vol. <volume>1</volume>, no. <issue>35</issue>, pp. <fpage>101206</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Xiao</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>A chinese lip-reading system based on convolutional block attention module</article-title>,&#x201D; <source>Mathematical Problems in Engineering</source>, vol. <volume>2021</volume>, no. <issue>12</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Ziafat</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Hafiz Farooq</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Iram</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Muhammad</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alhumam</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Kashif</surname></string-name></person-group>, &#x201C;<article-title>Correct pronunciation detection of the Arabic alphabet using deep learning</article-title>,&#x201D; <source>Applied Sciences</source>, vol. <volume>11</volume>, no. <issue>6</issue>, pp. <fpage>2508</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Asif</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Mukhtar</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Alqadheeb</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Hafiz Farooq</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Alhumam</surname></string-name>, </person-group>&#x201C;<article-title>An approach for pronunciation classification of classical Arabic phonemes using deep learning</article-title>,&#x201D; <source>Applied Sciences</source>, vol. <volume>12</volume>, no. <issue>1</issue>, pp. <fpage>238</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Damien</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Wakim</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Egea</surname></string-name></person-group>, &#x201C;<article-title>Phoneme-viseme mapping for modern, classical Arabic language</article-title>,&#x201D; in <conf-name>2009 Int. Conf. on Advances in Computational Tools for Engineering Applications</conf-name>, <conf-loc>Zouk Mosbeh, Lebanon</conf-loc>, vol. <volume>1</volume>, pp. <fpage>547</fpage>&#x2013;<lpage>552</lpage>, <year>2009</year>. </mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname> F. Z. Chelali</surname></string-name>, <string-name> <surname>K. Sadeddine</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Djeradi</surname></string-name></person-group>, &#x201C;<article-title>Visual speech analysis application to Arabic phonemes</article-title>,&#x201D; <source>Special Issue of International Journal of Computer Applications (0975-8887) on Software Engineering, Databases and Expert Systems-SEDEXS</source>, vol. <volume>102</volume>, no. <issue>1</issue>, pp. <fpage>29</fpage>&#x2013;<lpage>34</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Altalmas</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ammar</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ahmad</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Sediono</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Momoh</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Lips tracking identification of a correct Quranic letters pronunciation for Tajweed teaching and learning</article-title>,&#x201D; <source>IIUM Engineering Journal</source>, vol. <volume>18</volume>, no. <issue>1</issue>, pp. <fpage>177</fpage>&#x2013;<lpage>191</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Elrefaei</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Alhassan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Omar</surname></string-name></person-group>, &#x201C;<article-title>An Arabic visual dataset for visual speech recognition</article-title>,&#x201D; <source>Procedia Computer Science</source>, vol. <volume>163</volume>, no. <issue>10</issue>, pp. <fpage>400</fpage>&#x2013;<lpage>409</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Dweik</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Altorman</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Ashour</surname></string-name></person-group>, &#x201C;<article-title>Read my lips: Artificial intelligence word-level Arabic lip-reading system</article-title>,&#x201D; <source>Egyptian Informatics Journal</source>, vol. <volume>23</volume>, no. <issue>2</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Alsulami</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jamal</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Elrefaei</surname></string-name></person-group>, &#x201C;<article-title>Deep learning-based approach for Arabic visual speech recognition</article-title>,&#x201D; <source>CMC-Computers, Materials &#x0026; Continua</source>, vol. <volume>71</volume>, no. <issue>1</issue>, pp. <fpage>85</fpage>&#x2013;<lpage>108</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. Y.</given-names> <surname>El Amrani</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Rahman</surname></string-name>, <string-name><given-names>M. R.</given-names> <surname>Wahiddin</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Shah</surname></string-name></person-group>, &#x201C;<article-title>Building CMU Sphinx language model for the Holy Quran using simplified Arabic phonemes</article-title>,&#x201D; <source>Egyptian informatics journal</source>, vol. <volume>17</volume>, no. <issue>3</issue>, pp. <fpage>305</fpage>&#x2013;<lpage>314</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Abed</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Alshayeji</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Sultan</surname></string-name></person-group>, &#x201C;<article-title>Diacritics effect on Arabic speech recognition</article-title>,&#x201D; <source>Arabian Journal for Science and Engineering</source>, vol. <volume>44</volume>, no. <issue>11</issue>, pp. <fpage>9043</fpage>&#x2013;<lpage>9056</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Al-Kaf</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Sulong</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Joret</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Aminuddin</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Mohammad</surname></string-name></person-group>, &#x201C;<article-title>QVR: Quranic verses recitation recognition system using pocketsphinx</article-title>,&#x201D; <source>Journal of Quranic Sciences and Research</source>, vol. <volume>2</volume>, no. <issue>2</issue>, pp. <fpage>35</fpage>&#x2013;<lpage>41</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Rafi</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>A. W.</given-names> <surname>Usmani</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Zulqarnain</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Shuja</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Quran companion-A helping tool for huffaz</article-title>,&#x201D; <source>Journal of Information &#x0026; Communication Technology</source>, vol. <volume>13</volume>, no. <issue>2</issue>, pp. <fpage>21</fpage>&#x2013;<lpage>27</lpage>, <year>2019</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>