<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">IASC</journal-id>
<journal-id journal-id-type="nlm-ta">IASC</journal-id>
<journal-id journal-id-type="publisher-id">IASC</journal-id>
<journal-title-group>
<journal-title>Intelligent Automation &#x0026; Soft Computing</journal-title>
</journal-title-group>
<issn pub-type="epub">2326-005X</issn>
<issn pub-type="ppub">1079-8587</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">12775</article-id>
<article-id pub-id-type="doi">10.32604/iasc.2020.012775</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Text Detection and Classification from Low Quality Natural Images</article-title><alt-title alt-title-type="left-running-head">Text Detection and Classification from Low Quality Natural Images</alt-title><alt-title alt-title-type="right-running-head">Text Detection and Classification from Low Quality Natural Images</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western">
<surname>Yasmeen</surname>
<given-names>Ujala</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western">
<surname>Shah</surname>
<given-names>Jamal Hussain</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western">
<surname>Khan</surname>
<given-names>Muhammad Attique</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western">
<surname>Ansari</surname>
<given-names>Ghulam Jillani</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western">
<surname>ur Rehman</surname>
<given-names>Saeed</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western">
<surname>Sharif</surname>
<given-names>Muhammad</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-7" contrib-type="author">
<name name-style="western">
<surname>Kadry</surname>
<given-names>Seifedine</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib id="author-8" contrib-type="author" corresp="yes">
<name name-style="western">
<surname>Nam</surname>
<given-names>Yunyoung</given-names>
</name>
<xref ref-type="aff" rid="aff-4">4</xref>
<email>ynam@sch.ac.kr</email>
</contrib>
<aff id="aff-1">
<label>1</label><institution>Department of Computer Science, Wah Campus, COMSATS University Islamabad</institution>, <addr-line>Islamabad, 47040</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-2">
<label>2</label><institution>Department of Computer Science, HITEC University</institution>, <addr-line>Taxila, 47080</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-3">
<label>3</label><institution>Department of Mathematics and Computer Science, Beirut Arab University</institution>, <addr-line>Beirut</addr-line>, <country>Lebanon</country></aff>
<aff id="aff-4">
<label>4</label><institution>Department of Computer Science and Engineering, Soonchunhyang University</institution>, <addr-line>Asan, 31538</addr-line>, <country>South Korea</country></aff>
</contrib-group><author-notes><corresp id="cor1">&#x002A;Corresponding Author: Yunyoung Nam. Email: <email>ynam@sch.ac.kr</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2020-12-19">
<day>19</day>
<month>12</month>
<year>2020</year>
</pub-date>
<volume>26</volume>
<issue>6</issue>
<fpage>1251</fpage>
<lpage>1266</lpage>
<history>
<date date-type="received">
<day>02</day>
<month>7</month>
<year>2020</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>9</month>
<year>2020</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2020 Yasmeen et al.</copyright-statement>
<copyright-year>2020</copyright-year>
<copyright-holder>Yasmeen et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_IASC_12775.pdf"></self-uri>
<abstract>
<p>Detection of textual data from scene text images is a very thought-provoking issue in the field of computer graphics and visualization. This challenge is even more complicated when edge intelligent devices are involved in the process. The low-quality image having challenges such as blur, low resolution, and contrast make it more difficult for text detection and classification. Therefore, such exigent aspect is considered in the study. The technology proposed is comprised of three main contributions. (a) After synthetic blurring, the blurred image is preprocessed, and then the deblurring process is applied to recover the image. (b) Subsequently, the standard maximal stable extreme regions (MSER) technique is applied to localize and detect text. Soon after, K-Means is applied to get three different clusters of the query image to separate foreground and background and also incorporate character level grouping. (c) Finally, the segmented text is classified into textual and non-textual regions using a novel convolutional neural network (CNN) framework. The purpose of this task is to overcome the false positives. For evaluation of proposed technique, results are obtained on three mainstream datasets, including SVT, IIIT5K and ICDAR 2003. The achieved classification results of 90.3% for SVT dataset, 95.8% for IIIT5K dataset, and 94.0% for the ICDAR 2003 dataset, respectively. It shows the preeminence of the proposed methodology that it works fine for good model learning. Finally, the proposed methodology is compared with previous benchmark text-detection techniques to validate its contribution.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Feature points</kwd>
<kwd>K-means</kwd>
<kwd>deep learning</kwd>
<kwd>blur image</kwd>
<kwd>color spaces</kwd>
<kwd>classification</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The recent era has witnessed a lot of growth in smart cities. Especially in health care, banking, education and video surveillance as well as IoT based autonomous vehicles. Such systems demand edge and app intelligence mechanisms. Further, the invasion of deep learning (DL) based systems have also evolved and put many challenges on these autonomous systems [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. Textual data in images presents instrumental knowledge for content-based image repossession and many other applications of computer vision. This textual information varies because of disparities in font size, style, alignment, random orientation. Above all, the low contrast, low resolution, blur and complex background make its detection (localization and identification) and classification (verification) more challenging [<xref ref-type="bibr" rid="ref-3">3</xref>]. Generally, the scenic properties like foreground and background (Textual and non-textual data) often creates problems in text detection and classification. Foreground properties include variation in size, colour, font and orientation, which become a source of complication in the detection of textual data robustly from scene text images. Conversely, images with complex backgrounds contain a variety of objects with multiple colours along with sky, grass, bricks, and fences that decline the robustness and creates difficulty in extraction of textual features.</p>
<p>To counter low-quality property of an image for detecting text requires high-level skills and machine learning approaches to enhance the image for attaining good results. However, in recent years techniques based on deep learning have been introduced which acquire classified features from training data, provides new promising results on benchmarks as ICDAR series contents [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. Handcrafted features were also introduced to get the properties of textual areas as shape and texture of text. But most of them deal with standard quality images rather than low-quality images. As mentioned earlier, important information is stored in textual images. It is utilized in videos and image applications based on content, as in searching web images base on content, repossession of information from videos, text recognition and analysis based on mobile. A lot of projects recently have been using maximally stable extremal regions (MSER) based methods as character candidates. Although MSER is considered as the best method for localizing text in ICDAR 2011 Robust Reading Competition [<xref ref-type="bibr" rid="ref-6">6</xref>] and produces promising results. Nevertheless, at the same time, report some problems. The ability to read robust textual data from dissipated scene text images helped in multiple real-world applications. For example, advantageous technologies for visually harmed persons and also in geo-localization, urban and robot navigation, cross-lingual access and in autonomous vehicles [<xref ref-type="bibr" rid="ref-7">7</xref>] containing onboard units (OBU) and interacting with road side unit (RSU) as well as other edge and IoT based units [<xref ref-type="bibr" rid="ref-8">8</xref>]. Therefore, most recently, the problem of detection of textual data from natural scene text images of low quality has acquired increasing attention from computer vision. Moreover, such challenges are even more important when the systems also demand privacy, latency and scalability while trying to utilize the edge intelligent devices and nodes [<xref ref-type="bibr" rid="ref-9">9</xref>]. Scene text detection and classification became challenging because of two types of factors known as internal and external. External factors are based on the environment, which causes blur, noise, occlusions etc. to create problems in the detection of textual data. Internal factors are the properties and dissimilarities in textual data from scene text images [<xref ref-type="bibr" rid="ref-10">10</xref>]. The complications in scene text are due to the three major reasons which are: 1) Natural images containing text in orientations, so bounding boxes are oriented rectangles or quadrangles; 2) Significant variations in the aspect ratio of scene text bounding boxes; 3) Variations of text data as characters, words, and text lines may cause confusion for algorithms to locate the boundaries.</p>
<p>Alternatively, most of the traditional methods, such as MSE, are found credible in detecting text. In recent times, methods based on CNN achieve up to the mark results for classification of text and many other domains like surveillance [<xref ref-type="bibr" rid="ref-11">11</xref>], medicine [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>], biometric [<xref ref-type="bibr" rid="ref-14">14</xref>], and agriculture [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. Powerful extraction of characteristic features is the immense power of CNN based models, and these are further helpful for high-class model learning, which can respond effectively to unseen data [<xref ref-type="bibr" rid="ref-17">17</xref>]. Considering the effectiveness of such methods, a novel methodology is suggested in this paper for the detection and classification of textual data from low quality natural scene images. In this paper, K-Means and MSER are used for separating foreground from background and a CNN model for classification of regions containing text and non-text.</p>
<sec id="s1_1">
<label>1.1</label>
<title>Objective and Contribution</title>
<p>To address the app and edge level challenges in this paper, we have proposed a DL based intelligence system for a smart vehicle that is equipped with the smart cameras and is capable of processing text that is blurred and is challenging to process. Main points of the proposed work are as follows:<list list-type="bullet"><list-item>
<p>Firstly, input images are converted into blur images, as we do not have any benchmarked dataset of blur text images. So averaging filter is applied to the input images to make them blur. Then, preprocessing is applied on blur images for deblurring and improving the quality.</p></list-item><list-item>
<p>Further, histogram equalization is applied on blurry images for contrast enhancement, then L&#x002A;a&#x002A;b color space is applied for preprocessing as tones and colors are held distinctly, and one color can be adjusted without disturbing the other.</p></list-item><list-item>
<p>After visual enhancement of an image, the challenge of text detection having different foregrounds and backgrounds is encountered. The unsupervised learning algorithm k-means is exploited, where clusters are created using this algorithm. These clusters help in separation of foreground and background (text and non-text) from the query image.</p></list-item><list-item>
<p>Next, MSER is applied to detect textual data from each clustered image followed by character level grouping. Sometimes, the non-text connected components are discovered, which can be further discarded based on geometric properties.</p></list-item><list-item>
<p>Finally, obtained connected components are organized into text and non-text sections using the proposed CNN framework. By incorporating deep features, the false positive rate is significantly minimized.</p></list-item></list></p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Low-quality images are mostly affected by occlusions and blur effects. Occlusions in natural scene images are caused by overlapping of text, while blur effects are caused by capturing device issues and also because of uncontrolled light effects. Therefore, it becomes challenging for optical character recognition (OCR) to recognize or understand the textual data from low-quality images. Some previous works deal with the rebuilding of occluded characters by watermarks in textual images [<xref ref-type="bibr" rid="ref-18">18</xref>]. Others also exploited the texture synthesis to reconstruct the text that can in-expensively and efficiently create novel texture by selecting and duplicating the references from sources [<xref ref-type="bibr" rid="ref-19">19</xref>]. Exemplar-based methodologies rebuild the texture, but problems are found using linear structures. Some more algorithms are proposed to solve the image filling issue. Using digital methods of inpainting fill holes by promulgating linear structures through diffusion. The downside is that process of diffusion causes blur, that becomes obvious when filling enormous region. Criminist et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed a technique by combining both texture and image inpainting methods. The perceptiveness is that &#x2018;inpainting front&#x2019; should spread along linear isophotes. Zhao et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] introduced a methodology for removing noise from low-quality images. Their algorithm contained three steps that are a) Match filter was applied initially) Wiener filter was used to remove noise further and c) they have used an average filter to smooth the images. This algorithm has improved the quality of images by removing noise from images. Earlier Ittner et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] have proposed an algorithm to categorize textual data from low-quality images. This model was presented to read textual data from documents. The proposed method used OCR to categorize text as OCR can handle a large number of words utilizing low-quality images. The classifier consists of two parts: i) The prototype to compare the documents and ii) A function to transform similarity of a document to prototype for estimating the probability of the document. Iqbal et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] introduced a colour correction model for improving low-quality images by using contrast correction by efficiently removing bluish color and increase the low red illumination for achieving high-quality images. Their proposed technique has three main steps. Firstly, the image was equalized on red, green blue (RGB) colors. Then contrast correction was applied on the RGB color model, and finally, on HSI color model is used. Another framework [<xref ref-type="bibr" rid="ref-24">24</xref>] has been proposed to disable the degradation of images acquired by an outdoor camera. Neverova et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed for improving images with lightning conditions. They have shown that a combination of color images and corresponding depth maps allows recovering estimations of locations and colors of multiple lights from scene images. The optimization process has been used to find lightening conditions that allow minimizing the difference of original image and rendered one. Tuyet et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] presented a method for detecting edges of objects from medical images. This technique has used Bayesian thresholding for de-noising of images and, B-spine curve for smoothing. Zhu et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] introduced a method to detect animals from low quality images. This technique presented two-channeled pyramid network for image detection. For gaining local information, a depth indication extracted from original image and two channeled perceiving models were used as input for training of network. And finally, all information was merged to get full detection results. Al-Shemarry et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] has introduced a 3L-LBP method for reading number plates from low quality images. The method given by them used three levels of preprocessing pattern classifiers to detect license plate localities and showed good accuracy. On images having challenges such as uneven contrast, dirt, fog and distortion.</p>
<p>The methodologies mentioned above have improved the quality of images by noise removal. However, the main drawback while using OCR is that OCR may get confused among characters and words in case of low-quality textual images. As a result, the above-discussed techniques lack the state-of-the-art in detecting text from low quality natural scene text images. Hence, the proposed technique handles the challenges of detecting and classifying textual data accurately from low-quality scene text images.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methodology</title>
<p>A process plan of the proposed framework is presented in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The proposed work consists of two parts: (1) Following the preprocessing of low-quality images, the K-Means is employed to extract three-level clusters of the query image. Afterwards, the detection of textual data is performed using MSER followed by character level grouping; (2) The connected components obtained are then classified into textual and non-textual regions using a CNN framework.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed framework process</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="fig-1.png"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Blurring/Deblurring Process</title>
<p>Blurry dataset for natural scene text detection and recognition are currently not publically available for analysis purpose [<xref ref-type="bibr" rid="ref-29">29</xref>]. Therefore, in this paper, syntactic dataset was adopted and apply on natural scene text detection datasets. To begin with idea, the original image <inline-formula id="ieqn-1">
<alternatives><inline-graphic xlink:href="ieqn-1.png"/><tex-math id="tex-ieqn-1"><![CDATA[$I\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-1"><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> is converted into blur image <inline-formula id="ieqn-2">
<alternatives><inline-graphic xlink:href="ieqn-2.png"/><tex-math id="tex-ieqn-2"><![CDATA[$\; B\left( {x,y} \right)\;$]]></tex-math><mml:math id="mml-ieqn-2"><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>B</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thickmathspace"></mml:mspace></mml:math>
</alternatives></inline-formula> by using following generative model.</p>
<p><disp-formula id="eqn-1">
<label>(1)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-1.png"/><tex-math id="tex-eqn-1"><![CDATA[$$B\left( {x,y} \right) = G\left( {i,j} \right)*I\left( {x,y} \right) + N\left( {x,y} \right)$$]]></tex-math><mml:math id="mml-eqn-1" display="block"><mml:mi>B</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x002B;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></disp-formula></p>
<p>where, <inline-formula id="ieqn-3">
<alternatives><inline-graphic xlink:href="ieqn-3.png"/><tex-math id="tex-ieqn-3"><![CDATA[$N\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-3"><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> represents additive noise and <inline-formula id="ieqn-4">
<alternatives><inline-graphic xlink:href="ieqn-4.png"/><tex-math id="tex-ieqn-4"><![CDATA[$G\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-4"><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> is known as two-dimensional blur kernel/PSF (Point Spread Function) in our case it is Gaussian kernel, and it is calculated using [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]:</p>
<p><disp-formula id="eqn-2">
<label>(2)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-2.png"/><tex-math id="tex-eqn-2"><![CDATA[$$G\left( {i,j} \right) = \displaystyle{1 \over {2\pi {\sigma ^2}}}{e^{ - \displaystyle{{{i^2} + {j^2}} \over {2{\sigma ^2}}}}}$$]]></tex-math><mml:math id="mml-eqn-2" display="block"><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:msup><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>&#x002B;</mml:mo><mml:mrow><mml:msup><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:mrow></mml:msup></mml:mrow></mml:mstyle></mml:math>
</alternatives></disp-formula></p>
<p>where <inline-formula id="ieqn-5">
<alternatives><inline-graphic xlink:href="ieqn-5.png"/><tex-math id="tex-ieqn-5"><![CDATA[$i$]]></tex-math><mml:math id="mml-ieqn-5"><mml:mi>i</mml:mi></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-6">
<alternatives><inline-graphic xlink:href="ieqn-6.png"/><tex-math id="tex-ieqn-6"><![CDATA[$j$]]></tex-math><mml:math id="mml-ieqn-6"><mml:mi>j</mml:mi></mml:math>
</alternatives></inline-formula> represent horizontal and vertical distances from origin, and <inline-formula id="ieqn-7">
<alternatives><inline-graphic xlink:href="ieqn-7.png"/><tex-math id="tex-ieqn-7"><![CDATA[$\sigma$]]></tex-math><mml:math id="mml-ieqn-7"><mml:mi>&#x03C3;</mml:mi></mml:math>
</alternatives></inline-formula> is known as standard deviation of Gaussian distribution in the paper value of sigma (<inline-formula id="ieqn-8">
<alternatives><inline-graphic xlink:href="ieqn-8.png"/><tex-math id="tex-ieqn-8"><![CDATA[$\sigma$]]></tex-math><mml:math id="mml-ieqn-8"><mml:mi>&#x03C3;</mml:mi></mml:math>
</alternatives></inline-formula>) to 1.76. Output of the stanstatic blure using Gaussian blure is presented in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Degradation synthetic blur input image <inline-formula id="ieqn-9">
<alternatives><inline-graphic xlink:href="ieqn-9.png"/><tex-math id="tex-ieqn-9"><![CDATA[$B\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-9"><mml:mi>B</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula></title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="fig-2.png"/>
</fig>
<p>The next step is to reverse the synthetic blur process. For this purpose in this paper, the de-blurring is taken out in the frequency domain rather than in spatial domain via Fast Fourier Transform (FFT) and Inverse FFT. Wiener filter acts as a filtration function to restore the image into a clear image. Considering, <inline-formula id="ieqn-10">
<alternatives><inline-graphic xlink:href="ieqn-10.png"/><tex-math id="tex-ieqn-10"><![CDATA[$B\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-10"><mml:mi>B</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> is blurry image; it can be converted into <inline-formula id="ieqn-11">
<alternatives><inline-graphic xlink:href="ieqn-11.png"/><tex-math id="tex-ieqn-11"><![CDATA[$\hat F\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-11"><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> image by using the following Wiener filter process [<xref ref-type="bibr" rid="ref-32">32</xref>].</p>
<p><disp-formula id="eqn-3">
<label>(3)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-3.png"/><tex-math id="tex-eqn-3"><![CDATA[$$\hat F\left( {x,y} \right) = W\left( {x,y} \right)\; B\left( {x,y} \right)$$]]></tex-math><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mi>W</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>B</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></disp-formula></p>
<p><disp-formula id="eqn-13">
<label>(4)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-13.png"/><tex-math id="tex-eqn-13"><![CDATA[$$W\left( {u,v} \right) = \displaystyle{{G*\left( {x,y} \right)} \over {{{\left| {G\left( {x,y} \right)} \right|}^2} + K\left( {x,y} \right)}}\;{\rm i.e.,}$$]]></tex-math><mml:math id="mml-eqn-13"><mml:mi>W</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>G</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>&#x002B;</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mtext> </mml:mtext><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mo>.</mml:mo><mml:mi mathvariant="normal">e</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:mstyle></mml:math>
</alternatives></disp-formula></p>
<p><disp-formula id="eqn-4">
<label>(5)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-4.png"/><tex-math id="tex-eqn-4"><![CDATA[$$W\left( {u,v} \right) = \left\{ {\matrix{ {1/G\left( {x,y} \right),\; \; when,\; K = 0\; Inverse\; filter} \cr {Higher\; frequency\; attenuated,\; \; K \ge \left| {G\left( {x,y} \right)} \right|} \cr } } \right.$$]]></tex-math><mml:math id="mml-eqn-4" display="block"><mml:mi>W</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>w</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>K</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mn>0</mml:mn><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>q</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mspace width="thickmathspace"></mml:mspace><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>K</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math>
</alternatives></disp-formula></p>
<p>where, <inline-formula id="ieqn-13">
<alternatives><inline-graphic xlink:href="ieqn-13.png"/><tex-math id="tex-ieqn-13"><![CDATA[$\hat F\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-13"><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> is inverse process and estimation of <inline-formula id="ieqn-14">
<alternatives><inline-graphic xlink:href="ieqn-14.png"/><tex-math id="tex-ieqn-14"><![CDATA[$I\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-14"><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> computed from <inline-formula id="ieqn-15">
<alternatives><inline-graphic xlink:href="ieqn-15.png"/><tex-math id="tex-ieqn-15"><![CDATA[$B\left( {x,y} \right)$]]></tex-math><mml:math id="mml-ieqn-15"><mml:mi>B</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula>. The output of the de-blurring images is displayed in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Recovered images after deblurring process</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="fig-3.png"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Text Localization</title>
<p>Text localization and detection is important part of any scene detection recognition proposes. In this paper, we have extended the standards MSER for text localization and detection from scenes it can be defined as [<xref ref-type="bibr" rid="ref-33">33</xref>]:</p>
<p><bold>Enhanced Image</bold> <inline-formula id="ieqn-16">
<alternatives><inline-graphic xlink:href="ieqn-16.png"/><tex-math id="tex-ieqn-16"><![CDATA[$\hat F$]]></tex-math><mml:math id="mml-ieqn-16"><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow></mml:math>
</alternatives></inline-formula> is a map <inline-formula id="ieqn-17">
<alternatives><inline-graphic xlink:href="ieqn-17.png"/><tex-math id="tex-ieqn-17"><![CDATA[$\hat F:m \subset {z^2} \to E$]]></tex-math><mml:math id="mml-ieqn-17"><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mo>:</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2282;</mml:mo><mml:mrow><mml:msup><mml:mi>z</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>E</mml:mi></mml:math>
</alternatives></inline-formula>. External region of an image is defined as if: <inline-formula id="ieqn-18">
<alternatives><inline-graphic xlink:href="ieqn-18.png"/><tex-math id="tex-ieqn-18"><![CDATA[$E$]]></tex-math><mml:math id="mml-ieqn-18"><mml:mi>E</mml:mi></mml:math>
</alternatives></inline-formula> is ordered as reflexive and transitive binary relation <inline-formula id="ieqn-19">
<alternatives><inline-graphic xlink:href="ieqn-19.png"/><tex-math id="tex-ieqn-19"><![CDATA[$\le$]]></tex-math><mml:math id="mml-ieqn-19"><mml:mo>&#x2264;</mml:mo></mml:math>
</alternatives></inline-formula> exists.in proposed work only <inline-formula id="ieqn-20">
<alternatives><inline-graphic xlink:href="ieqn-20.png"/><tex-math id="tex-ieqn-20"><![CDATA[$E = \left\{ {0,1,2,3, \ldots ,\; 255} \right\}$]]></tex-math><mml:math id="mml-ieqn-20"><mml:mi>E</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mn>255</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> is considered. Neighborhood relation <inline-formula id="ieqn-21">
<alternatives><inline-graphic xlink:href="ieqn-21.png"/><tex-math id="tex-ieqn-21"><![CDATA[$B \subset m \times m$]]></tex-math><mml:math id="mml-ieqn-21"><mml:mi>B</mml:mi><mml:mo>&#x2282;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi></mml:math>
</alternatives></inline-formula> is defined.</p>
<p><bold>Region</bold> <inline-formula id="ieqn-22">
<alternatives><inline-graphic xlink:href="ieqn-22.png"/><tex-math id="tex-ieqn-22"><![CDATA[$Q$]]></tex-math><mml:math id="mml-ieqn-22"><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula> is a contiguous subset of <inline-formula id="ieqn-23">
<alternatives><inline-graphic xlink:href="ieqn-23.png"/><tex-math id="tex-ieqn-23"><![CDATA[$m$]]></tex-math><mml:math id="mml-ieqn-23"><mml:mi>m</mml:mi></mml:math>
</alternatives></inline-formula> for each <inline-formula id="ieqn-24">
<alternatives><inline-graphic xlink:href="ieqn-24.png"/><tex-math id="tex-ieqn-24"><![CDATA[$p,q \in Q$]]></tex-math><mml:math id="mml-ieqn-24"><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula>. For each there is a sequence <inline-formula id="ieqn-25">
<alternatives><inline-graphic xlink:href="ieqn-25.png"/><tex-math id="tex-ieqn-25"><![CDATA[$p,\; {\alpha _1},{\alpha _2}, \ldots ,\; {\alpha _n}$]]></tex-math><mml:math id="mml-ieqn-25"><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula>.</p>
<p><bold>Region Boundary</bold> <inline-formula id="ieqn-26">
<alternatives><inline-graphic xlink:href="ieqn-26.png"/><tex-math id="tex-ieqn-26"><![CDATA[$\partial Q = \left\{ {q \in m|Q:\exists \; p \in Q:qBp} \right\}$]]></tex-math><mml:math id="mml-ieqn-26"><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>Q</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>q</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>Q</mml:mi><mml:mo>:</mml:mo><mml:mi mathvariant="normal">&#x2203;</mml:mi><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>p</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Q</mml:mi><mml:mo>:</mml:mo><mml:mi>q</mml:mi><mml:mi>B</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> the boundary <inline-formula id="ieqn-27">
<alternatives><inline-graphic xlink:href="ieqn-27.png"/><tex-math id="tex-ieqn-27"><![CDATA[$\partial Q$]]></tex-math><mml:math id="mml-ieqn-27"><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula> of <inline-formula id="ieqn-28">
<alternatives><inline-graphic xlink:href="ieqn-28.png"/><tex-math id="tex-ieqn-28"><![CDATA[$Q$]]></tex-math><mml:math id="mml-ieqn-28"><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula> is the set of pixels that are adjacent to at least on pixel.</p>
<p><bold>External Region</bold> <inline-formula id="ieqn-29">
<alternatives><inline-graphic xlink:href="ieqn-29.png"/><tex-math id="tex-ieqn-29"><![CDATA[$Q \subset m$]]></tex-math><mml:math id="mml-ieqn-29"><mml:mi>Q</mml:mi><mml:mo>&#x2282;</mml:mo><mml:mi>m</mml:mi></mml:math>
</alternatives></inline-formula> is a region such that <inline-formula id="ieqn-30">
<alternatives><inline-graphic xlink:href="ieqn-30.png"/><tex-math id="tex-ieqn-30"><![CDATA[$p \in Q$]]></tex-math><mml:math id="mml-ieqn-30"><mml:mi>p</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-31">
<alternatives><inline-graphic xlink:href="ieqn-31.png"/><tex-math id="tex-ieqn-31"><![CDATA[$Q \in \partial Q$]]></tex-math><mml:math id="mml-ieqn-31"><mml:mi>Q</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula> the boundary <inline-formula id="ieqn-32">
<alternatives><inline-graphic xlink:href="ieqn-32.png"/><tex-math id="tex-ieqn-32"><![CDATA[$Q \in \partial Q:\hat F\left( p \right) > \hat F\left( q \right)$]]></tex-math><mml:math id="mml-ieqn-32"><mml:mi>Q</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>Q</mml:mi><mml:mo>:</mml:mo><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003E;</mml:mo><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> is maximum intensity region.</p>
<p><bold>Maximally Stable Extremal Regions (MSER):</bold> Let <inline-formula id="ieqn-33">
<alternatives><inline-graphic xlink:href="ieqn-33.png"/><tex-math id="tex-ieqn-33"><![CDATA[${Q_1}, \ldots .,{Q_{i - 1}},{Q_i}$]]></tex-math><mml:math id="mml-ieqn-33"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> be a sequence of nested external regions <inline-formula id="ieqn-34">
<alternatives><inline-graphic xlink:href="ieqn-34.png"/><tex-math id="tex-ieqn-34"><![CDATA[${Q_i} \subset {Q_{i + 1}}$]]></tex-math><mml:math id="mml-ieqn-34"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2282;</mml:mo><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula>. External region <inline-formula id="ieqn-35">
<alternatives><inline-graphic xlink:href="ieqn-35.png"/><tex-math id="tex-ieqn-35"><![CDATA[${Q_{i*}}$]]></tex-math><mml:math id="mml-ieqn-35"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> is maximally stable if <inline-formula id="ieqn-36">
<alternatives><inline-graphic xlink:href="ieqn-36.png"/><tex-math id="tex-ieqn-36"><![CDATA[${q_i} = \left| {{Q_{i + {\rm \Delta }\; }}|\; {Q_{i - {\rm \Delta }}}\; |{Q_i}} \right|$]]></tex-math><mml:math id="mml-ieqn-36"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mspace width="thickmathspace"></mml:mspace></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mspace width="thickmathspace"></mml:mspace><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mspace width="thickmathspace"></mml:mspace><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> has a local minimum <inline-formula id="ieqn-37">
<alternatives><inline-graphic xlink:href="ieqn-37.png"/><tex-math id="tex-ieqn-37"><![CDATA[${\rm \Delta } \in E$]]></tex-math><mml:math id="mml-ieqn-37"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mi>E</mml:mi></mml:math>
</alternatives></inline-formula> is a parameter of the method.</p>
<p>MSER algorithm that distinguishes the best-quality text candidates from stable areas that are extracted from different color channel images. Multi-resolution MSER gives better work to enormous scope changes and blurred images, which improves coordinating execution over large scale changes and for blurred images. To allow for detecting small letters in images of limited resolution, the complementary properties of canny edges and MSER are combined in our edge-enhanced MSER. In this paper, the K-means technique is extended before text detection to get more robust to distinct text from non-text for text detection. The output of MSER to K-means is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Separating foreground from background using K-means applied on the MSER images, where <inline-formula id="ieqn-38">
<alternatives><inline-graphic xlink:href="ieqn-38.png"/><tex-math id="tex-ieqn-38"><![CDATA[$k = 3$]]></tex-math><mml:math id="mml-ieqn-38"><mml:mi>k</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mn>3</mml:mn></mml:math>
</alternatives></inline-formula> is settled for obtaining clusters</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="fig-4.png"/>
</fig>
<p>In 2018 Yi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] have proposed a K-Means clustering for finding color of roadside replacement. They have used blur images, daylight or night time images, also images taken in bad weather, not bright or in the shadow. They have modified the images in LUV CIE color space. Inspired by this work, the proposed methodology employed to improve the quality of natural textual image. Since in many of scene text images we find blur which disrupts the character formation. Hence, in the proposed technique, K-Means seems to be the better choice not only to improve the image quality, but at the same time, it works fine for separating the foreground and background. K-Means converts the images into clusters based on colours. In the proposed methodology <inline-formula id="ieqn-39">
<alternatives><inline-graphic xlink:href="ieqn-39.png"/><tex-math id="tex-ieqn-39"><![CDATA[$k$]]></tex-math><mml:math id="mml-ieqn-39"><mml:mi>k</mml:mi></mml:math>
</alternatives></inline-formula> is settled to 3 that is <inline-formula id="ieqn-40">
<alternatives><inline-graphic xlink:href="ieqn-40.png"/><tex-math id="tex-ieqn-40"><![CDATA[${k_1}$]]></tex-math><mml:math id="mml-ieqn-40"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula>, <inline-formula id="ieqn-41">
<alternatives><inline-graphic xlink:href="ieqn-41.png"/><tex-math id="tex-ieqn-41"><![CDATA[${k_2}$]]></tex-math><mml:math id="mml-ieqn-41"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-42">
<alternatives><inline-graphic xlink:href="ieqn-42.png"/><tex-math id="tex-ieqn-42"><![CDATA[${k_3}$]]></tex-math><mml:math id="mml-ieqn-42"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula>. This approach is practical for text detection in the later stage.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Character Level Grouping and Text Detection</title>
<p>This section is based on the bounding box of detected characters from low-quality image. The standards are based on the geometric properties, which means widths <italic>(w</italic><sub><italic>1</italic></sub>, <italic>w</italic><sub><italic>2</italic></sub><italic>)</italic> and heights <italic>(h</italic><sub><italic>1</italic></sub>, <italic>h</italic><sub><italic>2</italic></sub><italic>)</italic> of bounding boxes will be considered for correct identification of the word. By definition of geometric properties, <xref ref-type="disp-formula" rid="eqn-6">Eqs. (6)</xref>, <xref ref-type="disp-formula" rid="eqn-7"> (7)</xref> and <xref ref-type="disp-formula" rid="eqn-8">(8)</xref> are given as:</p>
<p><disp-formula id="eqn-5">
<label>(6)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-5.png"/><tex-math id="tex-eqn-5"><![CDATA[$$h = min\left( {{h_1},{h_2}} \right)$$]]></tex-math><mml:math id="mml-eqn-5" display="block"><mml:mi>h</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></disp-formula></p>
<p><disp-formula id="eqn-6">
<label>(7)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-6.png"/><tex-math id="tex-eqn-6"><![CDATA[$$\partial x = \left| {{l_1} + {l_2}} \right| - \left( {{w_1} + {w_2}} \right)/2$$]]></tex-math><mml:math id="mml-eqn-6" display="block"><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>x</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x002B;</mml:mo><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x002B;</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:math>
</alternatives></disp-formula></p>
<p><disp-formula id="eqn-7">
<label>(8)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-7.png"/><tex-math id="tex-eqn-7"><![CDATA[$$\partial y = \left| {{m_1} - {m_2}} \right|$$]]></tex-math><mml:math id="mml-eqn-7" display="block"><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>y</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>m</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>m</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math>
</alternatives></disp-formula></p>
<p>where, <inline-formula id="ieqn-43">
<alternatives><inline-graphic xlink:href="ieqn-43.png"/><tex-math id="tex-ieqn-43"><![CDATA[$\partial x\;$]]></tex-math><mml:math id="mml-ieqn-43"><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>x</mml:mi><mml:mspace width="thickmathspace"></mml:mspace></mml:math>
</alternatives></inline-formula> will always be negative when two boxes correspond in <inline-formula id="ieqn-44">
<alternatives><inline-graphic xlink:href="ieqn-44.png"/><tex-math id="tex-ieqn-44"><![CDATA[$x$]]></tex-math><mml:math id="mml-ieqn-44"><mml:mi>x</mml:mi></mml:math>
</alternatives></inline-formula> direction. So they are well-suited and assumed as belonging to the same text. Hence, the explanation of character-bounding can easily be limited if conditions below are satisfied:</p>
<p><disp-formula id="eqn-8">
<label>(9)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-8.png"/><tex-math id="tex-eqn-8"><![CDATA[$$\left| {{h_1} - {h_2}} \right| < {k_1}h$$]]></tex-math><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mi>h</mml:mi></mml:math>
</alternatives></disp-formula></p>
<p><disp-formula id="eqn-9">
<label>(10)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-9.png"/><tex-math id="tex-eqn-9"><![CDATA[$$\partial x < {k_2}h$$]]></tex-math><mml:math id="mml-eqn-9" display="block"><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mi>h</mml:mi></mml:math>
</alternatives></disp-formula></p>
<p><disp-formula id="eqn-10">
<label>(11)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-10.png"/><tex-math id="tex-eqn-10"><![CDATA[$$\partial y < {k_3}h$$]]></tex-math><mml:math id="mml-eqn-10" display="block"><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>y</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mi>h</mml:mi></mml:math>
</alternatives></disp-formula></p>
<p>where, <inline-formula id="ieqn-45">
<alternatives><inline-graphic xlink:href="ieqn-45.png"/><tex-math id="tex-ieqn-45"><![CDATA[${k_1},\; {k_2}$]]></tex-math><mml:math id="mml-ieqn-45"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula>, and <inline-formula id="ieqn-46">
<alternatives><inline-graphic xlink:href="ieqn-46.png"/><tex-math id="tex-ieqn-46"><![CDATA[${k_3}$]]></tex-math><mml:math id="mml-ieqn-46"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> are parameters for elements of character bounding. The third parameter <inline-formula id="ieqn-47">
<alternatives><inline-graphic xlink:href="ieqn-47.png"/><tex-math id="tex-ieqn-47"><![CDATA[${k_3}$]]></tex-math><mml:math id="mml-ieqn-47"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> is necessary for determining the group as character and non-character. A similar process is directed on all detected characters. Character bounding output is presented in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> below:</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Character level grouping for different K-means images</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="fig-5.png"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Connected Components Classification Using CNN Framework</title>
<p>For straightforwardness, the primary function of this section is the classification of connected components in classes of text and non-text with labels (0/1). The main objective is to reduce FPR and eliminates FNR. Keeping the vague nature of images, a supervised CNN model with multi-layers is presented to obtain information of character, segmentation of character region, information of binary text and non-text. Additionally, precise features of textual data from low-level segmentation to high-level binary classification are found with the help of additional features. So the proposed model becomes powerful for understanding what, where and whether of character for taking advantage of formulating a consistent conclusion. Even though CNN is non-trivial because of information levels containing difficulties in learning and convergence rates. Hence, it proves to be suitable for sharing features. For <italic>N</italic> training examples represented by <inline-formula id="ieqn-48">
<alternatives><inline-graphic xlink:href="ieqn-48.png"/><tex-math id="tex-ieqn-48"><![CDATA[$\mathop \sum \limits_{k = 1}^N \left( {{x_k},{y_k}} \right)$]]></tex-math><mml:math id="mml-ieqn-48"><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> main goal of CNN is minimizing the following <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref> and also make sure that <inline-formula id="ieqn-49">
<alternatives><inline-graphic xlink:href="ieqn-49.png"/><tex-math id="tex-ieqn-49"><![CDATA[${x_k}$]]></tex-math><mml:math id="mml-ieqn-49"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> is the image patch and <inline-formula id="ieqn-50">
<alternatives><inline-graphic xlink:href="ieqn-50.png"/><tex-math id="tex-ieqn-50"><![CDATA[${y_k}$]]></tex-math><mml:math id="mml-ieqn-50"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> is the label that maps to &#x2018;0&#x2019; for non-text and &#x2018;1&#x2019; for text.</p>
<p><disp-formula id="eqn-11">
<label>(12)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-11.png"/><tex-math id="tex-eqn-11"><![CDATA[$$\arg \mathop {\min }\limits_W \mathop \sum \limits_{k = 1}^N \gamma \left( {{y_k},f\left( {{x_k},W} \right)} \right) + {\rm \Delta }W$$]]></tex-math><mml:math id="mml-eqn-11" display="block"><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mrow><mml:mo form="prefix">min</mml:mo></mml:mrow><mml:mi>W</mml:mi></mml:munder><mml:mo>&#x2061;</mml:mo><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>W</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x002B;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>W</mml:mi></mml:math>
</alternatives></disp-formula></p>
<p>where <inline-formula id="ieqn-51">
<alternatives><inline-graphic xlink:href="ieqn-51.png"/><tex-math id="tex-ieqn-51"><![CDATA[$f\left( {{x_k},W} \right)$]]></tex-math><mml:math id="mml-ieqn-51"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>W</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> in <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref> is a function whose parameter is <inline-formula id="ieqn-52">
<alternatives><inline-graphic xlink:href="ieqn-52.png"/><tex-math id="tex-ieqn-52"><![CDATA[$W.\; \gamma \left( . \right)$]]></tex-math><mml:math id="mml-ieqn-52"><mml:mi>W</mml:mi><mml:mo>.</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>.</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula> signifies loss function that is classically a soft-max loss for task of classification and least square loss for task of regression. <inline-formula id="ieqn-53">
<alternatives><inline-graphic xlink:href="ieqn-53.png"/><tex-math id="tex-ieqn-53"><![CDATA[$\gamma$]]></tex-math><mml:math id="mml-ieqn-53"><mml:mi>&#x03B3;</mml:mi></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-54">
<alternatives><inline-graphic xlink:href="ieqn-54.png"/><tex-math id="tex-ieqn-54"><![CDATA[${\rm \Delta }W$]]></tex-math><mml:math id="mml-ieqn-54"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>W</mml:mi></mml:math>
</alternatives></inline-formula> work as learning rate and regularization, respectively, and shown in <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>. The training procedure tries to implement binary classification, which finds a function of mapping to connect image patch as input with labels as output that is 0/1 lacking additional information.</p>
<p><disp-formula id="eqn-12">
<label>(13)</label>
<alternatives>
<graphic mimetype="image" mime-subtype="png" xlink:href="eqn-12.png"/><tex-math id="tex-eqn-12"><![CDATA[$${\rm \Delta }W = \left( {\mathop \sum \limits_{j = 1}^N \parallel{w^{\left[ j \right]}}\parallel^2} \right)\displaystyle{\varphi \over {2M}}$$]]></tex-math><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>W</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:munderover><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mo>&#x2061;</mml:mo><mml:msup><mml:mrow><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mi>j</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mi>&#x03C6;</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>M</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</alternatives></disp-formula></p>
<p>Here <inline-formula id="ieqn-55">
<alternatives><inline-graphic xlink:href="ieqn-55.png"/><tex-math id="tex-ieqn-55"><![CDATA[$M$]]></tex-math><mml:math id="mml-ieqn-55"><mml:mi>M</mml:mi></mml:math>
</alternatives></inline-formula> denotes the number of inputs, <inline-formula id="ieqn-56">
<alternatives><inline-graphic xlink:href="ieqn-56.png"/><tex-math id="tex-ieqn-56"><![CDATA[$N$]]></tex-math><mml:math id="mml-ieqn-56"><mml:mi>N</mml:mi></mml:math>
</alternatives></inline-formula> presents number of layers, <inline-formula id="ieqn-57">
<alternatives><inline-graphic xlink:href="ieqn-57.png"/><tex-math id="tex-ieqn-57"><![CDATA[$\varphi$]]></tex-math><mml:math id="mml-ieqn-57"><mml:mi>&#x03C6;</mml:mi></mml:math>
</alternatives></inline-formula> is regularization parameter, and <inline-formula id="ieqn-58">
<alternatives><inline-graphic xlink:href="ieqn-58.png"/><tex-math id="tex-ieqn-58"><![CDATA[${w^{\left[ j \right]}}$]]></tex-math><mml:math id="mml-ieqn-58"><mml:mrow><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mi>j</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:math>
</alternatives></inline-formula> denotes weight matrix of <inline-formula id="ieqn-59">
<alternatives><inline-graphic xlink:href="ieqn-59.png"/><tex-math id="tex-ieqn-59"><![CDATA[${j^{th}}$]]></tex-math><mml:math id="mml-ieqn-59"><mml:mrow><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math>
</alternatives></inline-formula> layer. The primary goal is to recognize image patch <inline-formula id="ieqn-60">
<alternatives><inline-graphic xlink:href="ieqn-60.png"/><tex-math id="tex-ieqn-60"><![CDATA[${x_k}$]]></tex-math><mml:math id="mml-ieqn-60"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> containing label over the labels <inline-formula id="ieqn-61">
<alternatives><inline-graphic xlink:href="ieqn-61.png"/><tex-math id="tex-ieqn-61"><![CDATA[${y_k}$]]></tex-math><mml:math id="mml-ieqn-61"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula>. The advantage of <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref> is to avoid over fitting. Sometimes, during training process, the training error is reduced but testing error remains constant. Conversely, model is trained well but not produces expected results. So, regularization is a technique, which allows making specific changes in the learning algorithm such that the model generalized better and work fine on the unseen data. A stochastic gradient learning procedure is used to train a CNN model. A large number of CNN models have used this algorithm. The sequential optimization from the regression of low level to a binary classification of high level is the main feature of CNN model. This approach is more appropriate for identifying text and non-text components. So training of CNN model is done with the positive samples taken from Char74k, cropped images from ICDAR 2003, SVT and IIIT5K. Significant amount of distracters are also part of the training process.</p>
<p>The proposed model of CNN shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> above as formulation used for binary classification. A series of image patches is used as input in the model, and every image is to be classified into text and non-text and labeled as &#x2018;1&#x2019; for text or &#x2018;0&#x2019; for non-text respectively. Two convolution layers with filters <italic>f</italic><sub><italic>1</italic></sub> and <italic>f</italic><sub><italic>2</italic></sub> are used by this network. The <inline-formula id="ieqn-62">
<alternatives><inline-graphic xlink:href="ieqn-62.png"/><tex-math id="tex-ieqn-62"><![CDATA[${f_1} = 78$]]></tex-math><mml:math id="mml-ieqn-62"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mn>78</mml:mn></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-63">
<alternatives><inline-graphic xlink:href="ieqn-63.png"/><tex-math id="tex-ieqn-63"><![CDATA[${f_2} = 216$]]></tex-math><mml:math id="mml-ieqn-63"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>&#x003D;</mml:mo><mml:mn>216</mml:mn></mml:math>
</alternatives></inline-formula> filters are used to extract deep feature. A supervised learning process is used for training network by stochastic gradient descent of binary image patches with a given size of 26 &#x00D7; 26. A window size of <inline-formula id="ieqn-64">
<alternatives><inline-graphic xlink:href="ieqn-64.png"/><tex-math id="tex-ieqn-64"><![CDATA[$11 \times 11$]]></tex-math><mml:math id="mml-ieqn-64"><mml:mn>11</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>11</mml:mn></mml:math>
</alternatives></inline-formula> is slide over the <inline-formula id="ieqn-65">
<alternatives><inline-graphic xlink:href="ieqn-65.png"/><tex-math id="tex-ieqn-65"><![CDATA[$26 \times 26$]]></tex-math><mml:math id="mml-ieqn-65"><mml:mn>26</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>26</mml:mn></mml:math>
</alternatives></inline-formula> image patch in order to extract features to create input vectors <inline-formula id="ieqn-66">
<alternatives><inline-graphic xlink:href="ieqn-66.png"/><tex-math id="tex-ieqn-66"><![CDATA[${x^{\left( k \right)}} \in {\rm {\mathbb{R}}\; }$]]></tex-math><mml:math id="mml-ieqn-66"><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mspace width="thickmathspace"></mml:mspace></mml:mrow></mml:math>
</alternatives></inline-formula> where <inline-formula id="ieqn-67">
<alternatives><inline-graphic xlink:href="ieqn-67.png"/><tex-math id="tex-ieqn-67"><![CDATA[$k \in \left\{ {1,2, \ldots n} \right\}$]]></tex-math><mml:math id="mml-ieqn-67"><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula>. Multiple low-level filters <inline-formula id="ieqn-68">
<alternatives><inline-graphic xlink:href="ieqn-68.png"/><tex-math id="tex-ieqn-68"><![CDATA[$D \in {{\rm {\mathbb{R}}}^{64 \times f1}}$]]></tex-math><mml:math id="mml-ieqn-68"><mml:mi>D</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>f</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math>
</alternatives></inline-formula> are trained by Stochastic Gradient Descent (SGD). For each 11 &#x00D7; 11 patch <inline-formula id="ieqn-69">
<alternatives><inline-graphic xlink:href="ieqn-69.png"/><tex-math id="tex-ieqn-69"><![CDATA[$x$]]></tex-math><mml:math id="mml-ieqn-69"><mml:mi>x</mml:mi></mml:math>
</alternatives></inline-formula>, the first layer response <inline-formula id="ieqn-70">
<alternatives><inline-graphic xlink:href="ieqn-70.png"/><tex-math id="tex-ieqn-70"><![CDATA[$Q$]]></tex-math><mml:math id="mml-ieqn-70"><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula> is calculated by the implementation of an inner product with filter pool followed by scalar activation function: <inline-formula id="ieqn-71">
<alternatives><inline-graphic xlink:href="ieqn-71.png"/><tex-math id="tex-ieqn-71"><![CDATA[$Q = maximum\; \left\{ {0,\left| {{D^T}x} \right| - \beta } \right\}$]]></tex-math><mml:math id="mml-ieqn-71"><mml:mi>Q</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi><mml:mspace width="thickmathspace"></mml:mspace><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mi>D</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math>
</alternatives></inline-formula>, where <inline-formula id="ieqn-72">
<alternatives><inline-graphic xlink:href="ieqn-72.png"/><tex-math id="tex-ieqn-72"><![CDATA[$\beta = 0.55$]]></tex-math><mml:math id="mml-ieqn-72"><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mn>0.55</mml:mn></mml:math>
</alternatives></inline-formula> denotes hyper parameter. For 26 &#x00D7; 26 image patch, <inline-formula id="ieqn-73">
<alternatives><inline-graphic xlink:href="ieqn-73.png"/><tex-math id="tex-ieqn-73"><![CDATA[$Q$]]></tex-math><mml:math id="mml-ieqn-73"><mml:mi>Q</mml:mi></mml:math>
</alternatives></inline-formula> for every 11 &#x00D7; 11 window size is computed to get a response map of <inline-formula id="ieqn-74">
<alternatives><inline-graphic xlink:href="ieqn-74.png"/><tex-math id="tex-ieqn-74"><![CDATA[$16\times16\times f1$]]></tex-math><mml:math id="mml-ieqn-74"><mml:mn>16</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>16</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>f</mml:mi><mml:mn>1</mml:mn></mml:math>
</alternatives></inline-formula>. Then, the reduction of response map up to <inline-formula id="ieqn-75">
<alternatives><inline-graphic xlink:href="ieqn-75.png"/><tex-math id="tex-ieqn-75"><![CDATA[$4\times4\times f1$]]></tex-math><mml:math id="mml-ieqn-75"><mml:mn>4</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>4</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>f</mml:mi><mml:mn>1</mml:mn></mml:math>
</alternatives></inline-formula> is made by applying average pool. Same process is directed on second convolutional layer to get reduced response map of <inline-formula id="ieqn-76">
<alternatives><inline-graphic xlink:href="ieqn-76.png"/><tex-math id="tex-ieqn-76"><![CDATA[$2\times2\times f2$]]></tex-math><mml:math id="mml-ieqn-76"><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>f</mml:mi><mml:mn>2</mml:mn></mml:math>
</alternatives></inline-formula>. Then, outputs are fully connected to classification layer. Training error is minimized with the help of back-propagating by applying SGD with an unchanged filter size throughout the classification process.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Text Correction after Recognition</title>
<p>In the proposed model presented in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> every time, a character label is the output of the CNN model. The extracted labels are collected into a text file to a complete word and need to process for text correction. It is necessary because sometimes, CNN recognized the false label for a character, which might change the semantic meaning of the extracted text. Thus, the aim raised to recognize the correct scene text digitally, which can be further used in various IoT based applications. To tackle this issue, a hamming distance (an error correction technique) is used to correct the semantic meaning of scene text. It is defined as given below:</p>
<p>Given two vectors <inline-formula id="ieqn-77">
<alternatives><inline-graphic xlink:href="ieqn-77.png"/><tex-math id="tex-ieqn-77"><![CDATA[${V_1}$]]></tex-math><mml:math id="mml-ieqn-77"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-78">
<alternatives><inline-graphic xlink:href="ieqn-78.png"/><tex-math id="tex-ieqn-78"><![CDATA[${V_2} \in {Z^n}$]]></tex-math><mml:math id="mml-ieqn-78"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:msup><mml:mi>Z</mml:mi><mml:mi>n</mml:mi></mml:msup></mml:mrow></mml:math>
</alternatives></inline-formula>, the hamming distance between <inline-formula id="ieqn-79">
<alternatives><inline-graphic xlink:href="ieqn-79.png"/><tex-math id="tex-ieqn-79"><![CDATA[${V_1}\;$]]></tex-math><mml:math id="mml-ieqn-79"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mspace width="thickmathspace"></mml:mspace></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-80">
<alternatives><inline-graphic xlink:href="ieqn-80.png"/><tex-math id="tex-ieqn-80"><![CDATA[${V_2}$]]></tex-math><mml:math id="mml-ieqn-80"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> is defined as <inline-formula id="ieqn-81">
<alternatives><inline-graphic xlink:href="ieqn-81.png"/><tex-math id="tex-ieqn-81"><![CDATA[$d({V_1},\; {V_2})$]]></tex-math><mml:math id="mml-ieqn-81"><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace"></mml:mspace><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math>
</alternatives></inline-formula> to the number of places where <inline-formula id="ieqn-82">
<alternatives><inline-graphic xlink:href="ieqn-82.png"/><tex-math id="tex-ieqn-82"><![CDATA[${V_1}$]]></tex-math><mml:math id="mml-ieqn-82"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-83">
<alternatives><inline-graphic xlink:href="ieqn-83.png"/><tex-math id="tex-ieqn-83"><![CDATA[${V_2}$]]></tex-math><mml:math id="mml-ieqn-83"><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> differ. As per definition, it is clear that hamming distance is the number of bits changed in the observed string. The stored labels (complete word) in the text file are related to any natural scene text image for they are recognized. These labels are treated as a string for further processing. Therefore, the string is searched using the lexicon to calculate the Hamming distance. If the value of Hamming distance is 0, then it concludes that labels in the text file that are a complete word match with natural scene text. On the other hand, if hamming distance value results &#x003E; 0, then it reflects that there exist false labels in the text, which do not match with the natural scene text. The hamming distance value also represents the total number of false recognized labels. In addition to this, the process also lists the optimized words which might have correct word scene text word. Finally, it clearly reflects that the proposed approach for scene text correction is very useful and worked effectively to generate correct scene text when few errors occur during recognizing and labelling in CNN.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Discussions</title>
<p>In this section, proposed methodologies are evaluated based on various tests for their performance. The following sections discuss the experimental setup and implementation details, datasets description and evaluation metrics used in this context. Furthermore, comparisons of the results with existing benchmark techniques are the central part of this section. All the results are based on average calculations of the evaluation standards.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Evaluation Standards</title>
<p>The following evaluation metrics are used in this study and are given in <xref ref-type="table" rid="table-1">Tab. 1</xref>.</p>

<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Evaluation metrics for scene text detection and classification</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Evaluation metric</th>
<th>Mathematical representation</th>
</tr>
</thead>
<tbody>
<tr>
<td>Accuracy</td>
<td><inline-formula id="ieqn-84">
<alternatives><inline-graphic xlink:href="ieqn-84.png"/><tex-math id="tex-ieqn-84"><![CDATA[$Accuracy = \displaystyle{{TP + TN} \over {TP + TN + FP + FN}}$]]></tex-math><mml:math id="mml-ieqn-84"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</alternatives></inline-formula></td>
</tr>
<tr>
<td>Precision or<break/>predictive positive</td>
<td><inline-formula id="ieqn-85">
<alternatives><inline-graphic xlink:href="ieqn-85.png"/><tex-math id="tex-ieqn-85"><![CDATA[$PPR = \displaystyle{{TP} \over {TP + FP}}$]]></tex-math><mml:math id="mml-ieqn-85"><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</alternatives></inline-formula></td>
</tr>
<tr>
<td>Recall or<break/>sensitivity or true positive</td>
<td><inline-formula id="ieqn-86">
<alternatives><inline-graphic xlink:href="ieqn-86.png"/><tex-math id="tex-ieqn-86"><![CDATA[$TPR = \displaystyle{{TP} \over {TP + FN}}$]]></tex-math><mml:math id="mml-ieqn-86"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</alternatives></inline-formula></td>
</tr>
<tr>
<td>F-measure or F1-score</td>
<td><inline-formula id="ieqn-87">
<alternatives><inline-graphic xlink:href="ieqn-87.png"/><tex-math id="tex-ieqn-87"><![CDATA[$F\; Measure = 2\displaystyle{{TPR.PPR} \over {TPR + PPT}}$]]></tex-math><mml:math id="mml-ieqn-87"><mml:mi>F</mml:mi><mml:mspace width="thickmathspace"></mml:mspace><mml:mi>M</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mn>2</mml:mn><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>.</mml:mo><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</alternatives></inline-formula></td>
</tr>
<tr>
<td>Specificity or<break/>true negative</td>
<td><inline-formula id="ieqn-88">
<alternatives><inline-graphic xlink:href="ieqn-88.png"/><tex-math id="tex-ieqn-88"><![CDATA[$TNR = \displaystyle{{TN} \over {TN + FP}}$]]></tex-math><mml:math id="mml-ieqn-88"><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</alternatives></inline-formula></td>
</tr>
<tr>
<td>False Positive</td>
<td><inline-formula id="ieqn-89">
<alternatives><inline-graphic xlink:href="ieqn-89.png"/><tex-math id="tex-ieqn-89"><![CDATA[$FPR = \displaystyle{{FP} \over {TN + FP}}$]]></tex-math><mml:math id="mml-ieqn-89"><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x003D;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x002B;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</alternatives></inline-formula></td>
</tr>
<tr>
<td>Area under the curve (AUC)</td>
<td><inline-formula id="ieqn-90">
<alternatives><inline-graphic xlink:href="ieqn-90.png"/><tex-math id="tex-ieqn-90"><![CDATA[$\mathop \int \nolimits_{ - \infty }^{ + \infty } TPR\left( T \right)FPR\left( T \right)dt$]]></tex-math><mml:math id="mml-ieqn-90"><mml:msubsup><mml:mrow><mml:mo largeop="false">&#x222B;</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x002B;</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:math>
</alternatives></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Datasets Description</title>
<p>The datasets evaluated in this study are the benchmark and publically available. These include ICDAR 2003, SVT, and IIIT5K. All datasets are challenging, which present text on the scene having different background scenic properties. Moreover, text also reflects various characteristics like random orientations, low contrast/resolution, blurry, hazy and arbitrary shapes and sizes. The descriptions of these datasets are given below:<list list-type="bullet"><list-item>
<p><bold><italic>Chars 74K:</italic></bold> The Chars 74k dataset is a collection of 7705 images comprised of English alphabetic characters, i.e., A to Z, 0 to 9 and a to z in Sixty-two (62) classes. Along with 647 classes of Kannada native language 3345 characters, which are segmented from 1922 scene text images manually [<xref ref-type="bibr" rid="ref-35">35</xref>].</p></list-item><list-item>
<p><bold><italic>ICDAR 2003/2005:</italic></bold> The ICDAR 2003 dataset was released for ICDAR 2003 Robust Reading Competition by Lucas et al. [<xref ref-type="bibr" rid="ref-36">36</xref>]. The same dataset with no change is used in the ICDAR 2005 Competition of Robust Reading. Therefore occasionally, the dataset is known as ICDAR 2005 [<xref ref-type="bibr" rid="ref-37">37</xref>]. It is a collection of 251 testing and 258 training character patches and word patches annotated by the bounded box and their text contents.</p></list-item><list-item>
<p><bold><italic>SVT:</italic></bold> The Street View Text dataset (SVT) was used explicitly for word spotting problems. This is a collection of 647 words from which 250 testing images (video frames) with the availability of bounding box locations and ground truth labels along with 100 training images (video frames). Also, the lexicon for each word is also available, and almost 50 words lexicon for each word is integrated. Each image is taken from Google Street View [<xref ref-type="bibr" rid="ref-38">38</xref>].</p></list-item><list-item>
<p><bold><italic>IIIT5K:</italic></bold> It is the biggest and most challenging dataset reported to date due to variation is font, color, layout, size and inclusion of noise, distortion, blur and varying illuminations. IIIT5K Word dataset [<xref ref-type="bibr" rid="ref-39">39</xref>] is a group of 5000 words collected from images found on the Internet, from which 3000 and 2000 words used to test and train subsets correspondingly.</p></list-item></list></p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Connected Components (CC) Classification Results Using Proposed CNN Framework</title>
<p>The classification of connected components is an important phase, which is done with the help of implemented CNN framework. For doing this, the classification is taken out with various distributions of the datasets into training and testing sets. Dataset is distributed with a ratio of 50&#x2013;50, 60&#x2013;40, 70&#x2013;30 and 80&#x2013;20 into training-testing samples. The results are computed on each dataset separately. The foremost objective is to test and monitor the classifier performance and also to overcome the false positives. It is observed that the classifier performance is improved gradually on all evaluation standards as an increase in the distribution samples. The reason is very simple that all deep learning algorithms are data-hungry algorithms. As a result, the model acts effectively for evaluation of unseen data.</p>
<p>In <xref ref-type="table" rid="table-2">Tab. 2</xref>, as per expectation, ICDAR 2003 attains 84.0% accuracy level with the highest distribution level that is 80&#x2013;20. At the same time, significant improvement is also recorded to overcome false positive rate (FPR). The improvements in parameters advance at once, when a training sample is increased, which support the above notion. The other noteworthy factor is deduced, that increased training samples help to reduce FPR gradually. Also, false negative rate (FNR) is minimized, which improves the classifier accuracy gradually. In <xref ref-type="table" rid="table-3">Tab. 3</xref>, SVT responses with the best accuracy level of 84.3%. However, significant improvement is gradually recorded in other parameters. The associated fact with SVT is that it is a complex dataset with almost a collection of all outdoor images. These images incorporate diversified scene text properties along with multiple sets of objects in the same image. On the other hand, IIIT5K performed well with the best result of accuracy level that is 90.8% shown in <xref ref-type="table" rid="table-4">Tab. 4</xref>. This fact is unexpected because IIIT5K is the most tricky and challenging dataset reported till now. IIIT5K is a collection of images with variant illumination, low contrast/resolution, and random orientations. Despite these challenges, only 0.086% FPR is reported, showing the performance of proposed method that it works fine.</p>

<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Results of connected components classification on ICDAR 2003 dataset, where PPR denote positive predictive value, TPR denote true positive value, ACC denotes accuracy, and AUC denotes area under the curve</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Training&#x2013;testing</th>
<th>PPR</th>
<th>TPR</th>
<th>F1</th>
<th>FPR</th>
<th>FNR</th>
<th>ACC</th>
<th>AUC</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>50&#x2013;50</bold></td>
<td>61.3</td>
<td>58.1</td>
<td>59.6</td>
<td>0.359</td>
<td>28.8</td>
<td>71.2</td>
<td>0.641</td>
</tr>
<tr>
<td><bold>60&#x2013;40</bold></td>
<td>64.9</td>
<td>60.2</td>
<td>62.4</td>
<td>0.292</td>
<td>24.3</td>
<td>75.7</td>
<td>0.708</td>
</tr>
<tr>
<td><bold>70&#x2013;30</bold></td>
<td>67.7</td>
<td>62.1</td>
<td>64.7</td>
<td>0.239</td>
<td>21.5</td>
<td>78.5</td>
<td>0.761</td>
</tr>
<tr>
<td><bold>80&#x2013;20</bold></td>
<td><bold>75.5</bold></td>
<td><bold>70.9</bold></td>
<td><bold>73.1</bold></td>
<td><bold>0.169</bold></td>
<td><bold>16.0</bold></td>
<td><bold>84.0</bold></td>
<td><bold>0.831</bold></td>
</tr>
</tbody>
</table>
</table-wrap>

<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Results of connected components classification on SVT dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Training&#x2013;testing</th>
<th>PPR</th>
<th>TPR</th>
<th>F1</th>
<th>FPR</th>
<th>FNR</th>
<th>ACC</th>
<th>AUC</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>50&#x2013;50</bold></td>
<td>62.1</td>
<td>58.2</td>
<td>60.1</td>
<td>0.399</td>
<td>29.1</td>
<td>70.9</td>
<td>0.601</td>
</tr>
<tr>
<td><bold>60&#x2013;40</bold></td>
<td>63.5</td>
<td>62.3</td>
<td>62.9</td>
<td>0.319</td>
<td>27.4</td>
<td>72.6</td>
<td>0.681</td>
</tr>
<tr>
<td><bold>70&#x2013;30</bold></td>
<td>68.2</td>
<td>65.9</td>
<td>67.0</td>
<td>0.223</td>
<td>23.1</td>
<td>76.9</td>
<td>0.777</td>
</tr>
<tr>
<td><bold>80&#x2013;20</bold></td>
<td><bold>80.1</bold></td>
<td><bold>76.2</bold></td>
<td><bold>78.1</bold></td>
<td><bold>0.129</bold></td>
<td><bold>15.7</bold></td>
<td><bold>84.3</bold></td>
<td><bold>0.871</bold></td>
</tr>
</tbody>
</table>
</table-wrap>

<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Results of connected components classification on IIIT5K dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Training&#x2013;testing</th>
<th>PPR</th>
<th>TPR</th>
<th>F1</th>
<th>FPR</th>
<th>FNR</th>
<th>ACC</th>
<th>AUC</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>50&#x2013;50</bold></td>
<td>65.1</td>
<td>62.2</td>
<td>63.6</td>
<td>0.289</td>
<td>27.7</td>
<td>72.3</td>
<td>0.711</td>
</tr>
<tr>
<td><bold>60&#x2013;40</bold></td>
<td>68.6</td>
<td>63.9</td>
<td>66.0</td>
<td>0.193</td>
<td>21.5</td>
<td>78.5</td>
<td>0.807</td>
</tr>
<tr>
<td><bold>70&#x2013;30</bold></td>
<td>75.2</td>
<td>70.4</td>
<td>72.7</td>
<td>0.180</td>
<td>18.6</td>
<td>81.4</td>
<td>0.820</td>
</tr>
<tr>
<td><bold>80&#x2013;20</bold></td>
<td><bold>88.5</bold></td>
<td><bold>82.7</bold></td>
<td><bold>85.5</bold></td>
<td><bold>0.086</bold></td>
<td><bold>9.2</bold></td>
<td><bold>90.8</bold></td>
<td><bold>0.914</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Connected Components (CC) Classification Using Deep and Separately Extracted Features</title>
<p>Taking a different scenario, the proposed technique is tested with separately extracted features on all selected mainstream datasets for connected component classifications. The purpose of this test is to present the legitimacy of the proposed technique, which extract deep features at two levels and then classify the connected components. The particular test (<italic>separately extracted features</italic>) is performed to visualize the classification process using support vector machine (SVM) and its variants (Linear-SVM, Cubic-SVM, and Quad-SVM) on extracted feature vector based on histogram oriented gradients (HOG), local binary patterns (LBP), and Geometric separately. Some other classifiers like Decision Tree (DT) and K-Nearest Neighbor (KNN) is also computed on the same set of features.</p>
<p>For this test, constant 80&#x2013;20 sample distribution is adapted for all the competitive classifiers. Moreover, two levels of deep features maps after pooling layer 1 and pooling layer 2 named <italic>Deep Features 1</italic> and <italic>Deep Features 2</italic> is extracted for this, respectively. Both of these features sets are also fed to predetermined benchmark classifiers for monitoring their performance along with other separately extracted features. The results show that the pre-determined benchmark classifier works fine when compared to other extracted feature vectors. During this test, it is observed that Linear-SVM is the second good performer among all the other benchmark classifiers on all mainstream datasets. However, none of the classifiers beats the proposed technique at any level of deep features. Moreover, to avoid the biasness, each experiment is repeated ten times, and the mean score is reported in the <xref ref-type="table" rid="table-5">Tabs. 5</xref>, and <xref ref-type="table" rid="table-7">7</xref>. In <xref ref-type="table" rid="table-5">Tab. 5</xref>, the extracted <italic>Deep Features 1</italic> and <italic>Deep Features 2</italic> outperforms and gives 86.2% and 89.6% accuracy on ICDAR 2003, while SVT presented 84.9% and 85.3% accuracy on both levels of CNN features which is shown in <xref ref-type="table" rid="table-6">Tab. 6</xref>. In <xref ref-type="table" rid="table-7">Tab. 7</xref>, IIIT5K exhibits an accuracy level of 88.2% and 90.8% correspondingly. Henceforth, it is also confirmed that IIIT5K produced the best results on two levels of CNN features. Additionally, <xref ref-type="table" rid="table-5">Tabs. 5</xref>, <xref ref-type="table" rid="table-6">6</xref>, and <xref ref-type="table" rid="table-7">7</xref>, clarify that all other benchmark classifiers also perform well on CNN features when compared to handcrafted features. Similarly, geometric features work fine and give intense competition to CNN features.</p>

<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Accuracy (%) results on ICDAR 2003 using separate feature extraction</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Classifiers</th>
<th>Appearance<break/>(HOG)</th>
<th>Texture<break/>(LBP)</th>
<th>Geometric<break/>features</th>
<th><italic>Deep</italic><break/><italic>features 1</italic></th>
<th><italic>Deep</italic><break/><italic>features 2</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td>L-SVM</td>
<td>71.2</td>
<td>73.3</td>
<td>81.1</td>
<td><bold>83.2</bold></td>
<td><bold>85.5</bold></td>
</tr>
<tr>
<td>C-SVM</td>
<td>70.3</td>
<td>71.3</td>
<td>79.6</td>
<td>82.7</td>
<td>83.6</td>
</tr>
<tr>
<td>Q-SVM</td>
<td>69.8</td>
<td>72.7</td>
<td>80.3</td>
<td>82.1</td>
<td>84.3</td>
</tr>
<tr>
<td>KNN</td>
<td>72.9</td>
<td>70.4</td>
<td>78.7</td>
<td>81.8</td>
<td>82.5</td>
</tr>
<tr>
<td>DT</td>
<td>68.6</td>
<td>69.4</td>
<td>77.5</td>
<td>80.7</td>
<td>80.9</td>
</tr>
<tr>
<td><bold>Proposed model (Softmax)</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td><bold>86.2</bold></td>
<td><bold>89.0</bold></td>
</tr>
</tbody>
</table>
</table-wrap>

<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Accuracy (%) results on SVT using separate feature extraction</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Classifiers</th>
<th>Appearance<break/>(HOG)</th>
<th>Texture<break/>(LBP)</th>
<th>Geometric<break/>features</th>
<th><italic>Deep</italic><break/><italic>features 1</italic></th>
<th><italic>Deep</italic><break/><italic>features 2</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td>L-SVM</td>
<td><bold>70.6</bold></td>
<td><bold>72.4</bold></td>
<td><bold>82.1</bold></td>
<td><bold>83.9</bold></td>
<td><bold>84.9</bold></td>
</tr>
<tr>
<td>C-SVM</td>
<td>69.1</td>
<td>71.3</td>
<td>80.2</td>
<td>82.4</td>
<td>83.1</td>
</tr>
<tr>
<td>Q-SVM</td>
<td>71.2</td>
<td>73.1</td>
<td>68.6</td>
<td>81.8</td>
<td>82.3</td>
</tr>
<tr>
<td>KNN</td>
<td>72.2</td>
<td>69.9</td>
<td>76.7</td>
<td>79.7</td>
<td>81.3</td>
</tr>
<tr>
<td>DT</td>
<td>67.3</td>
<td>70.2</td>
<td>77.7</td>
<td>80.1</td>
<td>80.9</td>
</tr>
<tr>
<td><bold>Proposed model (Softmax)</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td><bold>84.9</bold></td>
<td><bold>85.3</bold></td>
</tr>
</tbody>
</table>
</table-wrap>

<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Accuracy (%) results on IIIT5K using separate feature extraction</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Classifiers</th>
<th>Appearance<break/>(HOG)</th>
<th>Texture<break/>(LBP)</th>
<th>Geometric<break/>features</th>
<th><italic>CNN</italic><break/><italic>features 1</italic></th>
<th><italic>CNN</italic><break/><italic>features 2</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td>L-SVM</td>
<td><bold>71.3</bold></td>
<td><bold>72.1</bold></td>
<td><bold>81.2</bold></td>
<td><bold>85.1</bold></td>
<td><bold>85.7</bold></td>
</tr>
<tr>
<td>C-SVM</td>
<td>70.1</td>
<td>70.6</td>
<td>79.9</td>
<td>82.3</td>
<td>84.7</td>
</tr>
<tr>
<td>Q-SVM</td>
<td>69.6</td>
<td>71.2</td>
<td>78.1</td>
<td>80.5</td>
<td>83.3</td>
</tr>
<tr>
<td>KNN</td>
<td>73.3</td>
<td>72.3</td>
<td>78.4</td>
<td>83.1</td>
<td>84.5</td>
</tr>
<tr>
<td>DT</td>
<td>68.5</td>
<td>70.3</td>
<td>79.8</td>
<td>79.3</td>
<td>81.4</td>
</tr>
<tr>
<td><bold>Proposed model (Softmax)</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td><bold>88.2</bold></td>
<td><bold>90.8</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Comparison of Text Detection on Original Image and Clustered Image</title>
<p>The results in the above <xref ref-type="table" rid="table-7">Tab. 7</xref> that the detection of characters becomes easy after applying K-Means on images. In most of the images, MSER cannot detect all characters, but after applying clustering on the images, MSER detects characters without false positive values. The best results can be obtained using <inline-formula id="ieqn-91">
<alternatives><inline-graphic xlink:href="ieqn-91.png"/><tex-math id="tex-ieqn-91"><![CDATA[${k_1}$]]></tex-math><mml:math id="mml-ieqn-91"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> clustered images. However, in some cases, both <inline-formula id="ieqn-92">
<alternatives><inline-graphic xlink:href="ieqn-92.png"/><tex-math id="tex-ieqn-92"><![CDATA[${k_1}$]]></tex-math><mml:math id="mml-ieqn-92"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-93">
<alternatives><inline-graphic xlink:href="ieqn-93.png"/><tex-math id="tex-ieqn-93"><![CDATA[${k_2}$]]></tex-math><mml:math id="mml-ieqn-93"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> gave perfect results for detection. Therefore, it can be concluded that <inline-formula id="ieqn-94">
<alternatives><inline-graphic xlink:href="ieqn-94.png"/><tex-math id="tex-ieqn-94"><![CDATA[${k_1}$]]></tex-math><mml:math id="mml-ieqn-94"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> gives the best results while using <inline-formula id="ieqn-95">
<alternatives><inline-graphic xlink:href="ieqn-95.png"/><tex-math id="tex-ieqn-95"><![CDATA[${k_2}$]]></tex-math><mml:math id="mml-ieqn-95"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-96">
<alternatives><inline-graphic xlink:href="ieqn-96.png"/><tex-math id="tex-ieqn-96"><![CDATA[${k_3}$]]></tex-math><mml:math id="mml-ieqn-96"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> MSER can detect some characters only. It is also evident that detection becomes easier after applying clustering on images. In the <xref ref-type="table" rid="table-8">Tab. 8</xref> below, MSER works better on clustered images as compared to original images.</p>

<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>k-means based MSER character detection results</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="table-8.png"/>
</table-wrap>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Comparison of Proposed Text Extraction Method with Existing Benchmark Technologies</title>
<p>To check eminence of proposed method, it is compared with existing benchmark methods proposed by different researchers in recent times. However, it is also a challenging process due to the heterogeneous nature of datasets, parameters, and natural scene text characteristics. Cater to these challenges, the commonalities are spotted in the evaluation metrics. The most common parameters used in this regard are positive predicted rate (PPR), true positive rate (TPR) and F1 score. The results are obtained on three benchmark datasets and organized in <xref ref-type="table" rid="table-9">Tab. 9</xref>, which shows comparison results on ICDAR 2003, SVT and IIIT5K. This table worth mentioning comparisons of the proposed methodology with other benchmark techniques in the same domain with the same set of datasets. It is noticeable from <xref ref-type="table" rid="table-9">Tab. 9</xref> that few techniques are found to employ IIIT5K datasets for text extraction. The reason is IIIT5K is very challenging due to font size variation, color, layout, distortion occurrence, varying illumination, blur and noise. However, the proposed methodology produced PPR 0.85%, TPR 0.77% and F1 0.80% for IIIT5K. On widely used SVT complex dataset, this demonstrates high variability of outdoor scene text images. The proposed methodology remarkably performs by attaining the PPR 0.84%, TPR 0.76% and F1 0.80%. Similarly, for other datasets such ICDAR 2003, the proposed technique outperforms with values of PPR 0.85%, TPR 0.77% and F1 0.81%.</p>

<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>(%) Results of proposed text extraction method with benchmark existing methods on ICDAR 2003, SVT and IIIT5K</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th></th><th colspan="4">ICDAR 2003</th><th colspan="3">SVT</th><th colspan="3">IIIT5K</th>
</tr>
<tr>
<td>Methods</td>
<td>Year</td>
<td>PPR</td>
<td>TPR</td>
<td>F1</td>
<td>PPR</td>
<td>TPR</td>
<td>F1</td>
<td>PPR</td>
<td>TPR</td>
<td>F1</td>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>2017</td>
<td>0.83</td>
<td>0.69</td>
<td>0.75</td>
<td>0.37</td>
<td>0.47</td>
<td>0.41</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>2018</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>0.81</td>
<td>0.77</td>
<td>0.79</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>2018</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>0.69</td>
<td>0.61</td>
<td>0.65</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
<td>2018</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>0.73</td>
<td>0.69</td>
<td>0.71</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>2019</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>0.80</td>
<td>0.71</td>
<td>0.75</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Proposed</td>
<td>2020</td>
<td>0.85</td>
<td>0.77</td>
<td>0.81</td>
<td>0.84</td>
<td>0.76</td>
<td>0.80</td>
<td>0.85</td>
<td>0.77</td>
<td>0.80</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From the above-mentioned results, It is obvious that the proposed technique works fine and is more stable on the mainstream datasets when compared to other benchmark counterparts in terms of F1.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>In this paper, a state-of-the-art technique to extract text from low quality natural scene images is presented. The query image is synthetically blurred using averaging filter. Further L&#x002A;a&#x002A;b color space is adapted for enhancing contrast followed by deblurring of images, where wiener filter is utilized as filtration function. Then MSER is applied for localizing and detecting text regions, while non-text areas are discarded with geometric properties. K-Means clustering is applied for better separation of foreground from background and it also has less false positive rate. In this process, <inline-formula id="ieqn-100">
<alternatives><inline-graphic xlink:href="ieqn-100.png"/><tex-math id="tex-ieqn-100"><![CDATA[${\rm k}$]]></tex-math><mml:math id="mml-ieqn-100"><mml:mrow><mml:mi mathvariant="normal">k</mml:mi></mml:mrow></mml:math>
</alternatives></inline-formula> is settled to 3. Therefore, it can be seen that <inline-formula id="ieqn-101">
<alternatives><inline-graphic xlink:href="ieqn-101.png"/><tex-math id="tex-ieqn-101"><![CDATA[${k_1}$]]></tex-math><mml:math id="mml-ieqn-101"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> and <inline-formula id="ieqn-102">
<alternatives><inline-graphic xlink:href="ieqn-102.png"/><tex-math id="tex-ieqn-102"><![CDATA[${k_2}$]]></tex-math><mml:math id="mml-ieqn-102"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> gives the best results for text detection, but <inline-formula id="ieqn-103">
<alternatives><inline-graphic xlink:href="ieqn-103.png"/><tex-math id="tex-ieqn-103"><![CDATA[${k_3}$]]></tex-math><mml:math id="mml-ieqn-103"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> can detect only some characters. This problem can also be solved in future work by using SWT technique on <inline-formula id="ieqn-104">
<alternatives><inline-graphic xlink:href="ieqn-104.png"/><tex-math id="tex-ieqn-104"><![CDATA[${k_3}$]]></tex-math><mml:math id="mml-ieqn-104"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:math>
</alternatives></inline-formula> images to improve the results. The classification results are obtained on various distributions of training and testing sets. Furthermore, a different scenario that is separately extracted features and CNN features is also adapted to monitor the credibility of the CNN model. It is concluded that the proposed methodology works fine and responds well to all types of tests. Finally, it is observed that the proposed work outperforms in detecting text when compared with previous models in the same domain.</p>
</sec>
</body>
<back>
<ack>
<p>HITEC University Taxila</p>
</ack><fn-group>
<fn fn-type="conflict">
<p><bold>Funding Statement:</bold> This research was supported by the MSIT(Ministry of Science and ICT), Korea, under the ICAN(ICT Challenge and Advanced Network of HRD) program(IITP-2020-0-01832) supervised by the IITP (Institute of Information &#x0026; Communications Technology Planning &#x0026; Evaluation) and the Soonchunhyang University Research Fund.</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1">
<label>1</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>H.</given-names> 
<surname>Arshad</surname></string-name>, <string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>M. I.</given-names> 
<surname>Sharif</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Yasmin</surname></string-name>, <string-name>
<given-names>J. M. R. S.</given-names> 
<surname>Tavares</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>A multilevel paradigm for deep convolutional neural network features selection with an application to human gait recognition</article-title>,&#x201D; 
<source>Expert Systems</source>, vol. 
<volume>21</volume>, no. 
<issue>3</issue>, pp. 
<fpage>e12541</fpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-2">
<label>2</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>M. I.</given-names> 
<surname>Sharif</surname></string-name>, <string-name>
<given-names>J. P.</given-names> 
<surname>Li</surname></string-name>, <string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name> and <string-name>
<given-names>M. A.</given-names> 
<surname>Saleem</surname></string-name>
</person-group>, &#x201C;
<article-title>Active deep neural network features selection for segmentation and recognition of brain tumors using MRI images</article-title>,&#x201D; 
<source>Pattern Recognition Letters</source>, vol. 
<volume>129</volume>, pp. 
<fpage>181</fpage>&#x2013;
<lpage>189</lpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-3">
<label>3</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>M.</given-names> 
<surname>Rashid</surname></string-name>, <string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Alhaisoni</surname></string-name>, <string-name>
<given-names>S. H.</given-names> 
<surname>Wang</surname></string-name> and <string-name>
<given-names>S. R.</given-names> 
<surname>Naqvi</surname></string-name>
</person-group>, &#x201C;
<article-title>A sustainable deep learning framework for object recognition using multi-layers deep features fusion and selection</article-title>,&#x201D; 
<source>Sustainability</source>, vol. 
<volume>12</volume>, no. 
<issue>12</issue>, pp. 
<fpage>5037</fpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-4">
<label>4</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>D.</given-names> 
<surname>Karatzas</surname></string-name>, <string-name>
<given-names>L.</given-names> 
<surname>Gomez-Bigorda</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>Nicolaou</surname></string-name>, <string-name>
<given-names>S.</given-names> 
<surname>Ghosh</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>Bagdanov</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;<chapter-title>ICDAR 2015 competition on robust reading</chapter-title>,&#x201D; in <conf-name>2015 13th Int. Conf. on Document Analysis and Recognition</conf-name>, 
<publisher-loc>Tunis, Tunisia</publisher-loc>, pp. 
<fpage>1156</fpage>&#x2013;
<lpage>1160</lpage>, 
<year>2015</year>.</mixed-citation>
</ref>
<ref id="ref-5">
<label>5</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>D.</given-names> 
<surname>Karatzas</surname></string-name>, <string-name>
<given-names>F.</given-names> 
<surname>Shafait</surname></string-name>, <string-name>
<given-names>S.</given-names> 
<surname>Uchida</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Iwamura</surname></string-name>, <string-name>
<given-names>L. G. </given-names> 
<surname>i Bigorda</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>ICDAR 2013 robust reading competition</article-title>,&#x201D; in <conf-name>2013 12th Int. Conf. on Document Analysis and Recognition</conf-name>, 
<publisher-loc>Washington, DC, USA</publisher-loc>, pp. 
<fpage>1484</fpage>&#x2013;
<lpage>1493</lpage>, 
<year>2013</year>.</mixed-citation>
</ref>
<ref id="ref-6">
<label>6</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>A.</given-names> 
<surname>Shahab</surname></string-name>, <string-name>
<given-names>F.</given-names> 
<surname>Shafait</surname></string-name> and <string-name>
<given-names>A.</given-names> 
<surname>Dengel</surname></string-name>
</person-group>, &#x201C;
<article-title>ICDAR, 2011 robust reading competition challenge 2: Reading text in scene images</article-title>,&#x201D; in <conf-name>2011 Int. Conf. on Document Analysis and Recognition</conf-name>, 
<publisher-loc>Beijing, China</publisher-loc>, pp. 
<fpage>1491</fpage>&#x2013;
<lpage>1496</lpage>, 
<year>2011</year>. </mixed-citation>
</ref>
<ref id="ref-7">
<label>7</label><mixed-citation publication-type="other">
<person-group person-group-type="author"><string-name>
<given-names>M. E.</given-names> 
<surname>Maros</surname></string-name>, <string-name>
<given-names>C. G.</given-names> 
<surname>Cho</surname></string-name>, <string-name>
<given-names>A. G.</given-names> 
<surname>Junge</surname></string-name>, <string-name>
<given-names>B.</given-names> 
<surname>K&#x00E4;mpgen</surname></string-name>, <string-name>
<given-names>V.</given-names> 
<surname>Saase</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Comparative analysis of machine learning algorithms for computer-assisted reporting based on fully automated cross-lingual RadLex&#x00AE; mappings</article-title>,&#x201D; 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-8">
<label>8</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>Y.</given-names> 
<surname>Liu</surname></string-name>, <string-name>
<given-names>C.</given-names> 
<surname>Yang</surname></string-name>, <string-name>
<given-names>L.</given-names> 
<surname>Jiang</surname></string-name>, <string-name>
<given-names>S.</given-names> 
<surname>Xie</surname></string-name> and <string-name>
<given-names>Y.</given-names> 
<surname>Zhang</surname></string-name>
</person-group>, &#x201C;
<article-title>Intelligent edge computing for IoT-based energy management in smart cities</article-title>,&#x201D; 
<source>IEEE Network</source>, vol. 
<volume>33</volume>, no. 
<issue>2</issue>, pp. 
<fpage>111</fpage>&#x2013;
<lpage>117</lpage>, 
<year>2019</year>.</mixed-citation>
</ref>
<ref id="ref-9">
<label>9</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>R. S.</given-names> 
<surname>Alonso</surname></string-name>, <string-name>
<given-names>I.</given-names> 
<surname>Sitt&#x00F3;n-Candanedo</surname></string-name>, <string-name>
<given-names>&#x00D3;</given-names> 
<surname>Garc&#x00ED;a</surname></string-name>, <string-name>
<given-names>J.</given-names> 
<surname>Prieto</surname></string-name> and <string-name>
<given-names>S.</given-names> 
<surname>Rodr&#x00ED;guez-Gonz&#x00E1;lez</surname></string-name>
</person-group>, &#x201C;
<article-title>An intelligent edge-IoT platform for monitoring livestock and crops in a dairy farming scenario</article-title>,&#x201D; 
<source>Ad Hoc Networks</source>, vol. 
<volume>98</volume>, pp. 
<lpage>102047</lpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-10">
<label>10</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>P.</given-names> 
<surname>Lyu</surname></string-name>, <string-name>
<given-names>C.</given-names> 
<surname>Yao</surname></string-name>, <string-name>
<given-names>W.</given-names> 
<surname>Wu</surname></string-name>, <string-name>
<given-names>S.</given-names> 
<surname>Yan</surname></string-name> and <string-name>
<given-names>X.</given-names> 
<surname>Bai</surname></string-name>
</person-group>, &#x201C;
<article-title>Multi-oriented scene text detection via corner localization and region segmentation</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, 
<publisher-loc>Beijing, China</publisher-loc>, pp. 
<fpage>7553</fpage>&#x2013;
<lpage>7563</lpage>, 
<year>2018</year>. </mixed-citation>
</ref>
<ref id="ref-11">
<label>11</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>K.</given-names> 
<surname>Javed</surname></string-name>, <string-name>
<given-names>S. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>T.</given-names> 
<surname>Saba</surname></string-name>, <string-name>
<given-names>U.</given-names> 
<surname>Habib</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Human action recognition using fusion of multiview and deep features: An application to video surveillance</article-title>,&#x201D; 
<source>Multimedia Tools and Applications</source>, vol. 
<volume>10</volume>, pp. 
<fpage>335</fpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-12">
<label>12</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>A.</given-names> 
<surname>Majid</surname></string-name>, <string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Yasmin</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>Rehman</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>Yousafzai</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Classification of stomach infections: A paradigm of convolutional neural network along with classical features fusion and selection</article-title>,&#x201D; 
<source>Microscopy Research and Technique</source>, vol. 
<volume>83</volume>, no. 
<issue>5</issue>, pp. 
<fpage>562</fpage>&#x2013;
<lpage>576</lpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-13">
<label>13</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>F.</given-names> 
<surname>Ahmed</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Mittal</surname></string-name>, <string-name>
<given-names>L. M.</given-names> 
<surname>Goyal</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Gastrointestinal diseases segmentation and classification based on duo-deep architectures</article-title>,&#x201D; 
<source>Pattern Recognition Letters</source>, vol. 
<volume>131</volume>, pp. 
<fpage>193</fpage>&#x2013;
<lpage>204</lpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-14">
<label>14</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>F. E.</given-names> 
<surname>Batool</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Attique</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Sharif</surname></string-name>, <string-name>
<given-names>K.</given-names> 
<surname>Javed</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Nazir</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Offline signature verification system: A novel technique of fusion of GLCM and geometric features using SVM</article-title>,&#x201D; 
<source>Multimedia Tools and Applications</source>, pp. 
<fpage>1</fpage>&#x2013;
<lpage>20</lpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-15">
<label>15</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>T.</given-names> 
<surname>Akram</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Sharif</surname></string-name> and <string-name>
<given-names>T.</given-names> 
<surname>Saba</surname></string-name>
</person-group>, &#x201C;
<article-title>Fruits diseases classification: Exploiting a hierarchical framework for deep features fusion and selection</article-title>,&#x201D; 
<source>Multimedia Tools and Applications</source>, pp. 
<fpage>1</fpage>&#x2013;
<lpage>21</lpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-16">
<label>16</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>A.</given-names> 
<surname>Adeel</surname></string-name>, <string-name>
<given-names>M. A.</given-names> 
<surname>Khan</surname></string-name>, <string-name>
<given-names>T.</given-names> 
<surname>Akram</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>Sharif</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Yasmin</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Entropy&#x2010;controlled deep features selection framework for grape leaf diseases recognition</article-title>,&#x201D; 
<source>Expert Systems</source>, vol. 
<volume>1</volume>, no. 
<issue>1</issue>, pp. 
<fpage>1</fpage>, 
<year>2020</year>.</mixed-citation>
</ref>
<ref id="ref-17">
<label>17</label><mixed-citation publication-type="other">
<person-group person-group-type="author"><string-name>
<given-names>Q.</given-names> 
<surname>Yang</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Cheng</surname></string-name>, <string-name>
<given-names>W.</given-names> 
<surname>Zhou</surname></string-name>, <string-name>
<given-names>Y.</given-names> 
<surname>Chen</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Qiu</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Inceptext: A new inception-text module with deformable psroi pooling for multi-oriented scene text detection</article-title>,&#x201D; 
<comment>arXiv preprint arXiv: 1805. 01167</comment>, 
<year>2018</year>.</mixed-citation>
</ref>
<ref id="ref-18">
<label>18</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>M. S.</given-names> 
<surname>Das</surname></string-name>, <string-name>
<given-names>B. H.</given-names> 
<surname>Bindhu</surname></string-name> and <string-name>
<given-names>A.</given-names> 
<surname>Govardhan</surname></string-name>
</person-group>, &#x201C;
<article-title>Evaluation of text detection and localization methods in natural images</article-title>,&#x201D; 
<source>International Journal of Emerging Technology and Advanced Engineering</source>, vol. 
<volume>2</volume>, no. 
<issue>6</issue>, pp. 
<fpage>277</fpage>&#x2013;
<lpage>282</lpage>, 
<year>2012</year>.</mixed-citation>
</ref>
<ref id="ref-19">
<label>19</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>J.</given-names> 
<surname>Matas</surname></string-name>, <string-name>
<given-names>O.</given-names> 
<surname>Chum</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Urban</surname></string-name> and <string-name>
<given-names>T.</given-names> 
<surname>Pajdla</surname></string-name>
</person-group>, &#x201C;
<article-title>Robust wide-baseline stereo from maximally stable extremal regions</article-title>,&#x201D; 
<source>Image and Vision Computing</source>, vol. 
<volume>22</volume>, no. 
<issue>10</issue>, pp. 
<fpage>761</fpage>&#x2013;
<lpage>767</lpage>, 
<year>2004</year>.</mixed-citation>
</ref>
<ref id="ref-20">
<label>20</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>A.</given-names> 
<surname>Criminisi</surname></string-name>, <string-name>
<given-names>P.</given-names> 
<surname>P&#x00E9;rez</surname></string-name> and <string-name>
<given-names>K.</given-names> 
<surname>Toyama</surname></string-name>
</person-group>, &#x201C;
<article-title>Region filling and object removal by exemplar-based image inpainting</article-title>,&#x201D; 
<source>IEEE Transactions on Image Processing</source>, vol. 
<volume>13</volume>, no. 
<issue>9</issue>, pp. 
<fpage>1200</fpage>&#x2013;
<lpage>1212</lpage>, 
<year>2004</year>.</mixed-citation>
</ref>
<ref id="ref-21">
<label>21</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>S.</given-names> 
<surname>Zhao</surname></string-name>, <string-name>
<given-names>Y.</given-names> 
<surname>Wang</surname></string-name> and <string-name>
<given-names>Y.</given-names> 
<surname>Wang</surname></string-name>
</person-group>, &#x201C;
<article-title>Extracting hand vein patterns from low-quality images: A new biometric technique using low-cost devices</article-title>,&#x201D; in <conf-name>Fourth Int. Conf. on Image and Graphics</conf-name>, 
<publisher-loc>Sichuan, China</publisher-loc>, pp. 
<fpage>667</fpage>&#x2013;
<lpage>671</lpage>, 
<year>2007</year>. </mixed-citation>
</ref>
<ref id="ref-22">
<label>22</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>D. J.</given-names> 
<surname>Ittner</surname></string-name>, <string-name>
<given-names>D. D.</given-names> 
<surname>Lewis</surname></string-name> and <string-name>
<given-names>D. D.</given-names> 
<surname>Ahn</surname></string-name>
</person-group>, &#x201C;
<article-title>Text categorization of low quality images</article-title>,&#x201D; in <conf-name>Sym. on Document Analysis and Information Retrieval</conf-name>, pp. 
<fpage>301</fpage>&#x2013;
<lpage>315</lpage>, 
<year>1995</year>. </mixed-citation>
</ref>
<ref id="ref-23">
<label>23</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>K.</given-names> 
<surname>Iqbal</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Odetayo</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>James</surname></string-name>, <string-name>
<given-names>R. A.</given-names> 
<surname>Salam</surname></string-name> and <string-name>
<given-names>A. Z. H.</given-names> 
<surname>Talib</surname></string-name>
</person-group>, &#x201C;
<article-title>Enhancing the low quality images using unsupervised colour correction method</article-title>,&#x201D; in <conf-name>2010 IEEE Int. Conf. on Systems, Man and Cybernetics</conf-name>, 
<publisher-loc>Istanbul, Turkey</publisher-loc>, pp. 
<fpage>1703</fpage>&#x2013;
<lpage>1709</lpage>, 
<year>2010</year>. </mixed-citation>
</ref>
<ref id="ref-24">
<label>24</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>S.</given-names> 
<surname>Rudrani</surname></string-name> and <string-name>
<given-names>S.</given-names> 
<surname>Das</surname></string-name>
</person-group>, &#x201C;
<article-title>Face recognition on low quality surveillance images, by compensating degradation</article-title>,&#x201D; in <conf-name>Int. Conf. Image Analysis and Recognition</conf-name>, 
<publisher-loc>Berlin, Heidelberg</publisher-loc>: 
<publisher-name>Springer</publisher-name>, pp. 
<fpage>212</fpage>&#x2013;
<lpage>221</lpage>, 
<year>2011</year>. </mixed-citation>
</ref>
<ref id="ref-25">
<label>25</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>N.</given-names> 
<surname>Neverova</surname></string-name>, <string-name>
<given-names>D.</given-names> 
<surname>Muselet</surname></string-name> and <string-name>
<given-names>A.</given-names> 
<surname>Tr&#x00E9;meau</surname></string-name>
</person-group>, &#x201C;
<article-title>Lighting estimation in indoor environments from low-quality images</article-title>,&#x201D; in <conf-name>European Conf. on Computer Vision</conf-name>, 
<publisher-loc>Berlin, Heidelberg</publisher-loc>: 
<publisher-name>Springer</publisher-name>, pp. 
<fpage>380</fpage>&#x2013;
<lpage>389</lpage>, 
<year>2012</year>. </mixed-citation>
</ref>
<ref id="ref-26">
<label>26</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>V. T. H.</given-names> 
<surname>Tuyet</surname></string-name> and <string-name>
<given-names>N. T.</given-names> 
<surname>Binh</surname></string-name>
</person-group>, &#x201C;
<article-title>Edge detection in low quality medical images</article-title>,&#x201D; in <conf-name>Int. Conf. on Nature of Computation and Communication</conf-name>, 
<publisher-loc>Cham</publisher-loc>: 
<publisher-name>Springer</publisher-name>, pp. 
<fpage>351</fpage>&#x2013;
<lpage>362</lpage>, 
<year>2016</year>. </mixed-citation>
</ref>
<ref id="ref-27">
<label>27</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>C.</given-names> 
<surname>Zhu</surname></string-name>, <string-name>
<given-names>T. H.</given-names> 
<surname>Li</surname></string-name> and <string-name>
<given-names>G.</given-names> 
<surname>Li</surname></string-name>
</person-group>, &#x201C;
<article-title>Towards automatic wild animal detection in low quality camera-trap images using two-channeled perceiving residual pyramid networks</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Int. Conf. on Computer Vision Workshops</conf-name>, 
<publisher-loc>Sichuan, China</publisher-loc>, pp. 
<fpage>2860</fpage>&#x2013;
<lpage>2864</lpage>, 
<year>2017</year>. </mixed-citation>
</ref>
<ref id="ref-28">
<label>28</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>M. S.</given-names> 
<surname>Al-Shemarry</surname></string-name>, <string-name>
<given-names>Y.</given-names> 
<surname>Li</surname></string-name> and <string-name>
<given-names>S.</given-names> 
<surname>Abdulla</surname></string-name>
</person-group>, &#x201C;
<article-title>Ensemble of adaboost cascades of 3L-LBPs classifiers for license plates detection with low quality images</article-title>,&#x201D; 
<source>Expert Systems with Applications</source>, vol. 
<volume>92</volume>, pp. 
<fpage>216</fpage>&#x2013;
<lpage>235</lpage>, 
<year>2018</year>.</mixed-citation>
</ref>
<ref id="ref-29">
<label>29</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>G. J.</given-names> 
<surname>Ansari</surname></string-name>, <string-name>
<given-names>J. H.</given-names> 
<surname>Shah</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Sharif</surname></string-name> and <string-name>
<given-names>S.</given-names> 
<surname>ur Rehman</surname> </string-name>
</person-group>, &#x201C;
<article-title>A novel approach for scene text extraction from synthesized hazy natural images</article-title>,&#x201D; 
<source>Pattern Analysis and Applications</source>, vol. 
<volume>23</volume>, no. 
<issue>3</issue>, pp. 
<fpage>1</fpage>&#x2013;
<lpage>18</lpage>, 
<year>2019</year>.</mixed-citation>
</ref>
<ref id="ref-30">
<label>30</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>J.</given-names> 
<surname>Kostkov&#x00E1;</surname></string-name>, <string-name>
<given-names>J.</given-names> 
<surname>Flusser</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>L&#x00E9;bl</surname></string-name> and <string-name>
<given-names>M.</given-names> 
<surname>Pedone</surname></string-name>
</person-group>, &#x201C;
<article-title>Image invariants to anisotropic Gaussian blur</article-title>,&#x201D; in <conf-name>Scandinavian Conf. on Image Analysis</conf-name>, 
<publisher-loc>Cham</publisher-loc>: 
<publisher-name>Springer</publisher-name>, pp. 
<fpage>140</fpage>&#x2013;
<lpage>151</lpage>, 
<year>2019</year>. </mixed-citation>
</ref>
<ref id="ref-31">
<label>31</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>D. P. P.</given-names> 
<surname>Mesquita</surname></string-name>, <string-name>
<given-names>J&#x00E3;o P. P.</given-names> 
<surname>Gomes</surname></string-name>, <string-name>
<given-names>F.</given-names> 
<surname>Corona</surname></string-name>, <string-name>
<given-names>A. H.</given-names> 
<surname>Souza</surname> 
<suffix>Junior</suffix></string-name>, <string-name>
<given-names>J. A. S.</given-names> 
<surname>Nobre</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Gaussian kernels for incomplete data</article-title>,&#x201D; 
<source>Applied Soft Computing</source>, vol. 
<volume>77</volume>, pp. 
<fpage>356</fpage>&#x2013;
<lpage>365</lpage>, 
<year>2019</year>.</mixed-citation>
</ref>
<ref id="ref-32">
<label>32</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>P. K.</given-names> 
<surname>Rana</surname></string-name> and <string-name>
<given-names>D.</given-names> 
<surname>Jhanwar</surname></string-name>
</person-group>, &#x201C;
<article-title>Image deblurring methodology using wiener filter &#x0026; genetic algorithm</article-title>,&#x201D; 
<source>International Journal of Advanced Engineering Research and Science</source>, vol. 
<volume>6</volume>, no. 
<issue>9</issue>, pp. 
<fpage>1</fpage>&#x2013;
<lpage>18</lpage>, 
<year>2019</year>.</mixed-citation>
</ref>
<ref id="ref-33">
<label>33</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>H.</given-names> 
<surname>Chen</surname></string-name>, <string-name>
<given-names>S. S.</given-names> 
<surname>Tsai</surname></string-name>, <string-name>
<given-names>G.</given-names> 
<surname>Schroth</surname></string-name>, <string-name>
<given-names>D. M.</given-names> 
<surname>Chen</surname></string-name>, <string-name>
<given-names>R.</given-names> 
<surname>Grzeszczuk</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;
<article-title>Robust text detection in natural images with edge-enhanced maximally stable extremal regions</article-title>,&#x201D; in <conf-name>2011 18th IEEE Int. Conf. on Image Processing</conf-name>, 
<publisher-loc>Brussels, Belgium</publisher-loc>, pp. 
<fpage>2609</fpage>&#x2013;
<lpage>2612</lpage>, 
<year>2011</year>. </mixed-citation>
</ref>
<ref id="ref-34">
<label>34</label><mixed-citation publication-type="other">
<person-group person-group-type="author"><string-name>
<given-names>Q.</given-names> 
<surname>Yi</surname></string-name>, <string-name>
<given-names>D.</given-names> 
<surname>Shen</surname></string-name>, <string-name>
<given-names>J.</given-names> 
<surname>Lin</surname></string-name> and <string-name>
<given-names>S.</given-names> 
<surname>Chien</surname></string-name>
</person-group>, &#x201C;
<article-title>The color specification of surrogate roadside objects for the performance evaluation of roadway departure mitigation systems</article-title>,&#x201D; 
<comment>SAE Technical Paper</comment>, 
<fpage>01</fpage>&#x2013;
<lpage>0506</lpage>, 
<year>2018</year>.</mixed-citation>
</ref>
<ref id="ref-35">
<label>35</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>T.</given-names> 
<surname>de Campos</surname></string-name>, <string-name>
<given-names>B. R.</given-names> 
<surname>Babu</surname></string-name> and <string-name>
<given-names>M.</given-names> 
<surname>Varma</surname></string-name>
</person-group>, &#x201C;
<article-title>Character recognition in natural images</article-title>,&#x201D; 
<source>VISAPP</source>, vol. 
<volume>2</volume>, no. 
<issue>7</issue>, pp. 
<fpage>23</fpage>&#x2013;
<lpage>38</lpage>, 
<year>2009</year>.</mixed-citation>
</ref>
<ref id="ref-36">
<label>36</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>S. M.</given-names> 
<surname>Lucas</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>Panaretos</surname></string-name>, <string-name>
<given-names>L.</given-names> 
<surname>Sosa</surname></string-name>, <string-name>
<given-names>A.</given-names> 
<surname>Tang</surname></string-name>, <string-name>
<given-names>S.</given-names> 
<surname> Wong</surname></string-name> <etal>et al.</etal>
</person-group><italic>,</italic> &#x201C;<chapter-title>ICDAR 2003 robust reading competitions&#x201D;</chapter-title>, in <conf-name>Seventh Int. Conf. on Document Analysis and Recognition</conf-name>, 
<publisher-loc>Brussels, Belgium</publisher-loc>, pp. 
<fpage>682</fpage>&#x2013;
<lpage>687</lpage>, 
<year>2003</year>.</mixed-citation>
</ref>
<ref id="ref-37">
<label>37</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>S. M.</given-names> 
<surname>Lucas</surname></string-name>
</person-group>, &#x201C;
<article-title>ICDAR, 2005 text locating competition results</article-title>,&#x201D; in <conf-name>Eighth Int. Conf. on Document Analysis and Recognition</conf-name>, 
<publisher-loc>Seoul, South Korea</publisher-loc>, pp. 
<fpage>80</fpage>&#x2013;
<lpage>84</lpage>, 
<year>2005</year>. </mixed-citation>
</ref>
<ref id="ref-38">
<label>38</label><mixed-citation publication-type="conf-proc">
<person-group person-group-type="author"><string-name>
<given-names>K.</given-names> 
<surname>Wang</surname></string-name>, <string-name>
<given-names>B.</given-names> 
<surname>Babenko</surname></string-name> and <string-name>
<given-names>S.</given-names> 
<surname>Belongie</surname></string-name>
</person-group>, &#x201C;
<article-title>End-to-end scene text recognition</article-title>,&#x201D; in <conf-name>2011 Int. Conf. on Computer Vision</conf-name>, 
<publisher-loc>Barcelona, Spain</publisher-loc>: 
<publisher-name>IEEE</publisher-name>, pp. 
<fpage>1457</fpage>&#x2013;
<lpage>1464</lpage>, 
<year>2011</year>. </mixed-citation>
</ref>
<ref id="ref-39">
<label>39</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>B.</given-names> 
<surname>Shi</surname></string-name>, <string-name>
<given-names>X.</given-names> 
<surname>Bai</surname></string-name> and <string-name>
<given-names>C.</given-names> 
<surname>Yao</surname></string-name>
</person-group>, &#x201C;
<article-title>An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition</article-title>,&#x201D; 
<source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. 
<volume>39</volume>, no. 
<issue>7</issue>, pp. 
<fpage>2298</fpage>&#x2013;
<lpage>2304</lpage>, 
<year>2016</year>.</mixed-citation>
</ref>
<ref id="ref-40">
<label>40</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>Y.</given-names> 
<surname>Wang</surname></string-name>, <string-name>
<given-names>C.</given-names> 
<surname>Shi</surname></string-name>, <string-name>
<given-names>B.</given-names> 
<surname>Xiao</surname></string-name>, <string-name>
<given-names>C.</given-names> 
<surname>Wang</surname></string-name> and <string-name>
<given-names>C.</given-names> 
<surname>Qi</surname></string-name>
</person-group>, &#x201C;
<article-title>CRF based text detection for natural scene images using convolutional neural network and context information</article-title>,&#x201D; 
<source>Neurocomputing</source>, vol. 
<volume>295</volume>, pp. 
<fpage>46</fpage>&#x2013;
<lpage>58</lpage>, 
<year>2018</year>.</mixed-citation>
</ref>
<ref id="ref-41">
<label>41</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>K.</given-names> 
<surname>Fan</surname></string-name> and <string-name>
<given-names>S. J.</given-names> 
<surname>Baek</surname></string-name>
</person-group>, &#x201C;
<article-title>A robust proposal generation method for text lines in natural scene images</article-title>,&#x201D; 
<source>Neurocomputing</source>, vol. 
<volume>304</volume>, pp. 
<fpage>47</fpage>&#x2013;
<lpage>63</lpage>, 
<year>2018</year>.</mixed-citation>
</ref>
<ref id="ref-42">
<label>42</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>S.</given-names> 
<surname>Huang</surname></string-name>, <string-name>
<given-names>D.</given-names> 
<surname>Wu</surname></string-name>, <string-name>
<given-names>Y.</given-names> 
<surname>Yang</surname></string-name> and <string-name>
<given-names>H.</given-names> 
<surname>Zhu</surname></string-name>
</person-group>, &#x201C;
<article-title>Image dehazing based on robust sparse representation</article-title>,&#x201D; 
<source>IEEE Access</source>, vol. 
<volume>6</volume>, pp. 
<fpage>53907</fpage>&#x2013;
<lpage>53917</lpage>, 
<year>2018</year>.</mixed-citation>
</ref>
<ref id="ref-43">
<label>43</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>S.</given-names> 
<surname>Salazar-Colores</surname></string-name>, <string-name>
<given-names>I.</given-names> 
<surname>Cruz-Aceves</surname></string-name> and <string-name>
<given-names>J. M.</given-names> 
<surname>Ramos-Arreguin</surname></string-name>
</person-group>, &#x201C;
<article-title>Single image dehazing using a multilayer perceptron</article-title>,&#x201D; 
<source>Journal of Electronic Imaging</source>, vol. 
<volume>27</volume>, no. 
<issue>4</issue>, 
<lpage>043022</lpage>, 
<year>2018</year>.</mixed-citation>
</ref>
<ref id="ref-44">
<label>44</label><mixed-citation publication-type="journal">
<person-group person-group-type="author"><string-name>
<given-names>R.</given-names> 
<surname>Minetto</surname></string-name>, <string-name>
<given-names>N.</given-names> 
<surname>Thome</surname></string-name>, <string-name>
<given-names>M.</given-names> 
<surname>Cord</surname></string-name>, <string-name>
<given-names>N. J.</given-names> 
<surname>Leite</surname></string-name> and <string-name>
<given-names>J.</given-names> 
<surname>Stolfi</surname></string-name>
</person-group>, &#x201C;
<article-title>SnooperText: A text detection system for automatic indexing of urban scenes</article-title>,&#x201D; 
<source>Computer Vision and Image Understanding</source>, vol. 
<volume>122</volume>, pp. 
<fpage>92</fpage>&#x2013;
<lpage>104</lpage>, 
<year>2014</year>.</mixed-citation>
</ref>
</ref-list>
</back>
</article>