<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">34072</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.034072</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An Efficient Long Short-Term Memory Model for Digital Cross-Language Summarization</article-title>
<alt-title alt-title-type="left-running-head">An Efficient Long Short-Term Memory Model for Digital Cross-Language Summarization</alt-title>
<alt-title alt-title-type="right-running-head">An Efficient Long Short-Term Memory Model for Digital Cross-Language Summarization</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Padmanabha Reddy</surname><given-names>Y. C. A.</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Kasireddy</surname><given-names>Shyam Sunder Reddy</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Sirisala</surname><given-names>Nageswara Rao</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Kuchipudi</surname><given-names>Ramu</given-names></name><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kollapudi</surname><given-names>Purnachand</given-names></name><xref ref-type="aff" rid="aff-5">5</xref><email>purnachandkollapudi2022@gmail.com</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of CSE, B V Raju Institute of Technology</institution>, <addr-line>Narsapur, Medak, T.S, 502 313</addr-line>, <country>India</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of IT, Vasavi College of Engineering</institution>, <addr-line>Hyderabad, T.S, 500089</addr-line>, <country>India</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of CSE, K.S.R.M College of Engineering</institution>, <addr-line>Kadapa, A.P, 516003</addr-line>, <country>India</country></aff>
<aff id="aff-4"><label>4</label><institution>Department of IT, C.B.I.T</institution>, <addr-line>Gandipet, Hyderabad, Telangana, 500075</addr-line>, <country>India</country></aff>
<aff id="aff-5"><label>5</label><institution>Department of CSE, B V Raju Institute of Technology</institution>, <addr-line>Narsapur, Medak, T.S, 502 313</addr-line>, <country>India</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Purnachand Kollapudi. Email: <email>purnachandkollapudi2022@gmail.com</email></corresp>
</author-notes>
<pub-date publication-format="print" date-type="pub" iso-8601-date="2022-12-15"><day>15</day>
<month>12</month>
<year>2022</year></pub-date>
<volume>74</volume>
<issue>3</issue>
<fpage>6389</fpage>
<lpage>6409</lpage>
<history>
<date date-type="received"><day>05</day><month>7</month><year>2022</year></date>
<date date-type="accepted"><day>15</day><month>9</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Padmanabha Reddy et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Padmanabha Reddy et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_34072.pdf"></self-uri>
<abstract>
<p>The rise of social networking enables the development of multilingual Internet-accessible digital documents in several languages. The digital document needs to be evaluated physically through the Cross-Language Text Summarization (CLTS) involved in the disparate and generation of the source documents. Cross-language document processing is involved in the generation of documents from disparate language sources toward targeted documents. The digital documents need to be processed with the contextual semantic data with the decoding scheme. This paper presented a multilingual cross-language processing of the documents with the abstractive and summarising of the documents. The proposed model is represented as the Hidden Markov Model LSTM Reinforcement Learning (HMM<sub>lstm</sub>RL). First, the developed model uses the Hidden Markov model for the computation of keywords in the cross-language words for the clustering. In the second stage, bi-directional long-short-term memory networks are used for key word extraction in the cross-language process. Finally, the proposed HMM<sub>lstm</sub>RL uses the voting concept in reinforcement learning for the identification and extraction of the keywords. The performance of the proposed HMM<sub>lstm</sub>RL is 2&#x0025; better than that of the conventional bi-direction LSTM model.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Text summarization</kwd>
<kwd>reinforcement learning</kwd>
<kwd>hidden markov model</kwd>
<kwd>cross-language</kwd>
<kwd>multilingual</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Natural Language Processing (NLP) is an effective platform on computers for the efficient functioning of certain tasks in human languages. It involves the processing of the input provided by the human language in the conversion of the input information into an appropriate representation of the information in another language [<xref ref-type="bibr" rid="ref-1">1</xref>]. The NLP input or output can be either speech or text in the form of information disclosure, language knowledge, semantic, lexical, syntactic, or another knowledge [<xref ref-type="bibr" rid="ref-2">2</xref>]. The NLP process comprises two tasks, such as understanding the natural languages and generating natural languages. At present, the World Wide Web (WWW) is considered a rich and effective source for information processing with a huge development rate of 29.7 billion pages with exponential growth [<xref ref-type="bibr" rid="ref-3">3</xref>]. As per the survey conducted by Netcraft, English is identified as a dominant language on the web [<xref ref-type="bibr" rid="ref-4">4</xref>]. However, the increase in the number of users reveals that non-English users also increased with the information repository, and it is challenging. Those problems with different language users can be overcome with Cross-Lingual Information Retrieval (CLIR) [<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p>India comprises of diverse languages around 2000 dialiets are identified with stipulated use of Hindi and English for official communication with the national government. India has two official languages used by the national government, as well as 22 scheduled languages for administrative purposes [<xref ref-type="bibr" rid="ref-6">6</xref>]. The identified languages are Bengali, Bodo, Dogri, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Oriya, Punjabi, Sanskrit, Santali, Sindhi, Tamil, Telugu, and Urdu. In addition, English is widely used in India in a variety of contexts, including the media, science and technology, commerce, and education [<xref ref-type="bibr" rid="ref-7">7</xref>]. However, in India, Hindi is considered one of the scheduled languages [<xref ref-type="bibr" rid="ref-8">8</xref>]. Countries like India demand a multi-lingual society involved in the translation of documents from one language to another. Most of the state government documents are based on the regional languages of the Union government and the documents are in bilingual form in either Hindi or English. In order to achieve appropriate communication, it is necessary to convert those documents into languages based on the regional languages [<xref ref-type="bibr" rid="ref-9">9</xref>]. Additionally, regional languages demand the conversion of the daily news as per English through international news agencies. With the human translators, the limitations of information missing in reports and documents are observed. With appropriate machine-based assistance translation or workstation translators, summation is effective for cross-lingual schemes [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>CLIR is involved in the processing of the languages based on the query languages with other languages for the searched documents. This involved estimation of the information availability in the native languages within CLIR and relevant information about documents being identified [<xref ref-type="bibr" rid="ref-11">11</xref>]. With CLIR, an appropriate relationship is evolved with the information requirements and content availability. The systems are effective for users who evaluate estimates in different languages to gather appropriate information from different language documents to eliminate multiple requests. Cross-language information retrieval [<xref ref-type="bibr" rid="ref-12">12</xref>] facilitates users to search documents in their familiar languages and uses the translation method to retrieve information from other languages, stated as Multilingual Cross Language Information Retrieval (MCLIR). The cross-lingual information access uses the machine learning based translation paradigm of summarization, subsequent translation, and snippets for information extraction from the targeted languages [<xref ref-type="bibr" rid="ref-13">13</xref>]. The CLIR technique uses different techniques, those being a) document translation, b) query translation, and c) inter-lingual translations. Within the CLIR system, to improve performance efficiency, it is necessary to implement knowledge-based robust algorithms for search across many languages and translate the documents.</p>
<p>To improve the functionality of the CLIR, data mining is an effective tool for information processing [<xref ref-type="bibr" rid="ref-14">14</xref>]. Data mining entails estimating non-trivial, implicit information patterns from repositories such as Extensible Markup Language (XML) repositories, data warehouses, relational databases, and so on. At present, large datasets are available for the collection and processing of information better than humans. The evaluation is based on the extraction of the knowledge into a more accurate and information value within the software environment. With larger datasets, traditional techniques are improved for retrieval of information from raw data that is infeasible [<xref ref-type="bibr" rid="ref-15">15</xref>]. The implementation of the data mining concept with machine learning improves the process of CLIR with data selection, data cleaning, transformation of data, knowledge presentation, and evaluation of patterns. With the core concept of artificial intelligence, machine learning, and statistics, data mining concepts yield benefits of automated information collection and processing. The machine learning technique comprises various methods and techniques with the implementation of different algorithms. In the cross-language information retrieval platform, different techniques are K-nearest neighbour, Nave Bayes, Decision Tree, Neural Network, and Cluster analysis [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p><bold>Contribution of the Work</bold></p>
<p>In search of the World Wide Web, English is considered a primary language, with an increase in the number of users. Non-English native speakers also search for documents. To facilitate the ease of search for people, it is necessary to construct an appropriate domain with machine learning for identification and translation of the cross-language abstraction and summation model. The proposed model HMM<sub>lstm</sub>RL comprises the LSTM model with HMM integrated with reinforcement learning. The specific contribution is presented as follows:
<list list-type="bullet">
<list-item><p>With the Hidden Markov model, word count and number of keywords count are considered for analysis and processing. The LSTM model calculates the total number of words in the statement or document. Through a calculated number of words, the number of repeated words is updated in the neural network model.</p></list-item>
<list-item><p>To develop a bi-directional LSTM-based corpus model for multi-language encoding processing with keyword extraction.</p></list-item>
<list-item><p>To construct a reinforcement learning based machine learning model for feature extraction and word identification. Upon the testing and validation of the information, words are processed and updated on the network.</p></list-item>
<list-item><p>Within reinforcement learning, the MapReduce framework is applied for the clustering and removal of the same words in the text document. Finally, voting is integrated for the abstraction and summarization of the keywords.</p></list-item>
<list-item><p>The experimental analysis showed that the proposed HMM<sub>lstm</sub>RL achieves higher precision and recall value compared with the conventional techniques.</p></list-item>
</list></p>
<p>This paper is organised as follows: Section 2 investigates how the cross-lingual processing model works. In Section 3, research methods are adopted for the encoding and decoding of data in multi-lingual systems with the proposed HMM<sub>lstm</sub>RL. In Section 4, experimental analysis of the proposed HMM<sub>lstm</sub>RL model is comparatively examined with existing techniques. Finally, in Section 5, the overall conclusion is presented.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Works</title>
<p>The key challenge for English to Hindi statistical machine translation is that the Hindi language is richer in morphology than the English language. There are two strategies that facilitate reasonable performance in this language pair. Firstly, reordering of English source sentences in accordance with the Hindi language is the second strategy, which is by making use of suffixes of Hindi words. Either of these strategies or both strategies can be used during translation. The difference in word order between the Indian language and English makes these two strategies equally challenging. For example, the English sentence &#x201C;he went to the office&#x201D; has the Hindi translation &#x201C;vah (he) kaaryaalay (office) gaya (went to).&#x201D; In this example, it is evident that the position of words in English is not retained in the translated Hindi text. An author has developed an unsupervised part-of-speech tagger that makes use of target language information, and it has been proven that the results are better as compared with the Baum-Welch algorithm [<xref ref-type="bibr" rid="ref-17">17</xref>]. The main demerit of this approach is the increase in the required number of translations in an exponential manner. Thus, increasing the overall time of execution due to the increase in time for translation. The pruning method was used by the author to identify the unlikely disambiguation. Thus, reducing the number of translations required and also the overall time complexity of the algorithm.</p>
<p>In [<xref ref-type="bibr" rid="ref-18">18</xref>] proposed interpretable Hidden Markov Models (HMM)-based approaches for emotion recognition in text and analysed their performance under different architectures (training methods), ordering, and ensembles (e.g.,). The presented models are interpretable; they may show the emotional portions of a phrase and explain the progression of the overall feeling from the beginning to the end of the sentence. Experiments show that the new HMM algorithms and training methods are just as good as machine learning methods and, in some cases, even do a better job than older HMMs.</p>
<p>In [<xref ref-type="bibr" rid="ref-19">19</xref>] developed a cross-language text summarization (CLTS) technique to generate word summaries of different languages from different documents. The compressive CLTS approach examined the source text and target language for the computation of relevant sentences. The developed system comprises two-level sentences such as clustering of similar sentences with multi-sentence compression (MSC). The second technique is the use of compressed neural network models. The developed approach comprises the compression of the multi-sentence generated with the French-to-English summarization information for the extractive state of the art system. Moreover, the developed cross-lingual summarization exhibits the desired grammatical quality with the extractive approaches.</p>
<p>In [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed a framework for text summarization based on group with sematic clustering with text summarization based on grouping related sentences with the Semantic Link Network (SLN) based on group ranking with the summary. With the implementation of the SLN model with the two-layer model with semantic links, with part of a link, a sequential link, similar to a link, and a link for cause and effect, the experimental analysis expressed that the composed summaries involved in groups or paragraphs with composed sentences and summaries had an average source text length of 7000 words to 17,000 words usual length. Furthermore, seven clustering algorithms are generated and grouped with the five strategies through semantic links.</p>
<p>In [<xref ref-type="bibr" rid="ref-21">21</xref>] aimed to increase the Arabic summarisation accuracy through the novel topic-aware abstractive in the Arabic (TAAM) summarisation model with the integrated recurrent neural network. The quantitative and qualitative analysis of the TAAM model uses the ROUGE matrices, which provides a higher accuracy value of 10.8&#x0025; than the baseline models. Additionally, with the TAAM model, it generates coherent Arabic summaries that are captured and read for the input idea.</p>
<p>In [<xref ref-type="bibr" rid="ref-22">22</xref>] constructed a problem-based learning approach for natural language processing (NLP) processing in undergraduate and graduate students. The team of students uses sets for big data, guidelines, resources for cloud computing and other aids for the different teams with the collection of two big data those are related to web pages and electronic theses and dissertations (ETDs). The team of students is deployed based on consideration of the different methods, tools, libraries, and tasks for summarization. As the summarization is a process involved in addressing the NLP learning process for the different linguistic levels based on consideration of NLP practitioners&#x2019; tools and techniques. The results evaluation revealed that coherent teams are generated based on readable summarization. With different summarization techniques, the ETD process is evaluated through high quality and accurate utilisation of the NLP pipelines. Additionally, the developed technique expresses that the data and information are managed effectively in those fields for the developed approach or similar methods for the defined data sets through the provision of synergistic solutions.</p>
<p>In [<xref ref-type="bibr" rid="ref-23">23</xref>] presented a narrative abstractive summarization (NATSUM) technique for the chronological ordering of the target entity based on the consideration of the chronological order within the same topic. To achieve an effective cross-summarization process, a timeline for the cross-document is mentioned for the different point events for the same event. With the timeline-based arguments, the documents&#x2019; events are extracted and processed. Secondly, with the natural language generation, sentences are produced with event arguments. Specifically, through the realisation technique, hybrid surfaces are derived based on the technique for the generation and ranking of hybrid surfaces. The experimental analysis showed that the proposed NATSUM exhibits effective summarization through abstractive baselines to increase the F1-measure by 50&#x0025; for a simulated environment.</p>
<p>In [<xref ref-type="bibr" rid="ref-24">24</xref>] proposed a framework for the syntactic features using novel Syntax-Enriched Abstractive Summarization (SEASum) with graph attention networks (GATs) with the implementation of the source texts. With the implemented SEASum architecture, the sequence of the Pre-trained Language Models (PTM) semantic encoder is implemented with the graph attention networks (GAT)-based explicit syntax for the source document presentation. With the summarization framework, a feature fusion module is introduced through the syntactic features for the summarization models using the SEASum model for the semantic and syntactic encoder for multi-head fusion of the feature stream in the decoding process. Secondly, through the SEASum cascaded model, contextual embedding is performed for the semantic words for information flow. The experimental results show that the built SEASum and cascaded models outperform the abstractive summarization approach.</p>
</sec>
<sec id="s3"><label>3</label><title>Hidden Markov Model for the Extraction of Features</title>
<p>The proposed HMM<sub>lstm</sub>RL comprises of the coding scheme applied with the Hidden Markov Model for abstraction and summation. Initially, the character is estimated with HMM<sub>lstm</sub>RL perform coding scheme for the estimation of the stored characters in the computer in the form of bits. The detection of the charset is based on estimation of the Unicode transformation (UTF &#x2013; 8) due to presence of larger bytes in the sequences. The bytes are computed based on the implementation of the bytes with validity test through <italic>extremely</italic> unlikely scenario. With heuristics detection scheme it is necessary to assign label properly in the dataset label with appropriate encoding mechanism. Through hypertext markup language (HTML) documents encoding is in the meta elements as &#x003C;Metahttp-equiv=&#x201C;Content-Type&#x201D;content=&#x201C;text/html; charset=UTF-8&#x201D;/&#x003E;. In sequence, the document is processed with the Hypertext Transfer Protocol (HTTP) with the similar meta data for the content type header. Finally, the encoding is applied over the text filed for the initial bytes label explicit. The Unicode model comprises of the different unique number for each character those are stated as follows:</p>
<p>Accommodates more than 65,000.</p>
<p>Synchronized with the corresponding versions of ISO-10646.</p>
<p>Standards incorporated under Unicode</p>
<p>ISO 6937, ISO 8859 series</p>
<p>ISCII, KS C 5601, JIS X 0209, JIS X 0212, GB 2312, and CNS 11643 etc.</p>
<p>Scripts and Characters, European alphabetic scripts</p>
<p><italic>Middle Eastern right-to-left scripts, Scripts of Asia</italic></p>
<p><italic>Indian languages Devanagari, Bengali, Gurmukhi, Oriya, Tamil, Telugu, Kannada, Malayalam</italic>.</p>
<p><italic>Punctuation marks, diacritics, mathematical symbols, technical symbols, arrows, dingbats, etc</italic>.</p>
<p>The text processing system comprises of the dependent components based on the dictionaries in the targeted language document with the assigned character codes. The each elements code is defined as Unique number code point with the hexadecimal prefix number of U&#x201C;Ex&#x201D;, U&#x002B;0041 with value of &#x201C;A&#x201D;. The language character level is identified and processed with the statistical language identification for the training data in reinforcement learning [<xref ref-type="bibr" rid="ref-25">25</xref>&#x2013;<xref ref-type="bibr" rid="ref-31">31</xref>]. The performance is evaluated based on the consideration of the characteristic&#x0027;s samples defined as in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Characteristics of languages</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_34072-inline-1.png"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s3_1"><label>3.1</label><title>Hidden Markov Model for Position Estimation</title>
<p>Hidden Markov model has fewer assumptions of independence. In particular, HMM does not assume that probability that condemnation i is in summary is independent of whether condemnation i&#x2212;1 is in summary. The syntactic and semantic features are extracted from the source language text and are used in the transfer phase to generate the sentence in the target language. This information is extracted by the HMM source language analysis phase. This phase is further sub-divided into&#x2013;part-of-speech tagging and word sense disambiguation. Once this information is extracted, it is used in the transfer phase which makes use of Bayesian approach. This sense-based machine translation system makes use of Bayesian approach, which is based on the statistical analysis of existing bilingual parallel corpora. The label assigned with the HMM model is defined as follows based on features.</p>
<table-wrap id="table-14">
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Feature Name</th>
<th align="left">Label</th>
<th align="left">Value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Position of Paragraph</td>
<td align="left">O1</td>
<td align="left">1,2,3</td>
</tr>
<tr>
<td align="left">Number of terms</td>
<td align="left">O2</td>
<td align="left">log(wi&#x0002B;1)</td>
</tr>
<tr>
<td align="left">Baseline Term Probability</td>
<td align="left">O3</td>
<td align="left">log(Pr(terms in i|baseline))</td>
</tr>
<tr>
<td align="left">Document Term Probability</td>
<td align="left">O4</td>
<td align="left">log(Pr(terms in i|document))</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The position of condemnation in its paragraph. We assign each condemnation value o1(i) designating it as rest in paragraph (value 1), last in paragraph (value 3), or intermediate condemnation (value 2) condemnation in one-long paragraph is assigned value 1, &#x0026; condemnation s in two-long paragraph are assigned values of 1 &#x0026; 3. A number of terms in condemnation. value of this feature is computed as o2(i) &#x003D; log (number of terms &#x002B; 1).</p>
<p>The content processing with the HMM<sub>lstm</sub>RL with application of the HMM is defined as
<list list-type="order">
<list-item><p>The position of condemnation is incorporated with state-structure of HMM.</p></list-item>
<list-item><p>Estimation of this component is o1(i) &#x003D; log (number of terms &#x002B; 1)</p></list-item>
<list-item><p>Likely condemnation terms are, given report terms o2(i) &#x003D; log (P r (terms in condemnation i|D)).</p></list-item>
</list></p>
<p>The statistical analysis is also performed based on the local syntactic information available in the input sentence. Using the analysed data, the Bayesian approach is applied to predict the probable target word for the given input word. This proposed word sense-based statistical machine translation system may be mathematically expressed as in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mi>T</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mi>T</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>In the above <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>, source word is denoted as <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>; part of the source word is defined as <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>T</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and target word is denoted as <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.The probable target word <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>has dependency on the source word <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>and its part-of-speech <italic>tags</italic>. Thus, to find the probable target word <italic>Wt</italic> with the prior knowledge about the input source word and its predicted part-of-speech, a Bayesian approach is used. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> shows the pre-processing phase in proposed method.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>Pre-Processing phase in HMM<sub>lstm</sub>RL</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_34072-fig-1.png"/></fig>
<p>Since the proposed machine translation system works at word level, there is need for tokenization of source text at sentence level as well as word level. In Hindi language, there is tense, aspect and modality (TAM) information stored in the affixes of the words. These affixes also contribute to the accuracy of machine translation. To extract the TAM information stored in the affixes, longest affix matching algorithm is used to check the matching between affixes. Levenshtein distance is used to calculate the matching score. For example, consider the word <inline-graphic xlink:href="CMC_34072-inline-2.png"/> (khaane) and <inline-graphic xlink:href="CMC_34072-inline-3.png"/>(rahega) mentioned in input sentence &#x2013; <inline-graphic xlink:href="CMC_34072-inline-4.png"/> (raat ko ek rotee kam khaane se pet halka rahega)&#x201D;. Both these words have TAM information in it and it can be extracted using the longest affix matching algorithm. The word <inline-graphic xlink:href="CMC_34072-inline-5.png"/> (khaane) will be analyzed and the TAM information is found as masculine, plural verb. Similarly, the TAM information of the word <inline-graphic xlink:href="CMC_34072-inline-6.png"/> (rahega) is found to be masculine, singular verb. Once these affixes are extracted, the sequence of words along with its affixes are fed to the next phase of the translation system as shown in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2"><label>Table 2:</label><caption><title>Estimation of affix in statement</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_34072-inline-7.png"/></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2"><label>3.2</label><title>Word Count Extraction with Bi-Directional LSTM in HMM<sub>lstm</sub>RL</title>
<p>In the LSTM model, the probability value of the document drivers is calculated from the existing languages. The probability is computed based on the occurrence of the string value as S with the alphabet X sequence is represented in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. The defined probability for the string S in the character sequence for the words in document is represented in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:mstyle></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x2013; words total number occurred in the word m from the collection of N documents</p>
<p><inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x2013; in the document i number of words m appear</p>
<p>With the predefined categories for n values differentiate keywords are assigned with category <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for every other category value of <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In the LSTM model the combination of the probabilities in the language identified is estimated based on each letter. The selected language in the document highest probability is computed using the <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>n</mml:mi><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>r</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>The distribution of parameters is based on the consideration of the different variable n with the trails number and probability of occurrence of character in the document Unigram as p. The multilingual corpus texts for the characters are understand with the character set defined as follows:</p>
<p><inline-graphic xlink:href="CMC_34072-inline-8.png"/> so, on are charset of Kannada, A, B, C, D, so on are charset of Latin English <inline-graphic xlink:href="CMC_34072-inline-9.png"/> so on are charset of Telugu.</p>
</sec>
<sec id="s3_3"><label>3.3</label><title>Text Extraction and Clustering with MapReduce</title>
<p>The hidden markov model-based MapReduce tagger is a statistical approach which is used to identify the probable part-of-speech of each word in the sentence. The hidden markov model-based tagger basically finds the most probable sequence of part-of-speech for a given sentence by using the transition probability <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and emission probability <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. These transition and emission probabilities are learned by the tagger using the monolingual Hindi corpus. These probabilities are calculated using the expressions as in <xref ref-type="disp-formula" rid="eqn-5">Eqs. (5)</xref> and <xref ref-type="disp-formula" rid="eqn-6">(6)</xref>,
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula></p>
<p>The Input word sequences are denoted as W1, W2, W3, &#x2026;, n and the part-of-speech of each word is denoted as T1, T2, T3, &#x2026;, Tn. The part-of-speech of each word T1, T2, T3, &#x2026;, Tn acts as hidden states. Each of these hidden states is predicted using the emission and transition probabilities. For example, consider the input sentence as <inline-graphic xlink:href="CMC_34072-inline-10.png"/> (raat ko ek rotee kam khaane se pet halka rahega.)&#x201D; and the POS tagset as [JJ (adjective), N_NN (Noun), PSP (Postposition), QT_QTC (Cardinal), QT_QTF (Quantifiers), V_VM (Verb), RD_PUNC (Punctuations), CC_CCS (Conjuncts)]. The beginning of a sentence is denoted by ^ symbol. Considering the first word &#x201C;<inline-graphic xlink:href="CMC_34072-inline-11.png"/> (raat)&#x201D; in the sentence to calculate the probability of the word as a noun,</p>
<p><inline-graphic xlink:href="CMC_34072-inline-12.png"/></p>
<p><disp-formula id="eqn-8"><label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mo>&#x2227;</mml:mo></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2227;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x2227;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula></p>
<p>The frequency of occurrence count is found from the monolingual Hindi corpus. Consider there are 35 occurrence of word <inline-graphic xlink:href="CMC_34072-inline-50.png"/> (raat) as noun (N_NN) and the occurrence of noun (N_NN) to be 1000 out of which 600 times it is tagged with first word of the sentence. Number of sentences is also considered as 1000.</p>
<fig id="eqn-9">
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_34072-inline-52.png"/>
</fig>
<p><disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mo>&#x2227;</mml:mo></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>600</mml:mn><mml:mn>1000</mml:mn></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0.6</mml:mn></mml:math></disp-formula></p>
<p>In similar manner as in <xref ref-type="fig" rid="eqn-9">Eqs. (9)</xref> and <xref ref-type="disp-formula" rid="eqn-10">(10)</xref> each word probability is computed in the LSTM network. The computed words in the documents are aligned and processed as stated in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>Word alignment in HMM<sub>lstm</sub>RL</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_34072-inline-20.png"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Hindi Text: <inline-graphic xlink:href="CMC_34072-inline-16.png"/> (paanee peena svaasthy ke lie achchha hai.)</p>
<p>English Equivalent: Drinking water is good for health.</p>
<p>Tamil Text-1: <inline-graphic xlink:href="CMC_34072-inline-17.png"/></p>
<p>Tamil Text-2: <inline-graphic xlink:href="CMC_34072-inline-18.png"/></p>
<p>Tamil Text-3: <inline-graphic xlink:href="CMC_34072-inline-19.png"/>.</p>
<p>To disambiguate the appropriate sense of the word mentioned in source text, a word sense disambiguation is used.</p>
<p>The identified senses are used in the proposed statistical machine translation approach. The transfer phase basically predicts the probable target language word based on the source language word and its part-of-speech. Using Bayes rule, the transfer phase is mathematically expressed as below in <xref ref-type="disp-formula" rid="eqn-11">Eqs. (11)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-14">(14)</xref>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>The target word (<italic>Wt</italic>) is conditionally dependent on its previous word (<italic>wpre</italic>) in the text. Similarly, every target word has dependency with its preceding word. So, the probability of target word [60] is defined as a n-gram model and is mathematically expressed as in <xref ref-type="disp-formula" rid="eqn-15">Eqs. (15)</xref> and <xref ref-type="disp-formula" rid="eqn-16">&#x00A0;(16)</xref>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">count</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Wprev</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">count</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>v</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
</sec>
<sec id="s3_4"><label>3.4</label><title>Reinforcement Learning with MapReduce for Word Summation</title>
<p>In this proposed HMM<sub>lstm</sub>RL model, the part-of speech (pos) of the source text is also considered as a parameter to predict the probable alignment for a word. Thus, the probable position of target word can be calculated using its conditional dependence with position of source word, length of source text, length of target text and part-of speech of source word. Using Bayes theorem, the modified word alignment model is mathematically represented in below <xref ref-type="disp-formula" rid="eqn-17">Eq. (17)</xref>
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mi>j</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>In the above <xref ref-type="disp-formula" rid="eqn-17">Eq. (17)</xref>, target word position is represented as j, source word is denoted as i, source text of the word is represented as l, part of speech is denoted as pos and target text as m. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows the overall flow of proposed method.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Overall flow of the HMM<sub>lstm</sub>RL</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_34072-fig-2.png"/></fig>
<p>The hidden layer in the MapReduce tagger is activated using a sigmoid function which is expressed mathematically as below,
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>w</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><italic>word</italic> &#x2013; current word being processed</p>
<p><italic>tagi</italic> &#x2013; Part-of-speech tag of previous word</p>
<p><italic>tagj</italic> &#x2013; Part-of-speech tag from the considered POS tag set</p>
<p><italic>N</italic> &#x2013; number of distinct part-of-speech tags</p>
<p>The abstraction and summation phase of the proposed HMM<sub>lstm</sub>RL is estimated with reinforcement learning defined as in <xref ref-type="disp-formula" rid="eqn-20">Eq. (20)</xref>
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Using extended Bayes theorem, the above expression (20) is rewritten as,
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The HMM<sub>lstm</sub>RL approach for Hindi to Tamil machine translation is compared with the na&#x00EF;ve Bayes statistical machine translation system in terms of the features that are being used in both the system. The term-document frequency matrix is constructed using all the above sentences [S1, S2, &#x2026;, S6] along with the input sentence. Using these 7 sentences, the number of distinct words is identified as 39. Thus, the term-document frequency matrix (A) will be of size (39x7) and is as shown below,
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mtable columnalign="left center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mspace width="thinmathspace" /><mml:mtable columnalign="left center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtable columnalign="left center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mspace width="thinmathspace" /><mml:mtable columnalign="left center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The cosine similarity between vectors in right singular matrix and first row of singular diagonal matrix is calculated. The resultant vector after applying cosine similarity is as below,</p>
<p><italic>cosine</italic> <italic>similarity</italic> &#x003D; [0.13 0.3341 0.5423 0.4866 0.5335 0.4971 0.5086]</p>
<p>The sentence which has the cosine similarity nearer to 0 is closest to the input sentence. The first value in the vector denotes the cosine similarity with the input sentence itself and it is natural that it will be closest to zero. The next smallest value in the vector is 0.3341 which is the cosine similarity value of sentence S1.</p>
</sec>
<sec id="s3_5"><label>3.5</label><title>Word Extraction in the Sentence</title>
<p>This paper concentrated on the evaluation of language keywords through the implementation of a stacked classifier integrated with a voting scheme. This research utilizes machine learning for language estimation and classification. This research considers four classifiers such as AdaBoost, Artificial neural networks (ANN), Decision tree, and Support Vector Machine (SVM) integrated with the voting scheme. The developed HMM<sub>lstm</sub>RL is adopted through machine learning, which involves several steps those are input data, data pre-processing, feature extraction, feature selection, and classification. Based on the proposed HMM<sub>lstm</sub>RL classification of language classification. The steps implemented in the proposed mechanism are presented as follows:</p>
<p><italic>Input Data:</italic> The dataset collected from different data sources those are incorporated in machine learning for data processing.</p>
<p><italic>Data Pre-processing:</italic> In this stage collected dataset is evaluated for the elimination of redundant data.</p>
<p><italic>Feature Extraction:</italic> In this stage, language features are extracted for further processing of machine learning.</p>
<p><italic>Feature Selection:</italic> It eliminates the irrelevant features from the set of extracted data. To eliminate the irrelevant features for minimization of computational complexity with improved accuracy.</p>
<p><italic>Classification:</italic> This stage provides the results of the present data. This is adopted two factors such as training and testing:</p>
<p><italic>Training:</italic> This stage of machine learning involved training machine learning algorithms for making computers learn from the extracted features for the languages. Based on the classifier parameters are trained to machine learning to fit the classifier.</p>
<p><italic>Testing:</italic> The dataset for classification is evaluated based on the evaluation of specified classifier.</p>
<p>In the proposed HMM<sub>lstm</sub>RL mechanism collected data features are processed for machine learning stages. The proposed HMM<sub>lstm</sub>RL scheme uses 4 classifiers with stacking in machine learning. The 4 classifiers considered are AdaBoost, ANN, decision tree, and SVM in the machine learning process in HMM<sub>lstm</sub>RL. The proposed HMM<sub>lstm</sub>RL is involved in the classification of attacks. Initially, language classification is performed with consideration of 4 classifiers such as AdaBoost, SVM, decision, and ANN. With classification, if it is identified as an attack for 2 classifiers and not attack as another classifier then the proposed HMM<sub>lstm</sub>RL evaluates with a decision tree. The decision tree algorithm is involved in the decision-making of the language whether the file is an attack or not. Based on the classification results provided by the decision tree network system will estimate which language it belongs. The classification of the language is based on the voting mechanism.</p>
<p>Initially, the proposed HMM<sub>lstm</sub>RL train the model through classifiers such as AdaBoost, ANN, SVM, and decision tree. The model trained through the classifier is exported to the predictor for the computation of language. The predictors evaluate the predictor vector of each classifier and update to voting. The classification voting approach is based on the score value of the classifier; if the value obtained from the four classifiers is computed as a language, then the network is considered another language keywords. In case if two classifiers are stated as a keyword for the language then the proposed HMM<sub>lstm</sub>RL goes with the decision tree process for computation. To make a decision proposed HMM<sub>lstm</sub>RL utilizes a voting mechanism. The voting scheme is utilized for the estimation of languages. The analysis is based on the summation of classifier value for computation of languages belongs to keyword or not. The voting scheme is estimated based on the score of the classifier through consideration of condition such as:</p>
<p>if score &#x003E;&#x2009;2; then keyword language</p>
<p>else 0; other language</p>
<p>Based on the computation of classifier value voting score is computed. Through computed voting score proposed HMM<sub>lstm</sub>RL compute the data is belongs to particular language or not. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the voting process in HMM<sub>lstm</sub>RL.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Voting process in HMM<sub>lstm</sub>RL</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_34072-fig-3.png"/></fig>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Results and Discussions</title>
<p>The proposed HMM<sub>lstm</sub>RL comprises the sequence-to-sequence model and was developed using pytorch with TensorFlow at the backend. TensorFlow is an open-source platform used for developing deep neural networks. The proposed model learns the features from the input vector and target vector. These features are used to generate the target text based on the input vector fed to it. To make the model learn the features in an efficient way, there is a need for a huge amount of corpus. Due to this feature, a sentence in Tamil and Hindi can be shuffled in different combinations to generate variants of the given sentence. Since the Hindi language is a partially free word order language, all the combinations generated will not be grammatically correct. Thus, there is a need for verifying the grammatical correctness of the text being generated. Parsing the sentences will be helpful in checking the grammatical correctness of them. In this proposed approach, HMM<sub>lstm</sub>RL is used to verify the correctness of the generated Hindi sentence. The Hindi parser verifies the grammar by parsing the tagged text fed to it. In this way, the valid variants of Hindi text are generated along with its Tamil sentences and are maintained in the training dataset. Similarly, the Tamil sentences can also be shuffled, but there is no necessity for verification of grammar in them. This is due to the fully free word order nature of the language.</p>
<p>The proposed HMM<sub>lstm</sub>RL model was analysed with various dropout percentages and an optimal percentage value was found to be in the range of 20&#x0025; to 60&#x0025;. Since there is an encoder module and a decoder module, there is a need to analyse the dropout percentage in both these modules such that the performance of the overall system is good. The ideal dropout percentage for encoders is found to be 20&#x0025;, and the ideal dropout percentage for decoders is 60&#x0025;. The following are parameters that were used for the sequence-to-sequence model,</p>
<p>Number of epochs&#x2009;&#x003D;&#x2009;22</p>
<p>Learning rate&#x2009;&#x003D;&#x2009;0.01</p>
<p>Hidden layer size&#x2009;&#x003D;&#x2009;2</p>
<p>Dropout&#x2009;&#x003D;&#x2009;0.2 (in encoder) and 0.6 (in decoder)</p>
<p>After 22 epochs, the feature learning by model gets saturated and <inline-graphic xlink:href="CMC_34072-inline-21.png"/>the necessity for training becomes negligible. The learning rate is kept at 0.01 and the accuracy of the model fluctuates when the learning rate is made a 0.1. This is due to rapid change on the weights and thus the model was not able to learn the features. The number of hidden layers used in the proposed approach is two. The performance of proposed model is comparatively better as compared with the model having more than two hidden layers.</p>
<p>For the analysis taking an example from our published word [61], consider the input text as: <inline-graphic xlink:href="CMC_34072-inline-22.png"/> (English Equivalent: Fresh breath and shining teeth enhance your personality. After performing the initial pre-processing phase, the words in the input sentence are identified as unique token and these identified tokens are further used to extract the syntactic information in the sentence. The syntactic information is extracted using HMM based part-of-speech tagger. The sample part-of-speech tags used in this input sentence are described in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>Computation of tags in sentence</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Tag</th>
<th align="left">JJ</th>
<th align="left">N_NN</th>
<th align="left">CC_CCD</th>
<th align="left">PR_PRP</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Part-Of-Speech</td>
<td align="left">Adjective</td>
<td align="left">Noun</td>
<td align="left">Conjunction</td>
<td align="left">Pronoun</td>
</tr>
<tr>
<td align="left">Tag</td>
<td align="left">PSP</td>
<td align="left">V_VM></td>
<td align="left">V_VAUX</td>
<td align="left">RD_PUNC</td>
</tr>
<tr>
<td align="left">Part-Of-Speech</td>
<td align="left">Postposition</td>
<td align="left">Finite Verb</td>
<td align="left">Auxiliary Verb</td>
<td align="left">Punctuation</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The tagged output for the text considered will be as below, <inline-graphic xlink:href="CMC_34072-inline-23.png"/>\JJ <inline-graphic xlink:href="CMC_34072-inline-24.png"/>\N_NN <inline-graphic xlink:href="CMC_34072-inline-25.png"/>\CC_CCD <inline-graphic xlink:href="CMC_34072-inline-26.png"/>\JJ <inline-graphic xlink:href="CMC_34072-inline-27.png"/>\N_NN <inline-graphic xlink:href="CMC_34072-inline-28.png"/>\PR_PRP <inline-graphic xlink:href="CMC_34072-inline-29.png"/>\N_NN <inline-graphic xlink:href="CMC_34072-inline-30.png"/>\P SP <inline-graphic xlink:href="CMC_34072-inline-31.png"/>\V_VM <inline-graphic xlink:href="CMC_34072-inline-32.png"/>\V_VAUX <inline-graphic xlink:href="CMC_34072-inline-51.png"/>. \RD_PUNC. (taaja\JJ saansen\N_NN aur\CC_CCD chamachamaate\JJ daant\N_NN aapake \PR_PRP vyaktitv\N_NN ko\PSP nikhaarate\V_VM hain\V_VAUX <inline-graphic xlink:href="CMC_34072-inline-51.png"/>.\RD_PUNC.) This tagged text is fed to the word sense disambiguation phase to identify the various possible senses for the word used in that context. Considering the word <inline-graphic xlink:href="CMC_34072-inline-33.png"/>\JJ (taaja\JJ), the various senses for this word is retrieved from the Hindi wordnet. The various senses retrieved are &#x2013; <inline-graphic xlink:href="CMC_34072-inline-34.png"/> (taaza), <inline-graphic xlink:href="CMC_34072-inline-35.png"/> (taaja), <inline-graphic xlink:href="CMC_34072-inline-36.png"/> (amalaan), <inline-graphic xlink:href="CMC_34072-inline-37.png"/> (ashashuk), <inline-graphic xlink:href="CMC_34072-inline-38.png"/> (aala), <inline-graphic xlink:href="CMC_34072-inline-39.png"/> (garamaagaram), <inline-graphic xlink:href="CMC_34072-inline-40.png"/> (tataka). For all the retrieved senses, its corresponding usage sentences are retrieved which are used in the word sense disambiguation phase to identify the senses that match with the input text. The identified possible senses for the word <inline-graphic xlink:href="CMC_34072-inline-41.png"/> are <inline-graphic xlink:href="CMC_34072-inline-42.png"/> and <inline-graphic xlink:href="CMC_34072-inline-43.png"/> as shown in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5"><label>Table 5</label><caption><title>Estimated sentences with HMM<sub>lstm</sub>RL</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_34072-inline-44.png"/></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The proposed HMM<sub>lstm</sub>RL compute sequence to sequence model was tested with various test set. The generated target sentences were evaluated using Bilingual evaluation understudy (BLEU) score. The <xref ref-type="table" rid="table-6">Table 6</xref> shows the analysis of BLEU score with different training and testing pairs. The BLEU score calculated for this proposed system is the average of sentence level BLEU score. The BLEU score is found to improve with increase in the percentage of training set being used. The accuracy of proposed model is good when the training and testing corpus ratio is kept at 80:20. The estimation of machine learning parameters is shown in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6"><label>Table 6</label><caption><title>Estimation of machine learning parameters</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Training/Testing Corpus Size (in &#x0025;)</th>
<th align="left">BELU Score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">60/40</td>
<td align="left">0.7037</td>
</tr>
<tr>
<td align="left">70/30</td>
<td align="left">0.7234</td>
</tr>
<tr>
<td align="left">80/20</td>
<td align="left">0.7588</td>
</tr>
<tr>
<td align="left">90/10</td>
<td align="left">0.6628</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The neural machine translation system is also evaluated at different runs by keeping the ratio of training and testing pair as 80:20 as shown in <xref ref-type="table" rid="table-7">Table 7</xref>. At each of the runs the training and testing set has been changed and are chosen at random. It is also found that the neural machine translation system has accuracy better than the statistical machine translation system.</p>
<table-wrap id="table-7"><label>Table 7</label><caption><title>Computation of BELU values</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Number of Runs</th>
<th align="left">BELU score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">0.7478</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">0.7176</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">0.6784</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">0.7118</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">0.7588</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">0.7124</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">0.7022</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">0.6914</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">0.7156</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">0.7211</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The word embedding is performed using a continuous bag-of-words model and it is found to capture the semantics in the words. This in turn helped in improving the accuracy of the translation using the sequence-to-sequence model. Since Hindi and Tamil language are morphologically rich, there is need for semantic mapping which is made using this approach. The results are found to be far better than any state-of-art method for these two languages. It is found that BLEU score is 0.7588 and it can be improved further by using a properly aligned parallel corpora as shown in <xref ref-type="table" rid="table-8">Table 8</xref>. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the proposed HMM<sub>lstm</sub>RL recall and precision.</p>
<table-wrap id="table-8"><label>Table 8</label><caption><title>Performance analysis</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Corpus Size (In number of words)</th>
<th align="center" colspan="2">HMM<sub>lstm</sub>RL</th>
<th align="center" colspan="2">Without HMM<sub>lstm</sub>RL</th>
</tr>
<tr>
<th align="left">Precision (in &#x0025;)</th>
<th align="left">Recall (in &#x0025;)</th>
<th align="left">Precision (in &#x0025;)</th>
<th align="left">Recall (in &#x0025;)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">10000</td>
<td align="left">74</td>
<td align="left">77</td>
<td align="left">56</td>
<td align="left">51</td>
</tr>
<tr>
<td align="left">20000</td>
<td align="left">83</td>
<td align="left">86</td>
<td align="left">70</td>
<td align="left">73</td>
</tr>
<tr>
<td align="left">30000</td>
<td align="left">91</td>
<td align="left">89</td>
<td align="left">79</td>
<td align="left">82</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-4"><label>Figure 4</label><caption><title>HMM<sub>lstm</sub>RL recall and precision</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_34072-fig-4.png"/></fig>
<p>The target word is thus predicted to be &#x201C;<inline-graphic xlink:href="CMC_34072-inline-45.png"/>F (podhu)&#x201D; for the English word &#x201C;general&#x201D;. Similarly, all other words are translated to its probable Tamil word. In this case, the grammar of the input sentence matches with the target text. Thus, there is no need for structural rearrangements. This is identified by using the na&#x00EF;ve Bayes approach described before.</p>
<p>The generated sentence is compared with the reference sentence and is as shown in <xref ref-type="table" rid="table-9">Table 9</xref>. The proposed system has been analyzed on various other source sentence and its evaluation metrics are calculated. The results of evaluation metrics are as shown in <xref ref-type="table" rid="table-10">Table 10</xref>. Which details about the precision and recall of the proposed pivot language-based system. The table also shows the comparison of proposed system with the statistical machine translation without pivot language.</p>
<table-wrap id="table-9"><label>Table 9</label><caption><title>Comparison of generated target text with its reference text</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_34072-inline-46.png"/></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-10"><label>Table 10</label><caption><title>Comparison of proposed system with the statistical machine translation without pivot language</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">S.No.</th>
<th align="left" rowspan="2">Corpus Size<break/>(In number of words)</th>
<th align="center" colspan="2">Bi-directional LSTM</th>
<th align="center" colspan="2">Proposed HMM<sub>lstm</sub>RL</th>
</tr>
<tr>
<th/>
<th align="left">Precision (in &#x0025;)</th>
<th align="left">Recall (in &#x0025;)</th>
<th align="left">Precision (in &#x0025;)</th>
<th align="left">Recall(in&#x0025;)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">10000</td>
<td align="left">53</td>
<td align="left">52</td>
<td align="left">58</td>
<td align="left">57</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">20000</td>
<td align="left">64</td>
<td align="left">67</td>
<td align="left">71</td>
<td align="left">72</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">30000</td>
<td align="left">76</td>
<td align="left">74.5</td>
<td align="left">83</td>
<td align="left">79</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The performance of the bi-directional LSTM seems to be poor as compared with the statistical machine translation without pivot language which is evident from the graph shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The BLEU score of this pivot-based system is found to be 0.54.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Comparison of Precision</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_34072-fig-5.png"/></fig>
<p>Based on the above calculation, it is found that the word <inline-graphic xlink:href="CMC_34072-inline-48.png"/> (C&#x0101;t&#x0101;rana) will be most probable translation for the word <inline-graphic xlink:href="CMC_34072-inline-49.png"/> (saamaany). The structural transfer phase has identified that the sequence does not need any rearrangements to retain the target language grammar. The final translated text for this input sentence is shown in <xref ref-type="table" rid="table-11">Table 11</xref>. And it is also compared with the reference sentence mentioned in the corpus. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> shows the Recall of the proposed method.</p>
<table-wrap id="table-11"><label>Table 11</label><caption><title>Analysis of generated target text with its reference text</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
</colgroup>
<tbody>
<tr>
<td align="center"><inline-graphic xlink:href="CMC_34072-inline-47.png"/></td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-6"><label>Figure 6</label><caption><title>Comparison of Recall</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_34072-fig-6.png"/></fig>
<p>From <xref ref-type="table" rid="table-12">Table 12</xref>, the target sentence is found to be the same as it has to be with respect to the reference sentence. The proposed HMM<sub>lstm</sub>RL machine translation system has been evaluated using various sentences (restricted to health domain) and the Bilingual Evaluation Understudy (BLEU) score is calculated to be 0.68. The corpus size when increased, further leads to increase in distortion noise and thus producing a poor translation accuracy. It has been found ideally that the unique word should be kept at 30000 and further increase in the corpus size leads to poor translation due to increase in noise. The proposed system is also compared with the other statistical machine translation system in terms of BLEU (Bilingual Evaluation Understudy) score, which is shown in <xref ref-type="table" rid="table-13">Table 13</xref>. The HMM<sub>lstm</sub>RL approach is found to have a BLEU score of 0.7637 and it is also noted that the BLEU score has improved by few percentages when compared with the pivot-based approach.</p>
<table-wrap id="table-12"><label>Table 12</label><caption><title>Comparison with different corpus sizes</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">S. No.</th>
<th align="left" rowspan="2">Corpus Size (In number of words)</th>
<th align="center" colspan="2">Bi-directional HMM<sub>lstm</sub>RL</th>
<th align="center" colspan="2">Proposed HMM<sub>lstm</sub>RL</th>
</tr>
<tr>
<th/>
<th align="left">Precision (in&#x0025;)</th>
<th align="left">Recall (in&#x0025;)</th>
<th align="left">Precision (in&#x0025;)</th>
<th align="left">Recall (in &#x0025;)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">10000</td>
<td align="left">68</td>
<td align="left">58</td>
<td align="left">65</td>
<td align="left">63</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">20000</td>
<td align="left">70</td>
<td align="left">69</td>
<td align="left">76</td>
<td align="left">72.5</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">30000</td>
<td align="left">81</td>
<td align="left">76</td>
<td align="left">87</td>
<td align="left">86.5</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-13"><label>Table 13</label><caption><title>Comparison of various statistical machine translation system using BLEU score</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">S. No.</th>
<th align="left">Methodology</th>
<th align="left">Source (L)</th>
<th align="left">Target (L)</th>
<th align="left">BLEU Score</th>
</tr>
</thead>
<tbody>
<tr>
<th align="left">1</th>
<th align="left">HMM [<xref ref-type="bibr" rid="ref-18">18</xref>]</th>
<th align="left">English</th>
<th align="left">Part of Speech</th>
<th align="left">0.2287</th>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Topic-based coherence model [<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td align="left">Indonesian</td>
<td align="left">Japanese</td>
<td align="left">0.1723</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Proposed bi-directional HMM<sub>lstm</sub>RL</td>
<td align="left">Hindi</td>
<td align="left">Tamil</td>
<td align="left">0.7394</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Proposed HMM<sub>lstm</sub>RL</td>
<td align="left">Hindi</td>
<td align="left">Tamil</td>
<td align="left">0.7637</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5"><label>5</label><title>Conclusion</title>
<p>In today&#x0027;s multicultural business world, there is an increased use of natural languages for establishing communication in the business environment. As a result of globalisation, machine translation systems are required to aid in communication between various organisations. Thus, the demand for translation systems into various language pairs has increased. It can also help in improving communication between people of different origins. This paper presented a HMM<sub>lstm</sub>RL model HMM integrated with bi-directional LSTM with machine learning technique with the MapReduce framework for clustering. Finally, the proposed HMM<sub>lstm</sub>RL model uses voting for language whether it belongs to a keyword or not. The analysis of the results showed that the proposed model exhibits higher performance than the existing techniques.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to thank the research and development departments of B V Raju Institute of Technology, Vasavi College of Engineering, C.B.I.T and K.S.R.M College of Engineering for supporting this work.</p>
</ack>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> The author(s) received no specific funding for this study.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Chowdhary</surname></string-name></person-group>, &#x201C;<article-title>Natural language processing</article-title>,&#x201D; in <source>Fundamentals of Artificial Intelligence</source>, New Delhi: Springer, pp. <fpage>603</fpage>&#x2013;<lpage>649</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Qiu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Shao</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Dai</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Pre-trained models for natural language processing: A survey</article-title>,&#x201D; <source>Science China Technological Sciences</source>, vol. <volume>63</volume>, no. <issue>10</issue>, pp. <fpage>1872</fpage>&#x2013;<lpage>1897</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Xiong</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Mamon</surname></string-name></person-group>, &#x201C;<article-title>An enabling framework for automated extraction of signals from market information in real time</article-title>,&#x201D; <source>Knowledge-Based Systems</source>, vol. <volume>246</volume>, pp. <fpage>108612</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. W.</given-names> <surname>Otter</surname></string-name>, <string-name><given-names>J. R.</given-names> <surname>Medina</surname></string-name> and <string-name><given-names>J. K.</given-names> <surname>Kalita</surname></string-name></person-group>, &#x201C;<article-title>A survey of the usages of deep learning for natural language processing</article-title>,&#x201D; <source>IEEE Transactions on Neural Networks and Learning Systems</source>, vol. <volume>32</volume>, no. <issue>2</issue>, pp. <fpage>604</fpage>&#x2013;<lpage>624</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Barzilay</surname></string-name> and <string-name><given-names>K.</given-names> <surname>McKeown</surname></string-name></person-group>, &#x201C;<article-title>Sentence fusion for multi document news summarization</article-title>,&#x201D; <source>Computer Linguistic</source>, vol. <volume>31</volume>, no. <issue>3</issue>, pp. <fpage>297</fpage>&#x2013;<lpage>328</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Zhong</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Ghosh</surname></string-name></person-group>, &#x201C;<article-title>Generative model-based document clustering: A comparative study</article-title>,&#x201D; <source>Knowledge and Information Systems</source>, vol. <volume>8</volume>, no. <issue>3</issue>, pp. <fpage>374</fpage>&#x2013;<lpage>384</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>U.</given-names> <surname>Srilakshmi</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Alghamdi</surname></string-name>, <string-name><given-names>V. A.</given-names> <surname>Vuyyuru</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Veeraiah</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Alotaibi</surname></string-name></person-group>, &#x201C;<article-title>A secure optimization routing algorithm for mobile ad hoc networks</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>10</volume>, pp. <fpage>14260</fpage>&#x2013;<lpage>14269</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Siddharthan</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Nenkova</surname></string-name> and <string-name><given-names>K.</given-names> <surname>McKeown</surname></string-name></person-group>, &#x201C;<article-title>Information status distinctions and referring expressions: An empirical study of references to people in news summaries</article-title>,&#x201D; <source>Computer Linguistics</source>, vol. <volume>37</volume>, no. <issue>4</issue>, pp. <fpage>811</fpage>&#x2013;<lpage>842</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhou</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Zong</surname></string-name></person-group>, &#x201C;<article-title>Abstractive cross-language summarization via translation model enhanced predicate argument structure fusing</article-title>,&#x201D; <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>, vol. <volume>24</volume>, no. <issue>10</issue>, pp. <fpage>1842</fpage>&#x2013;<lpage>1853</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tong</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Song</surname></string-name></person-group>, &#x201C;<article-title>Cross-lingual topic discovery from multilingual search engine query log</article-title>,&#x201D; <source>ACM Transactions on Information Systems</source>, vol. <volume>35</volume>, no. <issue>2</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>28</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Vuli&#x0107;</surname></string-name>, <string-name><given-names>W.</given-names> <surname>De Smet</surname></string-name> and <string-name><given-names>M. F.</given-names> <surname>Moens</surname></string-name></person-group>, &#x201C;<article-title>Cross-language information retrieval models based on latent topic models trained with document-aligned comparable corpora</article-title>,&#x201D; <source>Information Retrieval</source>, vol. <volume>16</volume>, no. <issue>3</issue>, pp. <fpage>331</fpage>&#x2013;<lpage>368</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Sra</surname></string-name></person-group>, &#x201C;<article-title>A short note on parameter approximation for von mises-fisher distributions: And a fast implementation of <italic>I <sub>s</sub></italic> (<italic>x</italic>)</article-title>,&#x201D; <source>Computational Statistics</source>, vol. <volume>27</volume>, no. <issue>1</issue>, pp. <fpage>177</fpage>&#x2013;<lpage>190</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Esuli</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Moreo</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Sebastiani</surname></string-name></person-group>, &#x201C;<article-title>Funnelling: A new ensemble method for heterogeneous transfer learning and its application to cross-lingual text classification</article-title>,&#x201D; <source>ACM Transactions on Information Systems</source>, vol. <volume>37</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>30</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Heyman</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Vuli&#x0107;</surname></string-name> and <string-name><given-names>M. F.</given-names> <surname>Moens</surname></string-name></person-group>, &#x201C;<article-title>C-Bilda extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content</article-title>,&#x201D; <source>Data Mining and Knowledge Discovery</source>, vol. <volume>30</volume>, no. <issue>5</issue>, pp. <fpage>1299</fpage>&#x2013;<lpage>1323</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Hao</surname></string-name> and <string-name><given-names>M. J.</given-names> <surname>Paul</surname></string-name></person-group>, &#x201C;<article-title>An empirical study on cross lingual transfer in probabilistic topic models</article-title>,&#x201D; <source>Computer Linguistics</source>, vol. <volume>46</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>40</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. H.</given-names> <surname>Chang</surname></string-name> and <string-name><given-names>S. Y.</given-names> <surname>Hwang</surname></string-name></person-group>, &#x201C;<article-title>A word embedding-based approach to cross-lingual topic modelling</article-title>,&#x201D; <source>Knowledge and Information Systems</source>, vol. <volume>63</volume>, no. <issue>6</issue>, pp. <fpage>1529</fpage>&#x2013;<lpage>1555</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. V.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Yadav</surname></string-name></person-group>, &#x201C;<article-title>Word sense based hindi-tamil statistical machine translation</article-title>,&#x201D; <source>International Journal of Intelligent Information Technologies</source>, vol. <volume>14</volume>, no. <issue>1</issue>, pp. <fpage>17</fpage>&#x2013;<lpage>27</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Perikos</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Kardakis</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Hatzilygeroudis</surname></string-name></person-group>, &#x201C;<article-title>Sentiment analysis using novel and interpretable architectures of hidden markov models</article-title>,&#x201D; <source>Knowledge-Based Systems</source>, vol. <volume>229</volume>, pp. <fpage>107332</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. L.</given-names> <surname>Pontes</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Huet</surname></string-name>, <string-name><given-names>J. M.</given-names> <surname>Torres-Moreno</surname></string-name> and <string-name><given-names>A. C.</given-names> <surname>Linhares</surname></string-name></person-group>, &#x201C;<article-title>Compressive approaches for cross-language multi-document summarization</article-title>,&#x201D; <source>Data &#x0026; Knowledge Engineering</source>, vol. <volume>125</volume>, pp. <fpage>101763</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Cao</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Zhuge</surname></string-name></person-group>, &#x201C;<article-title>Grouping sentences as better language unit for extractive text summarization</article-title>,&#x201D; <source>Future Generation Computer Systems</source>, vol. <volume>109</volume>, pp. <fpage>331</fpage>&#x2013;<lpage>359</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Alahmadi</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Wali</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Alzahrani</surname></string-name></person-group>, &#x201C;<article-title>TAAM: Topic-aware abstractive arabic text summarisation using deep recurrent neural networks</article-title>,&#x201D; <source>Journal of King Saud University-Computer and Information Sciences</source>, vol. <volume>34</volume>, no. <issue>6</issue>, pp. <fpage>2651</fpage>&#x2013;<lpage>2665</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Geissinger</surname></string-name>, <string-name><given-names>W. A.</given-names> <surname>Ingram</surname></string-name> and <string-name><given-names>E. A.</given-names> <surname>Fox</surname></string-name></person-group>, &#x201C;<article-title>Teaching natural language processing through big data text summarization with problem-based learning</article-title>,&#x201D; <source>Data and Information Management</source>, vol. <volume>4</volume>, no. <issue>1</issue>, pp. <fpage>18</fpage>&#x2013;<lpage>43</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Barros</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Lloret</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Saquete</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Navarro-Colorado</surname></string-name></person-group>, &#x201C;<article-title>NATSUM: Narrative abstractive summarization through cross-document timeline generation</article-title>,&#x201D; <source>Information Processing &#x0026; Management</source>, vol. <volume>56</volume>, no. <issue>5</issue>, pp. <fpage>1775</fpage>&#x2013;<lpage>1793</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Cai</surname></string-name></person-group>, &#x201C;<article-title>Syntax-enriched abstractive summarization</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>199</volume>, pp. <fpage>116819</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>McCallum</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Corrada-Emmanuel</surname></string-name></person-group>, &#x201C;<article-title>Topic and role discovery in social networks with experiments on enron and academic email</article-title>,&#x201D; <source>Journal of Artificial Intelligence Research</source>, vol. <volume>30</volume>, pp. <fpage>249</fpage>&#x2013;<lpage>272</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. Q.</given-names> <surname>Nguyen</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Billingsley</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Du</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Johnson</surname></string-name></person-group>, &#x201C;<article-title>Improving topic models with latent feature word representations</article-title>,&#x201D; <source>Transactions of the Association for Computational Linguistics</source>, vol. <volume>3</volume>, pp. <fpage>299</fpage>&#x2013;<lpage>313</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ruder</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Vuli&#x0107;</surname></string-name> and <string-name><given-names>A.</given-names> <surname>S&#x00F8;gaard</surname></string-name></person-group>, &#x201C;<article-title>A survey of cross-lingual word embedding models</article-title>,&#x201D; <source>Journal of Artificial Intelligence Research</source>, vol. <volume>65</volume>, no. 1, pp. <issue>569&#x2013;630</issue>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Veeraiah</surname></string-name>, <string-name><given-names>O. I.</given-names> <surname>Khalaf</surname></string-name>, <string-name><given-names>C. V. P. R.</given-names> <surname>Prasad</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Alotaibi</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alsufyani</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Trust aware secure energy efficient hybrid protocol for manet</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>9</volume>, pp. <fpage>120996</fpage>&#x2013;<lpage>121005</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Veeraiah</surname></string-name> and <string-name><given-names>B. T.</given-names> <surname>Krishna</surname></string-name></person-group>, &#x201C;<article-title>Trust-aware fuzzyclus-fuzzy nb: Intrusion detection scheme based on fuzzy clustering and Bayesian rule</article-title>,&#x201D; <source>Wireless Networks</source>, vol. <volume>25</volume>, pp. <fpage>4021</fpage>&#x2013;<lpage>4035</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>U.</given-names> <surname>Srilakshmi</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Veeraiah</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Alotaibi</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Alghamdi</surname></string-name>, <string-name><given-names>O. I.</given-names> <surname>Khalaf</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>An improved hybrid secure multipath routing protocol for manet</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>9</volume>, pp. <fpage>163043</fpage>&#x2013;<lpage>163053</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Kollapudi</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Alghamdi</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Veeraiah</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Alotaibi</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Thotakura</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>A new method for scene classification from the remote sensing images</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>72</volume>, no. <issue>1</issue>, pp. <fpage>1339</fpage>&#x2013;<lpage>1355</lpage>, <year>2022</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>
