<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">16534</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2021.016534</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An Attention Based Neural Architecture for Arrhythmia Detection and Classification from ECG Signals</article-title>
<alt-title alt-title-type="left-running-head">An Attention Based Neural Architecture for Arrhythmia Detection and Classification from ECG Signals</alt-title>
<alt-title alt-title-type="right-running-head">An Attention Based Neural Architecture for Arrhythmia Detection and Classification from ECG Signals</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western">
<surname>Mangathayaru</surname>
<given-names>Nimmala</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western">
<surname>Rani</surname>
<given-names>Padmaja</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western">
<surname>Janaki</surname>
<given-names>Vinjamuri</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western">
<surname>Srinivas</surname>
<given-names>Kalyanapu</given-names>
</name>
<xref ref-type="aff" rid="aff-4">4</xref>
</contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western">
<surname>Mathura Bai</surname>
<given-names>B.</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western">
<surname>Sai Mohan</surname>
<given-names>G.</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-7" contrib-type="author">
<name name-style="western">
<surname>Lalith Bharadwaj</surname>
<given-names>B.</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<aff id="aff-1"><label>1</label><institution>Department of IT, VNR Vignana Jyothi Institute of Engineering &#x0026; Technology</institution>, <addr-line>Hyderabad, 500090</addr-line>, <country>India</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of CSE, JNTUH</institution>, <addr-line>Hyderabad, 500085</addr-line>, <country>India</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of CSE, Vaagdevi College of Engineering</institution>, <addr-line>Warangal, 506005</addr-line>, <country>India</country></aff>
<aff id="aff-4"><label>4</label><institution>Kakatiya Institute of Technology and Science</institution>, <addr-line>Warangal, 506015</addr-line>, <country>India</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1">&#x002A;Corresponding Author: Nimmala Mangathayaru. Email: <email>mangathayaru_n@vnrvjiet.in</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2021-07-13"><day>13</day><month>07</month><year>2021</year>
</pub-date>
<volume>69</volume>
<issue>2</issue>
<fpage>2425</fpage>
<lpage>2443</lpage>
<history>
<date date-type="received"><day>04</day><month>1</month><year>2021</year>
</date>
<date date-type="accepted"><day>11</day><month>4</month><year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2021 Mangathayaru et al.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Mangathayaru et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_16534.pdf"></self-uri>
<abstract>
<p>Arrhythmia is ubiquitous worldwide and cardiologists tend to provide solutions from the recent advancements in medicine. Detecting arrhythmia from ECG signals is considered a standard approach and hence, automating this process would aid the diagnosis by providing fast, cost-efficient, and accurate solutions at scale. This is executed by extracting the definite properties from the individual patterns collected from Electrocardiography (ECG) signals causing arrhythmia. In this era of applied intelligence, automated detection and diagnostic solutions are widely used for their spontaneous and robust solutions. In this research, our contributions are two-fold. Firstly, the Dual-Tree Complex Wavelet Transform (DT-CWT) method is implied to overhaul shift-invariance and aids signal reconstruction to extract significant features. Next, A neural attention mechanism is implied to capture temporal patterns from the extracted features of the ECG signal to discriminate distinct classes of arrhythmia and is trained end-to-end with the finest parameters. To ensure that the model&#x2019;s generalizability, a set of five train-test variants are implied. The proposed model attains the highest accuracy of 98.5% for classifying 8 variants of arrhythmia on the MIT-BIH dataset. To test the resilience of the model, the unseen (test) samples are increased by 5x and the deviation in accuracy score and MSE was 0.12% and 0.1% respectively. Further, to assess the diagnostic model performance, AUC-ROC curves are plotted. At every test level, the proposed model is capable of generalizing new samples and leverages the advantage to develop a real-world application. As a note, this research is the first attempt to provide neural attention in arrhythmia classification using MIT-BIH ECG signals data with state-of-the-art performance.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Arrhythmia classification</kwd>
<kwd>arrhythmia detection</kwd>
<kwd>MIT-BIH dataset</kwd>
<kwd>dual-tree complex wave transform</kwd>
<kwd>ECG classification</kwd>
<kwd>neural attention</kwd>
<kwd>neural networks</kwd>
<kwd>deep learning</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Arrhythmia is ubiquitous worldwide and there is still a major population at risk. The ECG signals are highly efficient and used as a gold-standard to detect the presence of arrhythmia and critical conditions such as cardiac arrest. So, detecting arrhythmia using ECG signals is challenging for researchers. In this era of automation, deep learning has made strides in various fields such as computer vision, language processing, and signal processing in developing state-of-the-art models on large scale databases. Further, deep learning has advanced in bio-medical imaging and bio-medical signal processing. Hence, it is aimed to develop a diagnostic model by extracting features from ECG using DT-CWT and processing them with help of the proposed neural architecture.</p>
<p>It is observed that the signals recorded by ECG are a combination of PQRS waves and these waves detect the heart functionality by identifying various characteristics. Certain features are extracted from the signal by pre-processing it with various transformation techniques in which continuous wavelet transform and discrete wavelet transform are frequently used. Sequentially these extracted features are then processed using various learning algorithms such as k-nearest neighbors (k-NN), Support vector machines (SVM), Multi-layer-perceptron (MLP), Singular value decomposition (SVD), etc. The pre-processing techniques are implied for extracting the features from ECG signal widely used Discrete wavelet transformation (DWT), which is sensitive to shift-invariance and downgrades the quality of the signal while reconstructing the decomposed signal. In the subsequent processing step, most of the methods include generic machine learning algorithms (classification or clustering algorithms) or an MLP which does not capture temporal invariances and eventually harms the performance by reducing the generalization. Hence the two key steps to provide a diagnostic model are, (a) an appropriate pre-processing of the signal (DT-CWT) (b) a processing step to prognosticate the disease (neural attention). To overcome the drawbacks of the existing research, the proposed diagnostic framework implies DT-CWT and neural attention mechanism to provide a significant solution.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Previous Works</title>
<p>The impact of MIT-BIH data was shown by George et al. [<xref ref-type="bibr" rid="ref-1">1</xref>], which acted as a catalyst and gave rise to numerous works that provide insights on automated devices to diagnose arrhythmia. Markos et al. [<xref ref-type="bibr" rid="ref-2">2</xref>] used time-domain analysis to extricate features and then arranged them in distinct combinations, which are utilized as input for neural networks. Sixty-three different types of neural networks were formed. The output of these networks was deployed to a decision tree to diagnose arrhythmia. Karimifard [<xref ref-type="bibr" rid="ref-3">3</xref>] worked on modelling of signals, who later used a Hermitian basis function to get a feature vector and sent to a k-nearest neighbor classifier to classify seven types of arrhythmia, which [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>] obtained sensitivity and specificity of 99.0% and 99.84% respectively. He also concluded that the size of the feature vector affected the training time of the model. Mohammadzadeh et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] took features from the signal by linear and non-linear methods, which were reduced by Gaussian discriminant analysis (GDA) and used an SVM to recognize six classes of arrhythmia with a sensitivity of 95.7% and specificity of 99.40%. Chi et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] focused on quick prediction of the disease by pulling out PQRST features from the ECG signal and later used Linear Discriminant Analysis for grouping five different classes of arrhythmia and achieved an accuracy of 96.23%.</p>
<p>A unique use of the kernel Adatron algorithm was combined with SVM by Majid et al. [<xref ref-type="bibr" rid="ref-6">6</xref>]. He explained the drawbacks of a multi-layered-perceptron (MLP) and differentiated the training and testing time of these two methods. Hamid et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] performed Complex wavelet transformation (CWT), Discrete wavelet transformation (DWT), Discrete cosine transformation (DCT) feature extraction methods separately on the signal and formed four different structures using MLP, then the other four using SVM and deduced the efficient use of a feature extraction method by the training time of the model [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. Oscar [<xref ref-type="bibr" rid="ref-8">8</xref>] preprocessed the signals by QRS extraction method and used fuzzy KNN, MLP with backpropagation, and MLP with scaled conjugate gradient backpropagation (GBP) to get the output matrix. Later these three matrices were combined and sent to the fuzzy inference system to get the result. This achieved an accuracy of 98%. Roland et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] presented an artificial neural network that took signals that are preprocessed by Fast Fourier Transformation (FFT) as an input and then categorized five classes of arrhythmia.</p>
<p>The importance of PQRST wave properties was also discussed here. Yeh et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] proposed a novel preprocessing method along with Cluster analysis (CA) to classify 5 distinct classes and attained a total classification accuracy (TCA) of 94.30%. Stefan [<xref ref-type="bibr" rid="ref-11">11</xref>] developed an android application for real-time detection of arrhythmia by Decision Trees (DT). This model clocked a sensitivity of 89.5% and specificity of 80.6%. Elgendi [<xref ref-type="bibr" rid="ref-12">12</xref>] introduced the application of the moving averages method on ECG signals for detection of P and T waves by addressing four sources of noise which altered the quality of the signal and obtained a sensitivity of 98.05% and specificity of 98.86%. Manu et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] extracted features using DTCWT and merged another four features (AC power, Kurtosis, Skewness, and timing information). This feature set was passed into an MLP and got an accuracy of 94.64% and a sensitivity of 94.6%. Ahmet et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] showed a comparison between the performance of bagged decision trees and a single decision tree with the input of nine features which was taken from ECG signal by applying Low Pass Filter, High Pass Filter, form factor (FF) computing, FF ratio to previous one (FFR), RR ration to the previous RR ration (RRR), RR difference from mean RR value (RRM), skewness, linear predictive coding (LPC) and the cumulated ensemble method outperformed a single decision tree with an accuracy of 99.15%. Mehrdad [<xref ref-type="bibr" rid="ref-15">15</xref>] improved the signal quality by un-decimated wavelet transformation (UWT) and then proposed a method that combined Negatively Correlated Learning (NCL) and Mixture of Experts (ME) which is known to provide an excellent recognition rate. The model is used to group premature ventricular contraction (PVC) arrhythmia and Normal heartbeat classes and achieved accuracy, sensitivity, the specificity of 96.02%, 92.27%, and 93.72% respectively. Ping [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed an adaptive feature extraction method based on wavelet transformation and a modified voting mechanism consists of K-means clustering, one against one SVM to enhance the recognition rate and got an accuracy of 89.2%. Joachim [<xref ref-type="bibr" rid="ref-17">17</xref>] worked on the categorization of poor and good signals. An alarm is set off when the parameters are not within a given scale. QRS extraction method along with SVM was used to reach this objective. Patricia [<xref ref-type="bibr" rid="ref-18">18</xref>] presented a Learning Vector Quantization (LVQ) algorithm with SVM to classify arrhythmia. However, the comparisons were made with simulated data. A total of 15 classes were grouped with three different architectures and the best architecture got an accuracy of 99.16%. Ali et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] performed a diagnosis of arrhythmia with the help of Alex net and prior QRS detection was done. The signal was converted to a 256 &#x00D7; 256 sized image and then passed into the network. The recognition rate and accuracy are 98.5% and 92%. Joy [<xref ref-type="bibr" rid="ref-20">20</xref>] implemented DCT transformation of waves and used Probabilistic Neural Network (PNN) for efficient detection of disease [<xref ref-type="bibr" rid="ref-11">11</xref>]. Vasileios [<xref ref-type="bibr" rid="ref-21">21</xref>] showed false beat detection effectively by detecting QRS peaks then filtering false beats using SVM and concluded that QRS peaks are very important for the detection. Rashid et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a new method that showed promising results by using Gaussian mixture modeling (GMM) with expectation maximization (EM), Combined with statistical and morphological features. Accuracy for class-oriented is 99.6% and for subject-oriented is 96.15%. Serkan et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] also made an android application using 1-D convolution to classify supraventricular ectopic beat (SVEB) and ventricular ectopic beat (VEB). FFT with DCT was used in feature extraction and obtained an accuracy of 99.0%, 97.2% for each class.</p>
<p>From the previous research works, it is observed that (pre-processing part) many methods do not capture temporal relationship among the data. If they capture temporal dependencies, they do not persist with long term dependencies. So, if long term dependencies are provided there are no sequential patterns, which provide attention to the network determining the importance factor. So, these loops are overhauled with the use of attention embedded neural architecture by capturing long term temporal dependencies. Further, some loops are addressed in the pre-processing section and they are overridden by utilising DT-CWT and are mentioned in the successive section.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Contributions</title>
<p>In this research, the contributions to the body of the knowledge are mentioned as,
<list list-type="bullet">
<list-item>
<p>Dual-Tree Complex Wavelet Transform (DT-CWT) method is implied to overhaul shift-invariance and aids the signal reconstruction to extract significant features. Further, a small set of features are extracted using the Pan-Tomkins algorithm and are adjoined with the features extracted from DT-CWT.</p></list-item>
<list-item>
<p>A neural attention mechanism is implied to capture temporal patterns from the extracted features of the ECG signal and to discriminate distinct classes of arrhythmia. The proposed attention model is end-to-end trained by carefully optimizing the hyperparameters.</p></list-item>
</list></p>
</sec>
<sec id="s4">
<label>4</label>
<title>Methodology</title>
<p>As mentioned, the two important steps are involved to complete the proposed automated system. This section aims to give a clear understanding of mathematics related to these two steps and explains the unique capabilities.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset Description</title>
<p>In this paper, ECG recordings acquired by the arrhythmia laboratory of Boston&#x2019;s Beth Israel Hospital are used and this database is known as the MIT-BIH arrhythmia database. ECG recordings are collected using Del Mar Avionics model 445 two-channel reel to reel Holter recorders. These signals were filtered using bandpass filters with frequency in the range of 0.1&#x2013;100 HZ and are digitized by Del Mar Avionics 660 playback unit with a sampling rate of 360 samples per second. This database consists of forty-eight half-hour excerpts of two-channel twenty-four-hour, ECG recordings from 47 subjects as record number 201. The first twenty-three records are drawn from a collection of four thousand Holter tapes and the other records include uncommon heartbeat irregularities but have great clinical significance. The subjects include twenty-five men and twenty-two women who are aged between twenty-three to eighty-nine.</p>
<p>The most frequently used ECG leads in this database are modified limb lead 2 (MLII) for channel one and v1 for the other channel. V2, V4 and V5 are also used occasionally, based on the subjects. Fusion Ventricular (FV), VEB, right bundle branch block (RBB), paced beat (PB), Normal (N), ventricular contraction (VC), left bundle branch block (LBB), atrial premature beat (APB) are the different classes of arrhythmia that are used in the task of classification to evaluate the model.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Feature Extraction (Pre-Processing)</title>
<p>As an insight, pre-processing step is important to capture appropriate features which in turn obliges prognosticate arrhythmia. After extensive research on what would be the best practice to get the features based on prior knowledge, it is to be deduced that wavelet transformation for the first step is the best practice for MIT-BIH arrhythmia. The QRS complex signal is pulled with a sample of 256 of which 128 samples are considered from the left side of the R peak and 128 samples from the right. Later wavelet transformation is applied to get a set of required features. As a note, the database reflects certain noise in the signal and the cause of it are described below,
<list list-type="bullet">
<list-item>
<p>The frequency from the power supply usually manipulates the signal, this is known as the powerline interface.</p></list-item>
<list-item>
<p>Our muscles often tend to contract and expand, which regularly gets combined with cardiac muscles and end up giving a signal with noise</p></list-item>
<list-item>
<p>A signal quality often depends on the contact between the lead and skin, there are some times that a movement by the patient corrupts the signal and this is described as motion artefact.</p></list-item>
</list></p>
<p>The above-discussed problems are solved by implementing the Pan-Tomkins and including it as one of the important features for signal pre-processing. These findings are addressed in discovering the QRS complex by Pan&#x2013;Tomkins [<xref ref-type="bibr" rid="ref-24">24</xref>]. A signal undergoes four steps in this algorithm. Initially, to attenuate noise, the signal is passed to a bandpass filter, which eliminates motion artefact and makes the signal more stable. A differentiator is used to get the slope of the signal and solve the baseline drift problem. This is followed by a squaring function that helps to remove get absolute value and limit false positives generated by T waves. Finally, moving-window integration is used to smooth the curve and get information about the slope of the signal. The steps are implemented for a signal (considered from the database) and are visually illustrated in <?A3B2 "fig1",5,"anchor"?><xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Wavelet transformation is applied to the extracted QRS complex signal where a wavelet acts as a window function. All the wavelet transformations are in the compressed or shifted form of the mother wavelet and the different versions of the mother wavelet are described in <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-3">(3)</xref>. In <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> [<xref ref-type="bibr" rid="ref-25">25</xref>], S is the inverse of the frequency of the signal, which can be used to get low and high-frequency signals and to make the wave thinner or broader. T is used to translate wavelet across the signal. Wavelet transformation helps in analysing the different frequencies at different locations; This is known as multi-resolution analysis. By changing the values of S, the wavelet can be obtained in expanded or in compressed form, which is known as scaling. For non-stationary waves, CWT is used, however, the upper limit and lower limit of CWT tends to infinity. This means that there would be a huge number of coefficients that are to be calculated at every possible position. (in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>) [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
<p><disp-formula id="eqn-1">
<label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:mrow><mml:mo>|</mml:mo><mml:mi>S</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:msqrt></mml:mfrac><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow><mml:mrow><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow></mml:msubsup><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mi>S</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-2">
<label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:mi>a</mml:mi></mml:msqrt></mml:mfrac><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mi>a</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi mathvariant="normal">w</mml:mi></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-3">
<label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:mi>b</mml:mi></mml:msqrt></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mi>f</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mi>b</mml:mi></mml:mfrac><mml:mo>]</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>where,</p>
<p><disp-formula id="eqn-4">
<mml:math id="mml-eqn-4" display="block"><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-22">
<mml:math id="mml-eqn-22" display="block"><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-5">
<label>(4)</label>
<mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi mathvariant="bold">&#x03A6;</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold">&#x03A6;</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold">&#x03A6;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-6">
<label>(5)</label>
<mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi mathvariant="bold">&#x03A8;</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold">&#x03A8;</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold">&#x03A8;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>To reduce the number of coefficients, DWT is used instead of CWT (<xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>) [<xref ref-type="bibr" rid="ref-27">27</xref>]. This is achieved by choosing a, b in powers of two and so, the DWT is calculated (computationally) by multilevel decomposition. The signal is further passed into a low pass and high pass filter and the two filters utilized are orthonormal by construction. Initially, the signal is passed to a low pass filter to get approximate coefficients and then again to a high pass filter to get detailed coefficients and are downsampled by 2 successively. The approximate coefficients are iteratively processed in the same way to get low pass portions as well as high pass portions. <?A3B2 "fig2",5,"anchor"?><xref ref-type="fig" rid="fig-2">Fig. 2</xref> visually explains the complete overview of the DWT decomposition process.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Steps in Pan&#x2013;Tomkins algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-1.png"/>
</fig>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Transformation of the signal after every individual step using pan Tomkins algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-2.png"/>
</fig>
<p>Yet after decomposition, it still lacks Perfect Reconstruction (PR) and does not provide a shift in-variance. To overcome this, Dual-Tree Complex Wave Transformation (DT-CWT) is used [<xref ref-type="bibr" rid="ref-28">28</xref>]. DTCWT employs a complex-valued scaling function and wavelet. The <xref ref-type="disp-formula" rid="eqn-4">Eqs. (4)</xref> and <xref ref-type="disp-formula" rid="eqn-5">(5)</xref> show the functions used in DT-CWT, where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a real part of the complex-valued function and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the imaginary part of the wavelet function. The main difference is that the <xref ref-type="disp-formula" rid="eqn-4">Eqs. (4)</xref> and <xref ref-type="disp-formula" rid="eqn-5">(5)</xref> has two distinct tree structures and multilevel decomposition is performed twice on the same signal. This is shown in <?A3B2 "fig3",5,"anchor"?><xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<p>In <?A3B2 "fig4",5,"anchor"?><xref ref-type="fig" rid="fig-4">Fig. 4</xref>, Tree(A) is used to acquire the real part coefficients and Tree (B) is utilized to get the imaginary part coefficients. The low pass filter is slowed down by one-fourth of the sample for non-symmetry, which helps in achieving the PR of the signal. The filters used in both the trees are orthonormal to each other and the reverse of decomposition provides the synthesis of the signal. Next, fourth and fifth level detailed coefficients are taken from both the trees. Then 1-D FFT is applied with the obtained features from these levels and another four features are appended to this feature set. The four features are AC power, kurtosis, skewness and timing information and cumulated twenty-eight features from DT-CWT are extracted from this process and are ready to be fed into the classifier.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Classification (Processing)</title>
<p>In the past few years, Feed Forward Neural Networks (FNN&#x2019;s) dominated the automation field. Even after their efficient performance throughout the years, they are still short of remembering long term dependencies and do not work well with the time-series data. Recurrent Neural Network&#x2019;s (RNN&#x2019;s) are used to overcome this drawback. The central theme of the architecture proposed in this section is based on Recurrent Neural Networks and before explaining the proposed architecture, a detailed summary of the Recurrent Neural Networks and their variants are explained below.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>DWT architecture and transformation of a signal at every level of decomposition</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-3.png"/>
</fig>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Architecture of DT-CWT</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-4.png"/>
</fig>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Recurrent Neural Networks (RNN)</title>
<p>Unlike FNN&#x2019;s, RNN&#x2019;s can consider the input of different lengths and provide an output of different lengths. This feature has increased the scope of applications of Deep Learning, such as image captioning and language translation. RNN&#x2019;s have a loop to their unit which helps to store information. These networks play on a recursive function that helps to generate a new state at the time (t) by the information of its old state at the time (t &#x2212; 1). <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref> [<xref ref-type="bibr" rid="ref-29">29</xref>] shows the recursive function which is a tanh function that has weights and linear operations in it.</p>
<p><disp-formula id="eqn-7">
<label>(6)</label>
<mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-8">
<label>(7)</label>
<mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-9">
<label>(8)</label>
<mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math>
</disp-formula></p>
<p>These networks use backpropagation through time to calculate the gradient. As the number of units in the network increases, the gradient value would come close to zero, because of this the weights would not add any information to the network and this problem is known as the <italic>vanishing gradients</italic>.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Long Short-Term Memory (LSTM)</title>
<p>Long Short-Term Memory (LSTM) is a different type of cell used in recurrent neural networks which were found by Sepp et al. [<xref ref-type="bibr" rid="ref-30">30</xref>]. In an LSTM cell, there are various gates, where each gate has its purpose. Lines that are connected to the gates carry a vector that is used to perform a linear operation and provide the output as required. These gates have full control over the information that has to be retained or removed. The gates in LSTM have sigmoid and pointwise operations and the information is initially passed through the LSTM network by the cell state. Only with the cell state, different operations can be performed from the information provided by the cell state to understand LSTM clearly. The working of an LSTM cell is divided into four steps.</p>
<p><bold>Step 1:</bold> To know what information should be forgotten from the previous state. This is done by the &#x2018;forget gate also known as the first layer of LSTM. It takes <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> (a previous hidden state at time t &#x2212; 1) and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (input at time t) and gives a value between 0 and 1, where &#x2018;0&#x2019; means to forget everything and &#x2018;1&#x2019; means to consider the complete information. The reason for getting two values is that this layer uses a sigmoid function. <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref> is used in this layer.</p>
<p><disp-formula id="eqn-10">
<label>(9)</label>
<mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><bold>Step 2:</bold> To take the required information which is done by the input gate that tells what to write to cell state. Two functions act in this gate.</p>
<p><bold>Step 3:</bold> First is the sigmoid layer known as the input gate layer <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>i</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which is used to take the required information.</p>
<p><bold>Step 4:</bold> A tanh function is used which produces a candidate set <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mover><mml:mi>C</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> that is added to the state.</p>
<p><disp-formula id="eqn-11">
<label>(10)</label>
<mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>i</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-12">
<label>(11)</label>
<mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mover><mml:mi>C</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><xref ref-type="disp-formula" rid="eqn-12">Eqs. (12)</xref> and <xref ref-type="disp-formula" rid="eqn-13">(13)</xref> are used to update the cell state; first, <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref> is multiplied by <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> to forget the information from the previous cell state and then add <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>i</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mover><mml:mi>C</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to add the information.</p>
<p><disp-formula id="eqn-13">
<label>(12)</label>
<mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>i</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mover><mml:mi>C</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-14">
<label>(13)</label>
<mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>After updating the cell state the output of the LSTM cell is calculated, this is calculated after the input passes through a sigmoid layer and then through a tanh function of <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. From <?A3B2 "fig5",5,"anchor"?><xref ref-type="fig" rid="fig-5">Fig. 5</xref>, We can say that the output of an LSTM cell depends on the previous state of the cell [<xref ref-type="bibr" rid="ref-31">31</xref>].</p>
<p><disp-formula id="eqn-15">
<label>(14)</label>
<mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">n</mml:mi><mml:mi mathvariant="italic">h</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>An LSTM unit</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-5.png"/>
</fig>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Gated Recurrent Units (GRU)</title>
<p>The gated recurrent unit (GRU) [<xref ref-type="bibr" rid="ref-32">32</xref>] is another variant of the Recurrent Neural Network and is similar to the LSTM with some changes. Due to the changes which are made, GRU tends to work faster than the LSTM network and gives an advantage over it and this can also be explained in 4 steps that are below:</p>
<p><bold>Update Gate</bold>: The functionality of the update gate is to decide what information is to be taken from the previous cell. It takes <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> (a previous hidden state at time t &#x2212; 1) and <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (input at time t) then uses the sigmoid function to give a value between 0 and 1.</p>
<p><disp-formula id="eqn-16">
<label>(15)</label>
<mml:math id="mml-eqn-16" display="block"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><bold>Reset Gate</bold>: Reset gate works the same as the forget gate in LSTM.</p>
<p><disp-formula id="eqn-17">
<label>(16)</label>
<mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>Current information: By using the reset gate, <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is calculated by performing Hadamard product (pointwise operation) with reset gate and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, then, it is added with the input (multiplied with its weight). Performing this would give the information that is taken from the previous hidden state using a reset gate.</p>
<p><disp-formula id="eqn-18">
<label>(17)</label>
<mml:math id="mml-eqn-18" display="block"><mml:mover><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p><bold>Output:</bold> update gate is employed to get the final output <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the GRU cell. Hadamard pointwise operation and sum operation are used to get the output [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
<p><disp-formula id="eqn-19">
<label>(18)</label>
<mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mover><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math>
</disp-formula></p>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Bi-Directional Units (Bi-LSTM and Bi-GRU)</title>
<p>It is seen that bi-directional RNN leverages performance compared to that of unidirectional RNN on speech data. Firstly, the state neurons are divided into two different time directions which are considered as forward and reverse states and the output from the reverse state is not connected to the input of forwarding states, and vice versa. With the assistance of these two sequential time directions, the input data assess the future and past dependencies. This helps to understand long-term dependencies. While training bi-directional RNN&#x2019;s the weighs are updated not only <italic>via</italic> forwarding pass but also through backward pass [<xref ref-type="bibr" rid="ref-34">34</xref>]. Additionally, it is observed that bi-directional LSTM units outperform in phenome classification and recognition tasks with fewer computations <italic>i.e</italic>., epochs [<xref ref-type="bibr" rid="ref-35">35</xref>]. Similarly, Bi-directional GRU&#x2019;s can draw desirable outcomes similar to that of bi-directional LSTM&#x2019;s [<xref ref-type="bibr" rid="ref-36">36</xref>]. Hence, it is aimed to connect sequential bi-directional LSTM and GRU units cautiously for outperforming the classification of ECG signals by acquiring their temporal patterns as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<?A3B2 "fig6",5,"anchor"?><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>A single gated recurrent unit (GRU)</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-6.png"/>
</fig>
</sec>
<sec id="s4_3_5">
<label>4.3.5</label>
<title>Proposed Neural Architecture</title>
<p>The attention mechanism in neural networks was first implied by Dzmitry et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] to memorize long sequences in decoder architecture. A neural architecture is proposed with an embedded attention mechanism for the classification of 8 distinct kinds of arrhythmia from ECG signals. The pre-processed signal from DT-CWT is fed into the proposed neural architecture. The input is merged into two sequential stacked layers with two variant patterns. At first pattern, bi-directional LSTM units are sequentially arranged with respective layer normalization and dropout layers [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>]. In the second pattern, bi-directional GRU units are stacked similar to that of the previous pattern. Then, the output sequence from pattern-i is multiplied with pattern-ii to imply attention.</p>
<p>For a definite time step &#x2018;t&#x2019;, both the bi-directional LSTM and GRU sequence units attempt to perform attention by a scalar product as mentioned in <xref ref-type="disp-formula" rid="eqn-19">Eq. (19)</xref>. This attention mechanism is proposed as global attention to extracting invariant temporal patterns [<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
<p><disp-formula id="eqn-20">
<label>(19)</label>
<mml:math id="mml-eqn-20" display="block"><mml:mrow><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math>
</disp-formula></p>
<p>Then the resultant multiplied output sequence proceeds as input to a GRU layer and next fed into a fully connected feed-forward network. The complete model architecture and its related parameters are depicted in <?A3B2 "fig7",5,"anchor"?><xref ref-type="fig" rid="fig-7">Fig. 7</xref>. The fully connected network consists of 128 unity in the first layer with ReLU as activation and the final layer consists of 8 neurons which are activated with softmax. The complete model consumed 106K trainable weights and negligible non-trainable weights consumed with the usage of layer normalization layers. The model was trained on 5 variant test patters.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Proposed neural architecture implied with an attention mechanism</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-7.png"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Results and Discussion</title>
<p>As mentioned above, to leverage the model performance, training and testing samples are split into multiple variants ranging from 10% to 50%. The complete analysis is carried out on various standard classification metrics and to study the proposed model behaviour, an accuracy score is chosen as the gold standard. Similarly, to study class wise performance precision, f-1 score recall is utilized. MSE is used to assess the predictability of the model which depicts the error attained due to imperfect predictions. Finally, AUC-ROC curves are generated to assess the diagnostic performance of the proposed model [<xref ref-type="bibr" rid="ref-41">41</xref>].</p>
<p>As AUC-ROC curves are sample invariant in nature as they are insensitive to the alterations implied in the class distributions. These curves are plotted class-wise to interpret the performance of the model at each class level. As a note, AUC-ROC visualizations can be obliged as they decouple the performance of the classifier from skewness in classes and error costs presented in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The proposed model is trained with two variant batches of 32 and 64 respectively. All the above-mentioned classification metrics are evaluated for all the test variants with two different batch sets. Large batches are acquired during training neural networks to minimize the generalization gap [<xref ref-type="bibr" rid="ref-42">42</xref>]. Hence, a large set of batches are considered with sizes of 32 and 64. (illustrated in the <?A3B2 "tbl1",5,"anchor"?><xref ref-type="table" rid="table-1">Tab. 1</xref>). In the feature extraction step, the Pan Tomkins method is employed to extract QRS points of the signal which play an important role in determining r-peak which helps to detect heartbeats. A large set of features are drawn out by using DT-CWT. This transformation is shift-invariant and provides PR. Most of the signal transformation methods lack these properties which can cause imperfect prediction and increase the chance of misclassification.</p>
<?A3B2 "fig8",5,"anchor"?><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Visualizations of AUC-ROC curves for 64 batch for test size from 10%, 20%, 30%, 40%, 50%</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_16534-fig-8.png"/>
</fig>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Proposed model performance on the variant test splits with 32, 64 batches</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Batch size</th>
<th>Test (%)</th>
<th>Accuracy (%)</th>
<th>MSE</th>
<th>Precision (avg)</th>
<th>F1-score (avg)</th>
<th>Recall (avg)</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">32</td>
<td>10</td>
<td>98.34</td>
<td>0.16268</td>
<td>0.94</td>
<td>0.92</td>
<td>0.93</td>
</tr>
<tr>
<td>20</td>
<td>98.43</td>
<td>0.14543</td>
<td>0.96</td>
<td>0.93</td>
<td>0.94</td>
</tr>
<tr>
<td>30</td>
<td>98.38</td>
<td>0.14788</td>
<td>0.94</td>
<td>0.94</td>
<td>0.94</td>
</tr>
<tr>
<td>40</td>
<td>98.45</td>
<td>0.14214</td>
<td>0.96</td>
<td>0.94</td>
<td>0.94</td>
</tr>
<tr>
<td>50</td>
<td>98.22</td>
<td>0.16842</td>
<td>0.95</td>
<td>0.94</td>
<td>0.94</td>
</tr>
<tr>
<td rowspan="5">64</td>
<td>10</td>
<td>98.44</td>
<td>0.15229</td>
<td>0.96</td>
<td>0.93</td>
<td>0.94</td>
</tr>
<tr>
<td>20</td>
<td>98.28</td>
<td>0.16694</td>
<td>0.95</td>
<td>0.94</td>
<td>0.94</td>
</tr>
<tr>
<td>30</td>
<td>98.30</td>
<td>0.16446</td>
<td>0.95</td>
<td>0.93</td>
<td>0.94</td>
</tr>
<tr>
<td>40</td>
<td>98.52</td>
<td>0.14089</td>
<td>0.96</td>
<td>0.93</td>
<td>0.94</td>
</tr>
<tr>
<td>50</td>
<td>98.26</td>
<td>0.17005</td>
<td>0.94</td>
<td>0.93</td>
<td>0.94</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Previous literature work on MIT-BIH dataset</title>
</caption>
<table>
<colgroup>
<col charoff="90pt"></col>
<col charoff="130pt"></col>
<col charoff="100pt"></col>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Authors</th>
<th>F.E. method</th>
<th>Classification method</th>
<th>Classes</th>
<th>ACC (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Markos et al. [<xref ref-type="bibr" rid="ref-2">2</xref>]</td>
<td>Time-frequency analysis</td>
<td>Neural network</td>
<td>2</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Karimifard [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>Hermitian basis function</td>
<td>KNN</td>
<td>7</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Mohammadzadeh et al. [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>Linear and non-linear analysis<break/>with GDA</td>
<td>SVM</td>
<td>6</td>
<td>99.16</td>
</tr>
<tr>
<td>Chi et al [<xref ref-type="bibr" rid="ref-5">5</xref>].</td>
<td>Qualitative feature selection</td>
<td>LDA</td>
<td>5</td>
<td>96.23</td>
</tr>
<tr>
<td>Oscar [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>Manual segmentation</td>
<td>NN with a fuzzy system</td>
<td>5</td>
<td>98</td>
</tr>
<tr>
<td>Roland et al. [<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td>FFT</td>
<td></td>
<td>6</td>
<td>98.6</td>
</tr>
<tr>
<td>Yeh et al. [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>Qualitative feature selection</td>
<td>Cluster analysis</td>
<td>5</td>
<td>94.30</td>
</tr>
<tr>
<td>Manu et al. [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>DTCWT</td>
<td>MLP</td>
<td>5</td>
<td>94.64</td>
</tr>
<tr>
<td>Ahmet et al. [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>Time-domain<break/>FE methods</td>
<td>Decision trees</td>
<td>2</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Mehrdad [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>SWT</td>
<td>Negative corelation learning</td>
<td>2</td>
<td>96.02</td>
</tr>
<tr>
<td>Ping [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>WT</td>
<td>K-mean &#x0026; SVM</td>
<td>12</td>
<td>98.92</td>
</tr>
<tr>
<td>Patricia [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Noise removal<break/>FPT<break/>Segmentation</td>
<td>Neural network</td>
<td>15</td>
<td>99.16</td>
</tr>
<tr>
<td>Ali et al. [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Transferred deep learning</td>
<td>Conventional ANN</td>
<td>1</td>
<td>92.4</td>
</tr>
<tr>
<td>Joy et al. [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>DCT<break/>PCA</td>
<td>FFNN</td>
<td>5</td>
<td>99.52</td>
</tr>
<tr>
<td>Vasileios [<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>DWT</td>
<td>SVM</td>
<td>3</td>
<td>95.35</td>
</tr>
<tr>
<td>Rashid et al. [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>RR interval FE,<break/>HOS FE,<break/>GMM FE</td>
<td>Ensemble method bootstrap</td>
<td>5</td>
<td>99.70</td>
</tr>
<tr>
<td>Proposed</td>
<td>Pan Tompkins<break/>DT-CWT</td>
<td>Neural attention<break/>BiGRU &#x002B; BiLSTM</td>
<td>8</td>
<td>98.5</td>
</tr>
<tr>
<td colspan="5">&#x2018;&#x2013;&#x2019;: Describes the unavailability of the concerned information.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The proposed network used adam [<xref ref-type="bibr" rid="ref-43">43</xref>] as an optimizer with a learning rate of <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The neural network uses categorical cross-entropy as the objective function for stochastic optimization of neural network with backpropagation.</p>
<p><disp-formula id="eqn-21">
<label>(20)</label>
<mml:math id="mml-eqn-21" display="block"><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>Generally, RNNs understand the temporal dependencies but they lack understanding in long term dependencies where LSTM overcomes the problems of RNN by understating long-term temporal relationships in the data. But LSTMs are computationally expensive and do have the problem of gradient vanishing with increasing units to a greater extent. Whereas GRU contains fewer gates compared to that of LSTM networks and overcomes the problems of LSTMs. GRUs are computationally faster compared to RNN&#x2019;s and LSTMs. As most of the research focuses on designing a neural architecture utilizing these units with changing the number of time steps, units and stacking pattern. To leverage predictive capability neural attention is implied by undersetting the temporal pattern extracted from the signal. So, by providing neural attention, the minute redundancies and noise captured during feature extraction can be regulated to a greater extent. This provides greater performance compared to that of remaining neural networks. Various previous work is studied and curated, and our method outperforms the existing literature, and the depicted results are tabulated (illustrated in <?A3B2 "tbl2",5,"anchor"?><xref ref-type="table" rid="table-2">Tab. 2</xref>).</p>
<p>As a note, the research was conducted to study arrhythmia without using signal patterns <italic>i.e</italic>., carried out by classifying variant attributes involved in predicting cardiovascular diseases [<xref ref-type="bibr" rid="ref-44">44</xref>] and also carried by implying PPG signals [<xref ref-type="bibr" rid="ref-45">45</xref>]. To see the future perspective of the proposed work, it can be figured out traditionally, the current deep learning applications have considered existing distance functions in the research literature for similarity computations but did not try to fit in new functions for similarity computations [<xref ref-type="bibr" rid="ref-46">46</xref>&#x2013;<xref ref-type="bibr" rid="ref-50">50</xref>]. There is a possibility to devise threshold and similarity functions to suit deep learning applications [<xref ref-type="bibr" rid="ref-51">51</xref>&#x2013;<xref ref-type="bibr" rid="ref-55">55</xref>]. For instance, recent research contributions propose various similarity and threshold functions for temporal pattern mining which can be redesigned to suit deep learning applications [<xref ref-type="bibr" rid="ref-56">56</xref>&#x2013;<xref ref-type="bibr" rid="ref-60">60</xref>].</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>In this research, a novel attention-based neural architecture is built to vanquish the loops of existing methods for classifying ECG signals. It can be stated that the proposed model is sample invariant as it has minute error variation when test samples increased five times. AUC-ROC plots are illustrated to provide a vivid understanding of the performance of the proposed diagnostic model. In a worst-case scenario, the model provides a micro averaged AUC of 0.9904. Even with numerous advantages, it is seen that the proposed model can consume high memory while embedding the model into a real-world application. The training procedure adapted is tested on two batches instead of a dynamic sampling is preferred to improve performance. In future, it is aimed to provide a salient model by acquiring humongous data with less computational capability and higher performance.</p>
</sec>
</body>
<back>
<ack>
<p>The authors acknowledge JNTUH/TEQIP-III, for providing research fund (Ref: No. JNTUH/TEQIP-III/CRS/2019/CSE/08).</p>
</ack>
<fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> This research was partially supported by JNTU Hyderabad, India under Grant proceeding number: JNTUH/TEQIP-III/CRS/2019/CSE/08. The authors are grateful for the support provided by the TEQIP-III team.</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. B.</given-names> <surname>George</surname></string-name> and <string-name><given-names>R. G.</given-names> <surname>Mark</surname></string-name></person-group>, &#x201C;<article-title>The impact of the MIT-BIH arrhythmia database</article-title>,&#x201D; <source>IEEE Engineering in Medicine and Biology Magazine</source>, vol. <volume>20</volume>, no. <issue>3</issue>, pp. <fpage>45</fpage>&#x2013;<lpage>50</lpage>, <year>2001</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. G.</given-names> <surname>Markos</surname></string-name> and <string-name><given-names>D. I.</given-names> <surname>Fotiadis</surname></string-name></person-group>, &#x201C;<article-title>Automatic arrhythmia detection based on time and time-frequency analysis of heart rate variability</article-title>,&#x201D; <source>Computer Methods and Programs in Biomedicine</source>, vol. <volume>74</volume>, no. <issue>2</issue>, pp. <fpage>95</fpage>&#x2013;<lpage>108</lpage>, <year>2004</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Karimifard</surname></string-name></person-group>, &#x201C;<article-title>Morphological heart arrhythmia detection using hermitian basis functions and kNN classifier</article-title>,&#x201D; in <conf-name>Int. Conf. of the IEEE Engineering in Medicine and Biology Society</conf-name>, <publisher-loc>New York, USA</publisher-loc>, pp. <fpage>1367</fpage>&#x2013;<lpage>1370</lpage>, <year>2006</year>. </mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. B.</given-names> <surname>Mohammadzadeh</surname></string-name>, <string-name><given-names>S. K.</given-names> <surname>Setarehdan</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Mohebbi</surname></string-name></person-group>, &#x201C;<article-title>Support vector machine-based arrhythmia classification using reduced features of heart rate variability signal</article-title>,&#x201D; <source>Artificial Intelligence in Medicine</source>, vol. <volume>44</volume>, no. <issue>1</issue>, pp. <fpage>51</fpage>&#x2013;<lpage>64</lpage>, <year>2008</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. Y.</given-names> <surname>Chi</surname></string-name>, <string-name><given-names>W. J.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>C. W.</given-names> <surname>Chiou</surname></string-name></person-group>, &#x201C;<article-title>Cardiac arrhythmia diagnosis method using linear discriminant analysis on ECG signals</article-title>,&#x201D; <source>Measurement</source>, vol. <volume>42</volume>, no. <issue>5</issue>, pp. <fpage>778</fpage>&#x2013;<lpage>789</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Majid</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Khorrami</surname></string-name></person-group>, &#x201C;<article-title>A qualitative comparison of artificial neural networks and support vector machines in ECG arrhythmias classification</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>37</volume>, no. <issue>4</issue>, pp. <fpage>3088</fpage>&#x2013;<lpage>3093</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Hamid</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Moavenian</surname></string-name></person-group>, &#x201C;<article-title>A comparative study of DWT, CWT and DCT transformations in ECG arrhythmias classification</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>37</volume>, no. <issue>8</issue>, pp. <fpage>5751</fpage>&#x2013;<lpage>5757</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Oscar</surname></string-name></person-group>, &#x201C;<article-title>Hybrid intelligent system for cardiac arrhythmia classification with fuzzy k-nearest neighbors and neural networks combined with a fuzzy system</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>39</volume>, no. <issue>3</issue>, pp. <fpage>2947</fpage>&#x2013;<lpage>2955</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A. E.</given-names> <surname>Roland</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Choi</surname></string-name></person-group>, &#x201C;<article-title>Using neural networks to predict cardiac arrhythmias</article-title>,&#x201D; in <conf-name>IEEE Int. Conf. on Systems, Man, and Cybernetics</conf-name>, <publisher-loc>Seoul, Korea</publisher-loc>, <publisher-name>IEEE</publisher-name>, pp. <fpage>402</fpage>&#x2013;<lpage>407</lpage>, <year>2012</year>. </mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.-C.</given-names> <surname>Yeh</surname></string-name>, <string-name><given-names>C. W.</given-names> <surname>Chiou</surname></string-name> and <string-name><given-names>H.-J.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>Analyzing ECG for cardiac arrhythmia using cluster analysis</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>39</volume>, no. <issue>1</issue>, pp. <fpage>1000</fpage>&#x2013;<lpage>1010</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Stefan</surname></string-name></person-group>, &#x201C;<article-title>Real-time ECG monitoring and arrhythmia detection using Android-based mobile devices</article-title>,&#x201D; in <conf-name>Annual Int. Conf. of the IEEE Engineering in Medicine and Biology Society</conf-name>, <publisher-loc>San Diego, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>, pp. <fpage>2452</fpage>&#x2013;<lpage>2455</lpage>, <year>2012</year>. </mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Elgendi</surname></string-name></person-group>, &#x201C;<article-title>P and T waves annotation and detection in MIT-BIH arrhythmia database</article-title>,&#x201D; <year>2012</year>. [Online]. Available: <uri>https://vixra.org/pdf/1301.0056v1.pdf</uri>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Manu</surname></string-name>, <string-name><given-names>M. K.</given-names> <surname>Das</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Ari</surname></string-name></person-group>, &#x201C;<article-title>Automatic ECG arrhythmia classification using dual tree complex wavelet-based features</article-title>,&#x201D; <source>AEU-International Journal of Electronics and Communications</source>, vol. <volume>69</volume>, no. <issue>4</issue>, pp. <fpage>715</fpage>&#x2013;<lpage>721</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Ahmet</surname></string-name>, <string-name><given-names>N.</given-names> <surname>K&#x0131;l&#x0131;&#x00E7;</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Akan</surname></string-name></person-group>, &#x201C;<article-title>Evaluation of bagging ensemble method with time-domain feature extraction for diagnosing of arrhythmia beats</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>24</volume>, no. <issue>2</issue>, pp. <fpage>317</fpage>&#x2013;<lpage>326</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Mehrdad</surname></string-name></person-group>, &#x201C;<article-title>Classification of ECG arrhythmia by a modular neural network based on mixture of experts and negatively correlated learning</article-title>,&#x201D; <source>Biomedical Signal Processing and Control</source>, vol. <volume>8</volume>, no. <issue>3</issue>, pp. <fpage>289</fpage>&#x2013;<lpage>296</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. C.</given-names> <surname>Ping</surname></string-name></person-group>, &#x201C;<article-title>Detection of cardiac arrhythmia in electrocardiograms using adaptive feature extraction and modified support vector machines</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>39</volume>, no. <issue>9</issue>, pp. <fpage>7845</fpage>&#x2013;<lpage>7852</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Joachim</surname></string-name></person-group>, &#x201C;<article-title>ECG signal quality during arrhythmia and its application to false alarm reduction</article-title>,&#x201D; <source>IEEE Transactions on Biomedical Engineering</source>, vol. <volume>60</volume>, no. <issue>6</issue>, pp. <fpage>1660</fpage>&#x2013;<lpage>1666</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Patricia</surname></string-name></person-group>, &#x201C;<article-title>A new neural network model based on the LVQ algorithm for multi-class classification of arrhythmias</article-title>,&#x201D; <source>Information Sciences</source>, vol. <volume>279</volume>, no. <issue>7&#x2013;9</issue>, pp. <fpage>483</fpage>&#x2013;<lpage>497</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Ali</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Ozdalili</surname></string-name></person-group>, &#x201C;<article-title>Cardiac arrhythmia detection using deep learning</article-title>,&#x201D; <source>Procedia Computer Science</source>, vol. <volume>120</volume>, pp. <fpage>268</fpage>&#x2013;<lpage>275</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. R.</given-names> <surname>Joy</surname></string-name></person-group>, &#x201C;<article-title>Characterization of ECG beats from cardiac arrhythmia using discrete cosine transform in PCA framework</article-title>,&#x201D; <source>Knowledge-Based Systems</source>, vol. <volume>45</volume>, no. <issue>9765</issue>, pp. <fpage>76</fpage>&#x2013;<lpage>82</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Vasileios</surname></string-name></person-group>, &#x201C;<article-title>Effective learning and filtering of faulty heart-beats for advanced ecg arrhythmia detection using mit-bih database</article-title>,&#x201D; in <conf-name>Proc. of the 5th EAI Int. Conf. on Wireless Mobile Communication and Healthcare</conf-name>, <publisher-loc>Brussels, Belgium</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2015</year>. </mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Rashid</surname></string-name>, <string-name><given-names>G. G.</given-names> <surname>Azarnia</surname></string-name> and <string-name><given-names>M. A.</given-names> <surname>Tinati</surname></string-name></person-group>, &#x201C;<article-title>Cardiac arrhythmia classification using statistical and mixture modeling features of ECG signals</article-title>,&#x201D; <source>Pattern Recognition Letters</source>, vol. <volume>70</volume>, no. <issue>3</issue>, pp. <fpage>45</fpage>&#x2013;<lpage>51</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Serkan</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Ince</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Gabbouj</surname></string-name></person-group>, &#x201C;<article-title>Real-time patient-specific ECG classification by 1-D convolutional neural networks</article-title>,&#x201D; <source>IEEE Transactions on Biomedical Engineering</source>, vol. <volume>63</volume>, no. <issue>3</issue>, pp. <fpage>664</fpage>&#x2013;<lpage>675</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Pan</surname></string-name> and <string-name><given-names>W. J.</given-names> <surname>Tompkins</surname></string-name></person-group>, &#x201C;<article-title>A real-time QRS detection algorithm</article-title>,&#x201D; <source>IEEE Transactions on Biomedical Engineering</source>, vol. <volume>3</volume>, no. <issue>3</issue>, pp. <fpage>230</fpage>&#x2013;<lpage>236</lpage>, <year>1985</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Abramovich</surname></string-name>, <string-name><given-names>T. C.</given-names> <surname>Bailey</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Sapatinas</surname></string-name></person-group>, &#x201C;<article-title>Wavelet analysis and its statistical applications</article-title>,&#x201D; <source>Journal of the Royal Statistical Society: Series D (The Statistician)</source>, vol. <volume>49</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>29</lpage>, <year>2000</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Fr&#x00E9;d&#x00E9;rique</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Gibert</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Holschneider</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Saracco</surname></string-name></person-group>, &#x201C;<article-title>Identification of sources of potential fields with the continuous wavelet transform: Basic theory</article-title>,&#x201D; <source>Journal of Geophysical Research: Solid Earth</source>, vol. <volume>104</volume>, no. <issue>B3</issue>, pp. <fpage>5003</fpage>&#x2013;<lpage>5013</lpage>, <year>1999</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Tim</surname></string-name></person-group>, &#x201C;<article-title>Discrete wavelet transforms: Theory and implementation</article-title>,&#x201D; <source>Universidad de</source>, pp. <fpage>28</fpage>&#x2013;<lpage>35</lpage>, <year>1991</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. W.</given-names> <surname>Ivan</surname></string-name>, <string-name><given-names>R. G.</given-names> <surname>Baraniuk</surname></string-name> and <string-name><given-names>N. C.</given-names> <surname>Kingsbury</surname></string-name></person-group>, &#x201C;<article-title>The dual-tree complex wavelet transforms</article-title>,&#x201D; <source>IEEE Signal Processing Magazine</source>, vol. <volume>22</volume>, no. <issue>6</issue>, pp. <fpage>123</fpage>&#x2013;<lpage>151</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Razvan</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Gulcehre</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Cho</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Bengio</surname></string-name></person-group>, &#x201C;<article-title>How to construct deep recurrent neural networks</article-title>,&#x201D; <year>2013</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1312.6026.pdf</uri>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Sepp</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Schmidhuber</surname></string-name></person-group>, &#x201C;<article-title>Long short-term memory</article-title>,&#x201D; <source>Neural Computation</source>, vol. <volume>9</volume>, no. <issue>8</issue>, pp. <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>, <year>1997</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Klaus</surname></string-name>, <string-name><given-names>R. K.</given-names> <surname>Srivastava</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Koutn&#x00ED;k</surname></string-name>, <string-name><given-names>B. R.</given-names> <surname>Steunebrink</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Schmidhuber</surname></string-name></person-group>, &#x201C;<article-title>LSTM: A search space odyssey</article-title>,&#x201D; <source>IEEE Transactions on Neural Networks and Learning Systems</source>, vol. <volume>28</volume>, no. <issue>10</issue>, pp. <fpage>2222</fpage>&#x2013;<lpage>2232</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Junyoung</surname></string-name>, <string-name><given-names>G. C. K.</given-names> <surname>Cho</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Bengio</surname></string-name></person-group>, &#x201C;<article-title>Empirical evaluation of gated recurrent neural networks on sequence modeling</article-title>,&#x201D; <year>2014</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1412.3555.pdf</uri>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Cho</surname></string-name>, <string-name><given-names>B. V.</given-names> <surname>Merri&#x00EB;nboer</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Bahdanau</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Bengio</surname></string-name></person-group>, &#x201C;<article-title>On the properties of neural machine translation: Encoder-decoder approaches</article-title>,&#x201D; <year>2014</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1409.1259.pdf</uri>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Mike</surname></string-name> and <string-name><given-names>K. K.</given-names> <surname>Paliwal</surname></string-name></person-group>, &#x201C;<article-title>Bidirectional recurrent neural networks</article-title>,&#x201D; <source>IEEE Transactions on Signal Processing</source>, vol. <volume>45</volume>, no. <issue>11</issue>, pp. <fpage>2673</fpage>&#x2013;<lpage>2681</lpage>, <year>1997</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Alex</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Fern&#x00E1;ndez</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Schmidhuber</surname></string-name></person-group>, &#x201C;<article-title>Bidirectional LSTM networks for improved phoneme classification and recognition</article-title>,&#x201D; in <conf-name>Conf. on Artificial Neural Networks</conf-name>, <publisher-loc>Berlin, Heidelberg</publisher-loc>, <publisher-name>Springer</publisher-name>, vol. <volume>2</volume>, pp. <fpage>799</fpage>&#x2013;<lpage>804</lpage>, <year>2005</year>. </mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Rui</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Duan</surname></string-name></person-group>, &#x201C;<article-title>Bidirectional GRU for sound event detection</article-title>,&#x201D; <year>2017</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1807.00129.pdf</uri>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Dzmitry</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Cho</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Bengio</surname></string-name></person-group>, &#x201C;<article-title>Neural machine translation by jointly learning to align and translate</article-title>,&#x201D; <year>2014</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1409.0473.pdf</uri>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Jingjing</surname></string-name></person-group>, &#x201C;<article-title>Understanding and improving layer normalization</article-title>,&#x201D; <year>2019</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1911.07013.pdf</uri>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>B. J.</given-names> <surname>Lei</surname></string-name>, <string-name><given-names>J. R.</given-names> <surname>Kiros</surname></string-name> and <string-name><given-names>G. E.</given-names> <surname>Hinton</surname></string-name></person-group>, &#x201C;<article-title>Layer normalization</article-title>,&#x201D; <year>2016</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1607.06450.pdf</uri>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>L. M.</given-names> <surname>Thang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Pham</surname></string-name> and <string-name><given-names>C. D.</given-names> <surname>Manning</surname></string-name></person-group>, &#x201C;<article-title>Effective approaches to attention-based neural machine translation</article-title>,&#x201D; <year>2015</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1508.04025.pdf</uri>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Tom</surname></string-name></person-group>, &#x201C;<article-title>An introduction to ROC analysis</article-title>,&#x201D; <source>Pattern Recognition Letters</source>, vol. <volume>27</volume>, no. <issue>8</issue>, pp. <fpage>861</fpage>&#x2013;<lpage>874</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>K. N.</given-names> <surname>Shirish</surname></string-name></person-group>, &#x201C;<article-title>On large-batch training for deep learning: Generalization gap and sharp minima</article-title>,&#x201D; <year>2006</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1609.04836.pdf</uri>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>K. P.</given-names> <surname>Diederik</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Ba</surname></string-name></person-group>, &#x201C;<article-title>Adam: A method for stochastic optimization</article-title>,&#x201D; <year>2014</year>. [Online]. Available: <uri>https://arxiv.org/pdf/1412.6980.pdf</uri>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Mangathayaru</surname></string-name>, <string-name><given-names>B. P.</given-names> <surname>Rani</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Janaki</surname></string-name>, <string-name><given-names>S. M.</given-names> <surname>Gajapaka</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Patel</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>An imperative diagnostic model for predicting CHD using deep learning</article-title>,&#x201D; in <conf-name>IEEE Int. Conf. for Innovation in Technology</conf-name>, <publisher-loc>Bangluru</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Mangathayaru</surname></string-name>, <string-name><given-names>B. P.</given-names> <surname>Rani</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Janaki</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Patel</surname></string-name>, <string-name><given-names>B. L.</given-names> <surname>Bharadwaj</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>An imperative diagnostic framework for PPG signal classification using GRU</article-title>,&#x201D; in <conf-name> Advanced Informatics for Computing Research. ICAICR</conf-name>, <conf-loc>Gurugram, Haryana, India</conf-loc>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Vangipuram</surname></string-name>, <string-name><given-names>R. K.</given-names> <surname>Gunupudi</surname></string-name>, <string-name><given-names>V. K.</given-names> <surname>Puligadda</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Vinjamuri</surname></string-name></person-group>, &#x201C;<article-title>A machine learning approach for imputation and anomaly detection in IoT environment</article-title>,&#x201D; <source>Expert Systems</source>, vol. <volume>37</volume>, no. <issue>5</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Aljawarneh</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Vangipuram</surname></string-name></person-group>, &#x201C;<article-title>GARUDA: Gaussian dissimilarity measure for feature representation and anomaly detection in Internet of things</article-title>,&#x201D; <source>Journal of Super Computing</source>, vol. <volume>76</volume>, pp. <fpage>4376</fpage>&#x2013;<lpage>4413</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Aljawarneh</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Vangipuram</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Cheruvu</surname></string-name></person-group>, &#x201C;<article-title>Nirnayam: Fusion of iterative rule based decisions to build decision trees for efficient classification</article-title>,&#x201D; in <conf-name>Proc. of the 5th Int. Conf. on Engineering and MIS</conf-name>, <publisher-loc>New York, NY, USA</publisher-loc>, <publisher-name>Association for Computing Machinery</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Aljawarneh</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Radhakrishna</surname></string-name> and <string-name><given-names>G. S.</given-names> <surname>Reddy</surname></string-name></person-group>, &#x201C;<article-title>Mantra: A novel imputation measure for disease classification and prediction</article-title>,&#x201D; in <conf-name>Proc. of the First Int. Conf. on Data Science, E-learning and Information Systems</conf-name>, <publisher-loc>New York, NY, USA</publisher-loc>, <publisher-name>Association for Computing Machinery</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Radhakrishna</surname></string-name>, <string-name><given-names>P. V.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Janaki</surname></string-name></person-group>, &#x201C;<article-title>SRIHASS-a similarity measure for discovery of hidden time profiled temporal associations</article-title>,&#x201D; <source>Multimed Tools Applications</source>, vol. <volume>77</volume>, no. <issue>14</issue>, pp. <fpage>17643</fpage>&#x2013;<lpage>17692</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Radhakrishna</surname></string-name>, <string-name><given-names>P. V.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Janaki</surname></string-name></person-group>, &#x201C;<article-title>Krishna Sudarsana: A Z-space similarity measure</article-title>,&#x201D; in <conf-name>Proc. of the Fourth Int. Conf. on Engineering &#x0026; MIS, 2018</conf-name>, <publisher-loc>New York, NY, USA</publisher-loc>, <publisher-name>Association for Computing Machinery</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Radhakrishna</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Aljawarneh</surname></string-name> and <string-name><given-names>P. V.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>ASTRA-A Novel interest measure for unearthing latent temporal associations and trends through extending basic gaussian membership function</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>78</volume>, no. <issue>4</issue>, pp. <fpage>4217</fpage>&#x2013;<lpage>4265</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Radhakrishna</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Aljawarneh</surname></string-name>, <string-name><given-names>P. V.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Janaki</surname></string-name></person-group>, &#x201C;<article-title>A novel fuzzy similarity measure and prevalence estimation approach for similarity profiled temporal association pattern mining</article-title>,&#x201D; <source>Future Generation Computer Systems</source>, vol. <volume>83</volume>, no. <issue>2</issue>, pp. <fpage>582</fpage>&#x2013;<lpage>595</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G. R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mangathayaru</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Narsimha</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Cheruvu</surname></string-name></person-group>, &#x201C;<article-title>Feature clustering for anomaly detection using improved fuzzy membership function</article-title>,&#x201D; in <conf-name>Proc. of the Fourth Int. Conf. on Engineering &#x0026; MIS-2018</conf-name>, <conf-loc>Istanbul, Turkey</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G. R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mangathayaru</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Narsimha</surname></string-name> and <string-name><given-names>G. S.</given-names> <surname>Reddy</surname></string-name></person-group>, &#x201C;<article-title>Evolutionary approach for intrusion detection</article-title>,&#x201D; in <conf-name>Int. Conf. on Engineering &#x0026; MIS</conf-name>, <publisher-loc>Monastir, Tunisia</publisher-loc>, <publisher-name>IEEE</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Mangathayaru</surname></string-name>, <string-name><given-names>G. R.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Narsimha</surname></string-name></person-group>, &#x201C;<article-title>Text mining based approach for intrusion detection</article-title>,&#x201D; in <conf-name>Int. Conf. on Engineering &#x0026; MIS</conf-name>, <publisher-loc>Agadir, Morocco</publisher-loc>, <publisher-name>IEEE</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2016</year>. </mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G. R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mangathayaru</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Narsimha</surname></string-name> and <string-name><given-names>G. S.</given-names> <surname>Reddy</surname></string-name></person-group>, &#x201C;<article-title>CLAPP: A self constructing feature clustering approach for anomaly detection</article-title>,&#x201D; <source>Future Generation Computer Systems</source>, vol. <volume>74</volume>, pp. <fpage>417</fpage>&#x2013;<lpage>429</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G. R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Mangathayaru</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Narsimha</surname></string-name></person-group>, &#x201C;<article-title>An approach for intrusion detection using novel gaussian based kernel function</article-title>,&#x201D; <source>Journal of Universal Computer Science</source>, vol. <volume>22</volume>, no. <issue>4</issue>, pp. <fpage>589</fpage>&#x2013;<lpage>604</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Radhakrishna</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Aljawarneh</surname></string-name>, <string-name><given-names>P. V.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>K. R.</given-names> <surname>Choo</surname></string-name></person-group>, &#x201C;<article-title>A novel fuzzy gaussian-based dissimilarity measure for discovering similarity temporal association patterns</article-title>,&#x201D; <source>Soft Computing</source>, vol. <volume>22</volume>, no. <issue>6</issue>, pp. <fpage>1903</fpage>&#x2013;<lpage>1919</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Vangipuram</surname></string-name>, <string-name><given-names>P. V.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Janaki</surname></string-name></person-group>, &#x201C;<article-title>Krishna Sudarsana&#x2014;A Z-space interest measure for mining similarity profiled temporal association patterns</article-title>,&#x201D; <source>Foundations of Science</source>, vol. <volume>25</volume>, no. <issue>4</issue>, pp. <fpage>1027</fpage>&#x2013;<lpage>1048</lpage>, <year>2020</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>
