<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">20389</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2022.020389</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>MNN-XSS: Modular Neural Network Based Approach for XSS Attack Detection</article-title>
<alt-title alt-title-type="left-running-head">MNN-XSS: Modular Neural Network Based Approach for XSS Attack Detection</alt-title>
<alt-title alt-title-type="right-running-head">MNN-XSS: Modular Neural Network Based Approach for XSS Attack Detection</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Alqarni</surname><given-names>Ahmed Abdullah</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Alsharif</surname><given-names>Nizar</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Khan</surname><given-names>Nayeem Ahmad</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref><email>nayeem@bu.edu.sa</email>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Georgieva</surname><given-names>Lilia</given-names></name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib> 
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Pardade</surname><given-names>Eric</given-names></name>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Alzahrani</surname><given-names>Mohammed Y.</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Sciences and Information Technology, AlBaha University</institution>, <addr-line>AlBaha</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Science, Heriot-Watt University</institution>, <addr-line>Edinburgh</addr-line>, <country>UK</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Science and Information Technology, La Trobe University</institution>, <addr-line>Melbourne, VIC 3086</addr-line>, <country>Australia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Nayeem Ahmad Khan. Email: <email>nayeem@bu.edu.sa</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2021-09-13"><day>13</day><month>9</month><year>2021</year></pub-date>
<volume>70</volume>
<issue>2</issue>
<fpage>4075</fpage>
<lpage>4085</lpage>
<history>
<date date-type="received"><day>22</day><month>5</month><year>2021</year></date>
<date date-type="accepted"><day>23</day><month>6</month><year>2021</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Alqarni et al.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Alqarni et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_20389.pdf"></self-uri>
<abstract>
<p>The rapid growth and uptake of network-based communication technologies have made cybersecurity a significant challenge as the number of cyber-attacks is also increasing. A number of detection systems are used in an attempt to detect known attacks using signatures in network traffic. In recent years, researchers have used different machine learning methods to detect network attacks without relying on those signatures. The methods generally have a high false-positive rate which is not adequate for an industry-ready intrusion detection product. In this study, we propose and implement a new method that relies on a modular deep neural network for reducing the false positive rate in the XSS attack detection system. Experiments were performed using a dataset consists of 1000 malicious and 10000 benign sample. The model uses 50 features selected by using Pearson correlation method and will be used in the detection and preventions of XSS attacks. The results obtained from the experiments depict improvement in the detection accuracy as high as 99.96&#x0025; compared to other approaches.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Cybersecurity</kwd>
<kwd>XSS</kwd>
<kwd>deep learning</kwd>
<kwd>modular neural network</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>The number of web services is growing exponentially. Web applications which are accessed via web browsers have become primary targets for cybercriminals [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. A report published by Symantec Corporation in 2019 implies that 439 million pieces of new varieties of malware were identified [<xref ref-type="bibr" rid="ref-3">3</xref>], With the distinct nature and behavior among the cybercriminals and cybersecurity solution provides to defend attacks thus making it very is challenging for cybersecurity defenders to discover in what manner the new kind of malware will appear [<xref ref-type="bibr" rid="ref-4">4</xref>]. The objective of conducting such types of attacks by cybercriminals is for individual, monetary and political benefits. Therefore, early detection of such malware attacks has emerged as the uppermost cybersecurity challenge.</p>
<p>The Open Web Applications Security Project (OWASP) has declared Cross-site scripting (XSS) as one of the top vulnerabilities which are exploited to perform attacks [<xref ref-type="bibr" rid="ref-5">5</xref>]. XSS attacks are malicious script code attacks which are injected into malicious or legitimate and trusted websites [<xref ref-type="bibr" rid="ref-6">6</xref>]. These malicious scripts are delivered illegitimately to the user&#x0027;s machine so that vulnerabilities are exploited in web applications, and attacks are performed. Mostly these malicious payloads are delivered to users through email attachments or on visiting a compromised website. Cybercriminals commit a breach on the legitimate but vulnerable website to infuse the malicious script inside or develop a phishing website. Poor programming practices which such as not covering all the security aspects is one of the significant causes of vulnerability in a web application [<xref ref-type="bibr" rid="ref-7">7</xref>]. Malicious JavaScript is often employed to perform XSS attacks. JavaScript being a scripting language has several advantages which include adding versatility, dynamism, interactive into the webpages. A significant advantage of using JavaScript is that it reduces the computation load on the server-side but executing the scripts on users&#x2019; side through a web browser [<xref ref-type="bibr" rid="ref-8">8</xref>]. JavaScript offers several advantages; however, the downside is that it provides a solid foundation to conduct the XSS attack. These attacks are performed on the users&#x2019; side, which is executed by the web browsers, but there is no mechanism in web browsers for detecting any malicious scripts. Web browsers run all the scripts sent by the server, whether malicious or benign. Execution of malicious JavaScript by web browser may lead to user session hijacking, manipulating the legitimate website, for example by injecting malicious code or content or phishing attacks [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Deep learning as a new field of research is a subset of machine learning and works on by imitating forming of connections in a human brain modelled as neural networks, and also been applied to detect malware [<xref ref-type="bibr" rid="ref-10">10</xref>]. With the Neural Network&#x0027;s advancement, the problems of previous machine learning approaches in terms of accuracy in malware detection have increased. The likelihood of enhanced classification accuracy appears by developing a neural network with a higher number of prospect layers, also known as deep learning. Preliminary studies in deep learning that have been employed to detect malware in Android mobiles confirm that malware was detected with high accuracy [<xref ref-type="bibr" rid="ref-11">11</xref>]. In this study, we propose using network-based neural networks to detect malicious XSS code attacks.</p>
<p>The motivation of using the Deep Neural Network (DNN) for the detection of XSS attacks is to remove the necessity of domain expertise in feature extraction, remove complexity and solve the problem end to end. A modular neural network works on the concept of implementing multiple individual neural networks. These neural networks are trained instantaneously for a particular subtask and the results achieved are combined to perform the single task.</p>
<p>The major contribution of this study is the widespread use of the Word2vec model for the detection of malicious JavaScript&#x0027;s using MNN. Our meticulous experiments using the Word2vec model and MNN reveals a much better performance. None of the existing studies has successfully employed Word2Vec and MNN to show this degree of performance. Hence determine that our method is an effective method for detecting malicious JavaScript attacks.</p>
<p>The rest of the paper is organized as: Section 2 gives the details about related work. In Section 3 the overview of the proposed approach is detailed. In Section 4 experimental details and evaluation are presented. Section 5 concludes the work.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Work</title>
<p>Existing solutions [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>] for detection of malicious code attacks are broadly based on two approaches: signature-based and heuristic-based. In the signature-based approach, the malicious script&#x0027;s detection is performed by comparing the unique string patters in the binary code. The unique strings are created from previously captured instances of malicious code. A security solution based on signature-based needs frequently updates of their database with new signatures [<xref ref-type="bibr" rid="ref-14">14</xref>]. There is a huge time gap between finding the new malware variant and updating the signature of the malware into the database on the client-side. The attackers take benefit of such time gap to launch the attack and may affect millions of devices. Signature-based approach for malware detection fails in such an environment where new malware variants are expected to arrive.</p>
<p>Another approach used for the detection of malware is heuristic-based detection [<xref ref-type="bibr" rid="ref-15">15</xref>]. In a heuristic-based approach, the detection is performed using an expert system based on expert decision rules. Based on the set criterion, the expert system will decide whether a piece of code is suspicious or benign. The major downside of this approach is the long scanning time deciding whether a code is malicious or benign. Another challenge with this approach is that it has a high false-positive rate. To overcome the challenges in signature-based and heuristic-based detection approaches such as of high false-positive rate, long scanning time, frequent updating of signatures at the client-side, researchers use machine learning.</p>
<p>Several alternative approaches have been proposed to detect malicious JavaScript attacks using machine learning and non-machine learning methods. In this section, the machine learning and deep learning-based approaches for detecting malicious code attacks are reviewed related to our approach. As a study by [<xref ref-type="bibr" rid="ref-16">16</xref>], proposed and implemented an approach for the detection of malicious code-based N-gram, and the classification was performed using SVM. N-gram was used to generate N tokens consecutively in a stream for feature extraction. The experiments were conducted using 1831 instances of SQL injection and XSS. The experimental result shows that precision of 98.04&#x0025; was obtained with a true false positive of 0.985&#x0025; and a false positive rate of 0.015&#x0025; when used on trigram. The downside of this approach is that the tokenizers need constant training to detect malicious code. A study by [<xref ref-type="bibr" rid="ref-17">17</xref>], used machine learning classifiers such as SVM, Na&#x00EF;ve Bayes, J48 and bagging for the classification of malicious and benign code. The features extracted from the user-input context were used along with some basic features particularly related to input, output, validation, and sanitization routines. Experimental results show that an accuracy of 92.6&#x0025; was achieved using bagging. Shar et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] proposed a predication model for detection on XSS vulnerabilities based on machine learning classification and clustering techniques. Hybrid attributes extracted using static and dynamic analysis were used for code for vulnerability prediction. The experiment was performed on six applications, and results show that an average of 90&#x0025; recall and 85&#x0025; precision was obtained. The downside of the approach is that it has huge performance overheads and high false positive rate. A study by Fang et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] presented an approach for detecting XSS using deep learning. In this study the features were extracted from XSS payload. They used Long Short-Term Memory (LSTM) recurrent neural network for detection. Experimental results show that precision of 99.5&#x0025; was achieved. The downside of this approach is that it has a high time complexity. A study by Stokes et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed and implemented a deep recurrent neural for the detection of malicious JavaScript&#x0027;s. A hybrid of static and dynamic analysis was used. The presented model is highly complex, and the results produced are not convincing. Experimental results show that the applied LaMP model achieved a 65.9&#x0025; true positive rate, and the best CPoLS model obtained 45.3&#x0025; true positive rate, with 1.0&#x0025; as a false positive rate.</p>
</sec>
<sec id="s3"><label>3</label><title>Proposed Approach</title>
<sec id="s3_1"><label>3.1</label><title>Overview</title>
<p>Our new proposed approach for the detection of malicious JavaScript is based on deep and modular neural network. The proposed approach works on a self-learning method capable of detecting known and unknown variants of malware. The property of using a system which is self-learning and utilised machine learning and deep learning models that enables to extract the convoluted features from the code snippets to differentiate between the benign and malware code. The processes involved in this detection approach is shown in <?A3B2 "fig1",5,"anchor"?><xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>Architecture of the proposed approach</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_20389-fig-1.png"/></fig>
</sec>
<sec id="s3_2"><label>3.2</label><title>Pre-Processing</title>
<p>The pre-processing in this approach includes decoding, generalisation and tokenizing. The first step towards detecting malicious XSS is performing decoding on the code segment, which needs to be tested for malicious or benign. The attacker&#x0027;s use obfuscation techniques to evade the detection bypassing the traditional filters and validation mechanisms. The obfuscation or encoding is done through techniques such as Unicode, Hex encoding Base64, UTF-7 encoding. Using all the possibilities, in this proposed approach the decoder will decode data and will bring the code to a normal format. The second step in pre-processing is generalisation. This step involves removing the data noise, meaningless and non-helpful information from the decoded code, and the from the normal code. The generalisation includes removing black spaces, special characters, http://, and conversion of function parameters to param_strings. The third step is performing the tokenizing on the data. The purpose of using tokenization here is to break the sequence of strings into pieces involving input characters, sub-characters, or subgroups. Another benefit of using tokenization is that it minimizes the length of data and reduces the complexity, leading to lessens the data handling cost. In tokenization only important word remain therefore increasing the accuracy.</p>
</sec>
<sec id="s3_3"><label>3.3</label><title>Word2vec and CBOW Model</title>
<p>In this approach, we have considered each code instances as a plain text and treating it like a natural language. For this purpose, we use word2vec [<xref ref-type="bibr" rid="ref-21">21</xref>]. Word2vec algorithm employs a neural network model to learn associations of words from the large text corpus. Once the model is trained; it will help detect the synonymous words and recommend some extra words. As the name indicates, word2vec shows every distinct word with a specific list of numbers towards a vector. The vectors are selected so that a simple mathematical cosine function will depict the words represented by those vectors. Continuous bag of words (CBOW) and continuous skip gram are two models which word2vec can use to generate distributed representations of words. CBOW architecture model is usually considered much faster than skip gram for construct word representation. Keeping in view of the efficiency, In this study we employed CBOW model. <?A3B2 "fig2",5,"anchor"?><xref ref-type="fig" rid="fig-2">Fig. 2</xref> [<xref ref-type="bibr" rid="ref-22">22</xref>] shows depicts the architecture of CBOW.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Architecture of CBOW model [<xref ref-type="bibr" rid="ref-22">22</xref>]</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_20389-fig-2.png"/></fig>
</sec>
<sec id="s3_4"><label>3.4</label><title>Feature Selection</title>
<p>Feature selection is a method in which the number of the features are reduced as input variables for generating a predictive model [<xref ref-type="bibr" rid="ref-23">23</xref>]. The objective of feature selection is to reduce the computational cost and performance overheads and enhance a model&#x0027;s prediction accuracy. Suppose we have a feature vector F, we have to find the most optimal feature set F&#x2019;, keeping in mind that not all the features will contribute to the prediction model&#x0027;s accuracy.</p>
<p>Given a set of features <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>F</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>f</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>f</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mi>f</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></inline-formula> find the subset <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2286;</mml:mo><mml:mi>F</mml:mi></mml:math></inline-formula> which will maximise the learner&#x0027;s ability to classify patterns.</p>
<p>In this study we used Pearson&#x0027;s Correlation method which is a filter-based selection method [<xref ref-type="bibr" rid="ref-24">24</xref>]. Pearson&#x0027;s Correlation method is useful in determining the association among the continuous features and the class [<xref ref-type="bibr" rid="ref-25">25</xref>]. The mathematical representation is shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mo stretchy="false">(</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
</p>
<p>Only top n features are selected in our dataset by determining the absolute value of correlation among the target and numerical features. The high association was calculated by using &#x2018;CorrelationAttributeEval&#x2019; package in WEKA [<xref ref-type="bibr" rid="ref-26">26</xref>]. The selected features are given in <?A3B2 "tbl1",5,"anchor"?><xref ref-type="table" rid="table-1">Tab. 1</xref>.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Selected feature list</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Feature number</th>
<th align="left">Feature name</th>
<th align="left">Feature number</th>
<th>Feature name</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">url_length</td>
<td align="left">26</td>
<td align="left">html_attr_profile</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">url_special_characters</td>
<td align="left">27</td>
<td align="left">html_attr_http-equiv</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">url_tag_script</td>
<td align="left">28</td>
<td align="left">html_event_onblur</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">url_attr_src</td>
<td align="left">29</td>
<td align="left">html_event_onchange</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">url_event_onload</td>
<td align="left">30</td>
<td align="left">html_event_onclick</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">url_event_onmouseover</td>
<td align="left">31</td>
<td align="left">html_event_onerror</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">url_cookie</td>
<td align="left">32</td>
<td align="left">html_number_keywords_evil</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">url_number_keywords_param</td>
<td align="left">33</td>
<td align="left">js_file</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">url_number_domain</td>
<td align="left">34</td>
<td align="left">js_pseudo_protocol</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">html_tag_script</td>
<td align="left">35</td>
<td align="left">js_dom_location</td>
</tr>
<tr>
<td align="left">11</td>
<td align="left">html_tag_iframe</td>
<td align="left">36</td>
<td align="left">js_dom_document</td>
</tr>
<tr>
<td align="left">12</td>
<td align="left">html_tag_meta</td>
<td align="left">37</td>
<td align="left">js_prop_cookie</td>
</tr>
<tr>
<td align="left">13</td>
<td align="left">html_tag_object</td>
<td align="left">38</td>
<td align="left">js_method_write</td>
</tr>
<tr>
<td align="left">14</td>
<td align="left">html_tag_embed</td>
<td align="left">39</td>
<td align="left">js_method_getElementById</td>
</tr>
<tr>
<td align="left">15</td>
<td align="left">html_tag_link</td>
<td align="left">40</td>
<td align="left">js_method_alert</td>
</tr>
<tr>
<td align="left">16</td>
<td align="left">html_tag_svg</td>
<td align="left">41</td>
<td align="left">js_method_eval</td>
</tr>
<tr>
<td align="left">17</td>
<td align="left">html_tag_frame</td>
<td align="left">42</td>
<td align="left">js_method_fromCharCode</td>
</tr>
<tr>
<td align="left">18</td>
<td align="left">html_tag_div</td>
<td align="left">43</td>
<td align="left">js_min_length</td>
</tr>
<tr>
<td align="left">19</td>
<td align="left">html_tag_style</td>
<td align="left">44</td>
<td align="left">js_min_define_function</td>
</tr>
<tr>
<td align="left">20</td>
<td align="left">html_tag_img</td>
<td align="left">45</td>
<td align="left">js_min_function_calls</td>
</tr>
<tr>
<td align="left">21</td>
<td align="left">html_tag_input</td>
<td align="left">46</td>
<td align="left">js_string_max_length</td>
</tr>
<tr>
<td align="left">22</td>
<td align="left">html_attr_classid</td>
<td align="left">47</td>
<td align="left">html_length</td>
</tr>
<tr>
<td align="left">23</td>
<td align="left">html_attr_codebase</td>
<td align="left">48</td>
<td align="left">js_method_getElementsByTagName</td>
</tr>
<tr>
<td align="left">24</td>
<td align="left">html_attr_href</td>
<td align="left">49</td>
<td align="left">js_prop_referrer</td>
</tr>
<tr>
<td align="left">25</td>
<td align="left">html_attr_longdesc</td>
<td align="left">50</td>
<td align="left">html_event_onmouseup</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_5"><label>3.5</label><title>Deep and Modular Neural Network</title>
<p>The last step in this detection approach is using modular neural network (MNN) to detect malicious JavaScript&#x0027;s. MNN is considered one of the most influential and independent artificial neural networks, which is changed with only a few intermediate values [<xref ref-type="bibr" rid="ref-27">27</xref>]. MNN are neural networks which symbolize the ideas and principles of modularity. The property of using the modularity is that it can be broken down into several generally free, replicable, and composite modules. MNN reduces the computational complexity and enhances system performance and robustness [<xref ref-type="bibr" rid="ref-28">28</xref>]. The results obtained have been highly desirable than monolithic system which is based on a rigid structure. The input features are analysed by MNN, which further breaks down features into sub-features and each network is processed independently. During the process, the output generated from individual networks are consumed by the intermediary process as an input to generate the final output. The intermediary process has a characteristic of taking each process individually and perform the required action without getting distracted from other signals and doesn&#x0027;t interrelate without other networks. Basically, the strategy used by MNN to solve the problem is based on &#x201C;divide and conquer method&#x201D; [<xref ref-type="bibr" rid="ref-29">29</xref>]. MNN divides the highly complex task into a multiple subtask and each subtask is handled individually by each module. The solution produced from subtasks are combined through a unified multi module decision making strategy. Keeping in view of the advantages of using modular network, in this study optimised neural network is implemented for the detection of malicious JavaScript code attacks. <?A3B2 "fig3",5,"anchor"?><xref ref-type="fig" rid="fig-3">Fig. 3</xref> [<xref ref-type="bibr" rid="ref-30">30</xref>], depicts the basic structure of MNN and considered as a collection of monolithic neural networks that each deal with a subset of a problem and then have their separate outputs merged by an integration unit to form a comprehensive solution to the entire issue. The basic principle is that a complex problem can be broken down into simpler subsets that simpler neural networks can solve. The entire solution can be a blend of the outputs of the simple monolithic neural networks.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Basic structure of MNN [<xref ref-type="bibr" rid="ref-30">30</xref>]</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_20389-fig-3.png"/></fig>
<p>The output &#x201C;<inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>O</mml:mi></mml:math></inline-formula>&#x201D; generated by each independent network is combined to produce the final optimised output and is mathematically represented as given in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. The presence or absence of the module is known through the coefficient of the network module.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>O</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>
where,
<inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msubsup><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>=</mml:mo></mml:math></inline-formula> Module output in which <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:math></inline-formula>
<inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo></mml:math></inline-formula> Average deviation of the generated output by module
<inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo></mml:math></inline-formula> Coefficient of module <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>i</mml:mi></mml:math></inline-formula>.</p>
<p>In this proposed approach for XSS detection using MNN, each module in neural network takes as its input from the dataset. Each module in this study a 2-layer multilayer preceptor where the output generated by the second layer on the neural network is <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msubsup><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <italic>n</italic> is the maximum number of modules. In this study we are using 2 modules as given in <?A3B2 "fig4",5,"anchor"?><xref ref-type="fig" rid="fig-4">Fig. 4</xref> and module integration is done using <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. The reason of selecting only 2 modules is to simplify calculations. The output generated by modules is given as input to module integrator for making the final decision. The integrator works based on a threshold for deciding whether a code is malicious or benign. If the module integrator generated an output greater than 0.5 the code is classified as malicious, and if it is less than 0.5 the code is benign as shown in <?A3B2 "fig5",5,"anchor"?><xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The efficiency of the proposed approach is evaluated using experimental results.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>Modular neural network with 2 modules</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_20389-fig-4.png"/></fig>
<fig id="fig-5"><label>Figure 5</label><caption><title>Working of module integrator</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_20389-fig-5.png"/></fig>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Experimental Setup</title>
<sec id="s4_1"><label>4.1</label><title>Dataset</title>
<p>The dataset used in this study was obtained from figshare.com [<xref ref-type="bibr" rid="ref-31">31</xref>] developed by authors [<xref ref-type="bibr" rid="ref-32">32</xref>]. The dataset consists of 101000 instances with 1000 as malicious and 100000 benign instances. The dataset contains 67 features based on three categories viz HTML, JavaScript, and URL. The sample class is represented by [0, 1], 0 for benign and 1 for malicious benign.</p>
</sec>
<sec id="s4_2"><label>4.2</label><title>Experimental Environment &#x0026; Evaluation</title>
<p>The experiments to confirm the effectiveness of this approach was conducting on i5, 3.5 Ghz, 8 GB RAM. MATLAB platform was used to develop and experiment modular neural network. WEKA was used for the feature selection based on Pearson correlation. In first the dataset obtained was divided into 80&#x0025; training and 20&#x0025; testing data. Further 10-fold cross validation scheme was used for resampling of the data and generate predictions from unseen data. In this study only 50 features were selected based on Pearson correlation among the total of 68 features as given in the main dataset. The features selected are given in <xref ref-type="table" rid="table-1">Tab. 1</xref>. To compare the results generated from proposed modular neural network we used two approach: Backpropagation Neural Network (BPNN) and Radial Basis Function Network (RBFN) on the same data using same parameters. The only change was the number of features selected. in both approach we used 25 features each. The first 25 features were used in Backpropagation Neural Networking and second 25 features were used in and Radial Basis Function Network. The hidden layer in our experiments was set 20 and the activation function used were &#x201C;tansig&#x201D; and &#x201C;purelin&#x201D;. The training was continued till 500 epochs to achieve best accuracy result without any additional performance overheads. The confusion matric was used as for basic evaluation. Accuracy, Precision, Recall and F-1 score was measured in this study. These evaluation metrics are widely known and accepted by the research community. The formula for calculating each metrics is given in <xref ref-type="disp-formula" rid="eqn-3">Eqs. (3)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-5">(5)</xref> respectively. The results obtained are given in <?A3B2 "tbl2",5,"anchor"?><xref ref-type="table" rid="table-2">Tab. 2</xref>.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mspace width="thickmathspace" /><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
where, TP &#x003D; True Positive, TN &#x003D; True Negative, FP &#x003D; False Positive, FN &#x003D; false Negative.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Experimental results and comparison with other methods</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="3">Metrices</th>
<th align="center" colspan="6">Methods</th>
</tr>
<tr>
<th align="left" colspan="2">BPNN</th>
<th align="center" colspan="2">RBFN</th>
<th align="center" colspan="2">Proposed MNN</th>
</tr>
<tr>
<th align="left">Benign</th>
<th align="left">Malicious</th>
<th align="left">Benign</th>
<th align="center">Malicious</th>
<th align="left">Benign</th>
<th>Malicious</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Accuracy (&#x0025;)</td>
<td align="left" colspan="2">98.66</td>
<td align="left" colspan="2">97.34</td>
<td colspan="2">99.96</td>
</tr>
<tr>
<td align="left">Precision (&#x0025;)</td>
<td align="left">96.95</td>
<td align="left">100</td>
<td align="left">95</td>
<td align="left">100</td>
<td>99.95</td>
<td>99.95</td>
</tr>
<tr>
<td align="left">Recall (&#x0025;)</td>
<td align="left">100</td>
<td align="left">96.38</td>
<td align="left">100</td>
<td align="left">92.0</td>
<td>100</td>
<td>99.91</td>
</tr>
<tr>
<td align="left">F1-score (&#x0025;)</td>
<td align="left">98.32</td>
<td align="left">97</td>
<td align="left">96.45</td>
<td align="left">95.67</td>
<td>99.95</td>
<td>99.92</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Based on the insight obtained from the experimental results it is evident that our proposed MNN achieved an accuracy of 99.96&#x0025; which was high compared to BPNN and RBFN which achieved 98.66&#x0025; and 97.34&#x0025; respectively using all the features mentioned in <xref ref-type="table" rid="table-1">Tab. 1</xref>. An increase of 1&#x0025; in any detection approach means hundreds of attackers can be detected. Thus, the proposed MNN-XSS approach has proven to be highly effective in detecting XSS attacks.</p>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Conclusion</title>
<p>In this study, we propose MNN for XSS detection and conducted experiments. The experimental results show that the proposed approach achieved an accuracy of 99.96&#x0025; in detecting novel malicious JavaScript based XSS attack upon learning. The selection of number on neurons is one of the main paraments that are required to be specified the realization of the MNN. This further leads to decrease in complexity and thus allows implementation of neural network with limited resources. The Pearson correlation played an important role in feature selection thus lead to the higher accuracy.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Abdullah</surname></string-name> and <string-name><given-names>S. A.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Defending malicious script attacks using machine learning classifiers</article-title>,&#x201D; <source>Wireless Communications and Mobile Computing</source><italic>,</italic> vol. <volume>17</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Yar</surname></string-name>, and <string-name><given-names>K. F.</given-names> <surname>Steinmetz</surname></string-name></person-group>, &#x201C;<chapter-title>Cybercrime and the Internet</chapter-title>,&#x201D; in <source>Cybercrime and Society</source><italic>,</italic> <edition>3rd Edition</edition>, <publisher-loc>London</publisher-loc>, <publisher-loc>UK</publisher-loc>, <publisher-name>SAGE</publisher-name>, <year>2019</year>. [Online]. Available: <uri xlink:href="http://www.books.google.com">http://www.books.google.com</uri>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other">&#x201C;<article-title>Symantec corporation annual report-2019</article-title>,&#x201D; <year>2019</year>. [Online]. Available: <uri xlink:href="https://docs.broadcom.com/doc/istr-24-2019-en">https://docs.broadcom.com/doc/istr-24-2019-en</uri>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. J.</given-names> <surname>Holt</surname></string-name> and <string-name><given-names>M. G.</given-names> <surname>Turner</surname></string-name></person-group>, &#x201C;<article-title>Examining risks and protective factors of on-line identity theft</article-title>,&#x201D; <source>Deviant Behaviour</source><italic>,</italic> vol. <volume>33</volume>, pp. <fpage>308</fpage>&#x2013;<lpage>323</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other">&#x201C;<article-title>Open Web Application Security Project report 2020: Top 10 Web application security risks, 2020</article-title>,&#x201D; <year>2020</year>. [Online]. Available: <uri xlink:href="https://owasp.org/www-project-top-ten/">https://owasp.org/www-project-top-ten/</uri>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N. A.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Abdullah</surname></string-name> and <string-name><given-names>A. S.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>Towards vulnerability prevention model for web browser using interceptor approach</article-title>,&#x201D; in <conf-name>9th Int. Conf. on IT in Asia, (IEEE-CITA-15)</conf-name>, <conf-loc>Kuching, Malaysia</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. A.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Abdullah</surname></string-name> and <string-name><given-names>A. S.</given-names> <surname>Khan</surname></string-name></person-group>, &#x201C;<article-title>A dynamic method of detecting malicious scripts using classifiers</article-title>,&#x201D; <source>Advance Science Letters</source><italic>,</italic> vol. <volume>23</volume>, no. <issue>7</issue>, pp. <fpage>5352</fpage>&#x2013;<lpage>5355</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Laskov</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Srndic</surname></string-name></person-group>, &#x201C;<article-title>Static detection of malicious JavaScript-bearing PDF documents</article-title>,&#x201D; in <conf-name>27th Annual Computer Security Applications Conf.</conf-name>, <conf-loc>Florida, USA</conf-loc>, pp. <fpage>373</fpage>&#x2013;<lpage>382</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Likarish</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Jung</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Jo</surname></string-name></person-group>, &#x201C;<article-title>Obfuscated malicious JavaScript detection using classification techniques</article-title>,&#x201D; in <conf-name>4th IEEE Int. Conf. on Malicious and Unwanted Software (MALWARE)</conf-name>, <conf-loc>Montreal, Quebec, Canada</conf-loc>, pp. <fpage>47</fpage>&#x2013;<lpage>54</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. A.</given-names> <surname>Khan</surname></string-name></person-group>. <person-group person-group-type="author"><string-name><given-names>M. Y.</given-names> <surname>Alzahrani</surname></string-name> and <string-name><given-names>H. A.</given-names> <surname>Kar</surname></string-name></person-group>, &#x201C;<article-title>Hybrid feature classification approach for malicious javaScript attack detection using deep learning</article-title>,&#x201D; <source>International Journal of Computer Science and Information Security</source><italic>,</italic> vol. <volume>18</volume>, no. <issue>5</issue>, pp. <fpage>79</fpage>&#x2013;<lpage>90</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Singh</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Singh</surname></string-name></person-group>, &#x201C;<article-title>A survey on machine learning-based malware detection in executable files</article-title>,&#x201D; <source>Journal of Systems Architecture</source><italic>,</italic> vol. <volume>112</volume>, no. <issue>101861</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>24</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Singh</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Dixit</surname></string-name></person-group>, &#x201C;<article-title>Cross site scripting (XSS) attack detection using intrusion detection system</article-title>,&#x201D; in <conf-name>Int. Conf. on Intelligent Computing Systems</conf-name>, <conf-loc>Madurai, India</conf-loc>, pp. <fpage>199</fpage>&#x2013;<lpage>203</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Canfora</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sorbo</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Mercaldo</surname></string-name> and <string-name><given-names>C. A.</given-names> <surname>Visaggio</surname></string-name></person-group>, &#x201C;<article-title>Obfuscation techniques against signature-based detection</article-title>,&#x201D; in <conf-name>Proc. of Mobile Systems Technologies Workshop</conf-name>, <conf-loc>Milan, Italy</conf-loc>, pp. <fpage>21</fpage>&#x2013;<lpage>26</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C. M. R.</given-names> <surname>Silva</surname></string-name>, <string-name><given-names>E. L.</given-names> <surname>Feitosa</surname></string-name> and <string-name><given-names>V. C.</given-names> <surname>Garcia</surname></string-name></person-group>, &#x201C;<article-title>Heuristic-based strategy for phishing prediction: A survey of URL-based approach</article-title>,&#x201D; <source>Computers &#x0026; Security</source><italic>,</italic> vol. <volume>88</volume>, no. <issue>101613</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>19</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Han</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Deng</surname></string-name></person-group>, &#x201C;<article-title>Review on the research and practice of deep learning and reinforcement learning in smart grids</article-title>,&#x201D; <source>Journal of Power and Energy Systems</source><italic>,</italic> vol. <volume>4</volume>, no. <issue>3</issue>, pp. <fpage>362</fpage>&#x2013;<lpage>370</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Choi</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Choi</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Efficient malicious code detection using n-gram analysis and SVM</article-title>,&#x201D; in <conf-name>14th Int. Conf. on Network-Based Information Systems</conf-name>, <conf-loc>Tirana, Albania</conf-loc>, pp. <fpage>618</fpage>&#x2013;<lpage>21</lpage>, <year>2011</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M. K.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>M. C.</given-names> <surname>Govil</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Singh</surname></string-name></person-group>, &#x201C;<article-title>Predicting cross-cite scripting (XSS) security vulnerabilities in web applications</article-title>,&#x201D; in <conf-name>12th Int. Joint Conf. on Computer Science and Software Engineering</conf-name>, <conf-loc>USA</conf-loc>, pp. <fpage>162</fpage>&#x2013;<lpage>167</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L. K.</given-names> <surname>Shar</surname></string-name>, <string-name><given-names>H. B. K.</given-names> <surname>Tan</surname></string-name> and <string-name><given-names>L. C.</given-names> <surname>Briand</surname></string-name></person-group>, &#x201C;<article-title>Mining SQL injection and cross site scripting vulnerabilities using hybrid program analysis</article-title>,&#x201D; in <conf-name>35th Int. Conf. on Software Engineering</conf-name>, <conf-loc>San Francisco, USA</conf-loc>, pp. <fpage>642</fpage>&#x2013;<lpage>651</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>DeepXSS: Cross site scripting detection based on deep learning</article-title>,&#x201D; in <conf-name>Int. Conf. on Computing and Artificial Intelligence</conf-name>, <conf-loc>Chengdu, China</conf-loc>, pp. <fpage>47</fpage>&#x2013;<lpage>51</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J. W.</given-names> <surname>Stokes</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Agrawal</surname></string-name> and <string-name><given-names>G.</given-names> <surname>McDonald</surname></string-name></person-group>, &#x201C;<article-title>Neural classification of malicious scripts: A study with JavaScript and vbscript</article-title>,&#x201D; <comment>arXiv preprint arXiv: 1805.05603</comment>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. W.</given-names> <surname>Church</surname></string-name></person-group>, &#x201C;<article-title>Word2vec</article-title>,&#x201D; <source>Natural Language Engineering</source><italic>,</italic> vol. <volume>23</volume>, no. <issue>1</issue>, pp. <fpage>155</fpage>&#x2013;<lpage>162</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Bouala</surname></string-name></person-group>, &#x201C;<article-title>Word embedding models: Word2vec, Camembert and USE</article-title>,&#x201D; <source>Le Blog de Baamtu</source><italic>,</italic> <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Guyon</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Elisseeff</surname></string-name></person-group>, &#x201C;<article-title>An introduction to variable and feature selection</article-title>,&#x201D; <source>Journal of Machine Learning Research</source><italic>,</italic> vol. <volume>3</volume>, pp. <fpage>1157</fpage>&#x2013;<lpage>1182</lpage>, <year>2003</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other">&#x201C;<article-title>Spss tutorials: Pearson correlation</article-title>,&#x201D; <publisher-name>Kent University Library Tutorial</publisher-name>, <year>2021</year>. [Online]. Available: <uri xlink:href="https://libguies.lides.library.kent.edu/SPSS/PearsonCorr">https://libguies.lides.library.kent.edu/SPSS/PearsonCorr</uri>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. C.</given-names> <surname>Blessie</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Karthikeyan</surname></string-name></person-group>, &#x201C;<article-title>Sigmis: A feature selection algorithm using correlation based method</article-title>,&#x201D; <source>Journal of Algorithms &#x0026; Computational Technology</source><italic>,</italic> vol. <volume>6</volume>, no. <issue>3</issue>, pp. <fpage>385</fpage>&#x2013;<lpage>394</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other">&#x201C;<article-title>Class correlationAttributeEval</article-title>,&#x201D; <publisher-name>WEKA, University of Waikato</publisher-name>, <publisher-loc>New Zealand</publisher-loc>, <year>2020</year>. [Online]. Available: <uri xlink:href="https://weka.sourceforge.io/doc.dev/weka/attributeSelection/CorrelationAttributeEval.html">https://weka.sourceforge.io/doc.dev/weka/attributeSelection/CorrelationAttributeEval.html</uri>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Golovko</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bezobrazov</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Kachurka</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Vaitsekhovich</surname></string-name></person-group>, &#x201C;<article-title>Neural network and artificial immune systems for malware and network intrusion detection</article-title>,&#x201D; <source>Advances in Machine Learning</source><italic>,</italic> vol. <volume>2</volume>, pp. <fpage>485</fpage>&#x2013;<lpage>513</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Kourakos</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Mantoglou</surname></string-name></person-group>, &#x201C;<article-title>Pumping optimization of coastal aquifers based on evolutionary algorithms and surrogate modular neural network models</article-title>,&#x201D; <source>Advances in Water Resources</source><italic>,</italic> vol. <volume>32</volume>, no. <issue>4</issue>, pp. <fpage>507</fpage>&#x2013;<lpage>521</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>U.</given-names> <surname>Lahiri</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Pradhan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Mkhopadhyaya</surname></string-name></person-group>, &#x201C;<article-title>Modular neural network-based directional relay for transmission line protection</article-title>,&#x201D; <source>IEEE Transactions on Power Systems</source><italic>,</italic> vol. <volume>20</volume>, no. <issue>4</issue>, pp. <fpage>2154</fpage>&#x2013;<lpage>2155</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Abugabah</surname></string-name>, <string-name><given-names>A.</given-names> <surname>AlZubi</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Al-Obeidat</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alarifi</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Alwadain</surname></string-name></person-group>, &#x201C;<article-title>Data mining techniques for analyzing healthcare conditions of urban space-person lung meta-heuristic optimized neural networks</article-title>,&#x201D; <source>Cluster Computing</source><italic>,</italic> vol. <volume>23</volume>, pp. <fpage>1781</fpage>&#x2013;<lpage>1794</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F. M. M.</given-names> <surname>Mokbal</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Fu</surname></string-name></person-group>, &#x201C;<article-title>Data augmentation-based conditional wasserstein generative adversarial network-gradient penalty for XSS attack detection system</article-title>,&#x201D; <source>PeerJ Computer Science</source><italic>,</italic> vol. <volume>6</volume>, no. <issue>e328</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>20</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Makbal</surname></string-name></person-group>, &#x201C;<article-title>Cross-ite scripting attack (XSS) dataset</article-title>,&#x201D; <source>Figshare</source>. <year>2021</year>. [Online]. Available: <uri xlink:href="https://figshare.com/articles/dataset/XSS_dataset1_csv/13046138?file=24959207">https://figshare.com/articles/dataset/XSS_dataset1_csv/13046138?file&#x2009;&#x003D;&#x2009;24959207</uri>.</mixed-citation></ref>
</ref-list>
</back>
</article>