<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">57211</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.057211</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>ML-SPAs: Fortifying Healthcare Cybersecurity Leveraging Varied Machine Learning Approaches against Spear Phishing Attacks</article-title>
<alt-title alt-title-type="left-running-head">ML-SPAs: Fortifying Healthcare Cybersecurity Leveraging Varied Machine Learning Approaches against Spear Phishing Attacks</alt-title>
<alt-title alt-title-type="right-running-head">ML-SPAs: Fortifying Healthcare Cybersecurity Leveraging Varied Machine Learning Approaches against Spear Phishing Attacks</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Alanazi</surname><given-names>Saad Awadh</given-names></name><email>sanazi@ju.edu.sa</email></contrib>
<aff>
<institution>Department of Computer Science, College of Computer and Information Sciences, Jouf University</institution>, <addr-line>Sakaka, 72341, Aljouf</addr-line>, <country>Saudi Arabia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Saad Awadh Alanazi. Email: <email>sanazi@ju.edu.sa</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>19</day><month>12</month><year>2024</year>
</pub-date>
<volume>81</volume>
<issue>3</issue>
<fpage>4049</fpage>
<lpage>4080</lpage>
<history>
<date date-type="received">
<day>10</day>
<month>8</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>10</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Author.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_57211.pdf"></self-uri>
<abstract>
<p>Spear Phishing Attacks (SPAs) pose a significant threat to the healthcare sector, resulting in data breaches, financial losses, and compromised patient confidentiality. Traditional defenses, such as firewalls and antivirus software, often fail to counter these sophisticated attacks, which target human vulnerabilities. To strengthen defenses, healthcare organizations are increasingly adopting Machine Learning (ML) techniques. ML-based SPA defenses use advanced algorithms to analyze various features, including email content, sender behavior, and attachments, to detect potential threats. This capability enables proactive security measures that address risks in real-time. The interpretability of ML models fosters trust and allows security teams to continuously refine these algorithms as new attack methods emerge. Implementing ML techniques requires integrating diverse data sources, such as electronic health records, email logs, and incident reports, which enhance the algorithms&#x2019; learning environment. Feedback from end-users further improves model performance. Among tested models, the hierarchical models, Convolutional Neural Network (CNN) achieved the highest accuracy at 99.99%, followed closely by the sequential Bidirectional Long Short-Term Memory (BiLSTM) model at 99.94%. In contrast, the traditional Multi-Layer Perceptron (MLP) model showed an accuracy of 98.46%. This difference underscores the superior performance of advanced sequential and hierarchical models in detecting SPAs compared to traditional approaches.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Spear phishing attack</kwd>
<kwd>cybersecurity</kwd>
<kwd>healthcare security</kwd>
<kwd>data privacy</kwd>
<kwd>machine learning</kwd>
<kwd>sequential</kwd>
<kwd>hierarchal</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Deanship of Graduate Studies and Scientific Research at Jouf University</funding-source>
<award-id>DGSSR-2023-02-02513</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The healthcare industry&#x2019;s transformation through digital technology has significantly improved patient care and data management. However, this digital evolution also brings forth serious cybersecurity challenges, especially Spear Phishing Attacks (SPAs). These attacks are particularly sophisticated, utilizing emails that mimic legitimate sources to steal sensitive information, such as login credentials and financial details. The prevalence of these attacks is notably high in the healthcare sector due to the abundance of sensitive personal and medical information [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>]. As these threats become more common, healthcare organizations are increasingly relying on Machine Learning (ML) to detect and prevent them, showcasing the dual-edged nature of technological advancement in healthcare [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p>ML in healthcare is primarily employed to analyse large volumes of data to identify patterns that may indicate potential SPAs. Traditional ML models, however, often lack transparency in their decision-making processes, which can be a significant barrier in regulated industries like healthcare [<xref ref-type="bibr" rid="ref-6">6</xref>]. These models need to provide understandable and interpretable results to ensure that healthcare professionals can trust and effectively use the technology. Without transparency, the adoption of ML in sensitive environments would be limited, underscoring the need for models that healthcare workers can interpret and validate.</p>
<p>The application of ML extends beyond mere identification of threats. It includes developing features that can detect signs of SPAs and employing models that healthcare staff can easily understand. For example, ML can analyse email metadata, content, and user behaviour to identify unusual patterns [<xref ref-type="bibr" rid="ref-7">7</xref>]. Furthermore, using interpretable models like Decision Tree (DT) or rule-based systems enhances the transparency of the analytical process, allowing healthcare providers to trust and effectively act on the insights provided by the models [<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>Another innovative application of ML in this field is user profiling. This technique involves analysing historical data on how users interact with emails and the internet. By understanding normal behaviour patterns, ML models can flag actions that deviate from the norm, potentially identifying malicious attempts before they cause harm [<xref ref-type="bibr" rid="ref-9">9</xref>]. Additionally, integrating threat intelligence with ML models provides up-to-date information on current SPAs and malicious domains, further enhancing the model&#x2019;s accuracy and reliability [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>Lastly, the continuous learning aspect of ML is crucial for adapting to the evolving nature of cyber threats. By integrating feedback mechanisms, healthcare organizations can continuously refine their ML models. This ongoing improvement helps the models stay effective against new and changing SPAs, ensuring that the healthcare industry can maintain a robust defence against these targeted attacks [<xref ref-type="bibr" rid="ref-11">11</xref>]. Through these comprehensive strategies, ML not only strengthens cybersecurity in healthcare but also builds a foundation for future advancements in protecting sensitive patient data.</p>
<sec id="s1_1">
<label>1.1</label>
<title>Research Problem or Gap</title>
<p>SPAs have become highly sophisticated, posing a substantial threat to the healthcare sector by targeting employees to extract sensitive information [<xref ref-type="bibr" rid="ref-12">12</xref>]. Although ML presents a promising solution by learning to detect and mitigate these threats, the lack of transparency in these models raises significant concerns, particularly in a sector as sensitive as healthcare [<xref ref-type="bibr" rid="ref-13">13</xref>]. To address these challenges, several research gaps have been identified that align with the aims and objectives of enhancing healthcare defence:
<list list-type="bullet">
<list-item>
<p><bold>Development of Transparent Machine Learning Models:</bold> There is a pressing need to develop ML models that are not only accurate but also transparent, providing clear explanations for their decisions within the healthcare context.</p></list-item>
<list-item>
<p><bold>Contextual Analysis of Spear Phishing Attacks:</bold> It is critical to conduct in-depth analyses of SPAs specific to healthcare, considering the unique vulnerabilities and social engineering techniques that could be exploited.</p></list-item>
<list-item>
<p><bold>Compliance with Regulatory Requirements:</bold> Research should focus on creating ML solutions that adhere to strict privacy regulations like HIPAA, ensuring that the interpretability of these models does not compromise patient confidentiality.</p></list-item>
</list></p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Research Aims and Objectives</title>
<p>The research aims to fortify the healthcare industry&#x2019;s defence against SPAs through strategic data analysis and ML innovations. SPAs pose significant threats to healthcare data security, prompting the need for advanced detection systems that are both effective and understandable to those who operate them. The specific objectives to achieve this goal are:
<list list-type="bullet">
<list-item>
<p><bold>Dataset Collection and Pre-processing:</bold> Gather and refine a comprehensive dataset of SPAs aimed at the healthcare sector, ensuring data quality through standardization and noise reduction.</p></list-item>
<list-item>
<p><bold>Feature Selection and Engineering:</bold> Identify and engineer key features from the dataset that effectively distinguish between legitimate and SPAs within the healthcare context.</p></list-item>
<list-item>
<p><bold>Model Development and Interpretability:</bold> Develop accurate ML models that not only detect SPAs efficiently but also incorporate methodologies that enhance the interpretability of model decisions, making it easier to understand and trust the system&#x2019;s predictions.</p></list-item>
</list></p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>Research Contribution</title>
<p>This research enhances defence mechanisms against SPAs in the healthcare industry using ML to tackle complex security challenges. The key contributions are:
<list list-type="bullet">
<list-item>
<p><bold>Enhanced D&#x00E9;fense against Spear Phishing Attacks:</bold> The proposed solution significantly outperforms traditional security measures, reducing the prevalence and success of SPAs in healthcare environments.</p></list-item>
<list-item>
<p><bold>Improved Transparency and Interpretability:</bold> The ML techniques provide healthcare professionals with insights into the system&#x2019;s decision-making process, ensuring understandable and verifiable alerts and actions using diverse performance measures.</p></list-item>
</list></p>
<p>The structure of this study is carefully organized to ensure a clear understanding of the research. <xref ref-type="sec" rid="s2">Section 2</xref> presents the Related Work, establishing the background necessary for understanding the field. <xref ref-type="sec" rid="s3">Section 3</xref> describes the Methodologies used for detecting SPAs. <xref ref-type="sec" rid="s4">Section 4</xref> offers a detailed Performance Analysis, assessing how the proposed model stacks up against existing benchmarks. <xref ref-type="sec" rid="s5">Section 5</xref> discusses the significance of the proposed techniques and their advantages over previous methods. <xref ref-type="sec" rid="s6">Section 6</xref> concludes the study by summarizing the main findings and contributions, addressing any limitations, and suggesting avenues for future research.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>This literature review critically examines the advancements in SPAs detection, initially focusing on traditional ML models before delving into the application of neural network-based sequential models. The review further explores innovative techniques that enhance the detection capabilities, offering a comprehensive overview of current methodologies and their effectiveness in identifying and mitigating SPAs. This thorough analysis aims to identify gaps in current research and suggest directions for future studies to improve SPAs defence mechanisms.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Vulnerability of Healthcare to Spear Phishing Attacks</title>
<p>The healthcare sector&#x2019;s rapid digital transformation has significantly increased its vulnerability to cyber-attacks, particularly SPAs. These attacks target healthcare facilities due to the rich, sensitive data they manage, including patient records and financial information. The researchers in [<xref ref-type="bibr" rid="ref-14">14</xref>] pointed out that such data is highly valuable on the black market, making healthcare institutions prime targets. The risk is compounded by the sector&#x2019;s need to maintain continual access to critical information, making downtime caused by SPAs particularly damaging. As healthcare continues to integrate more digital technologies, the potential entry points for cyber-attacks multiply, necessitating robust defence to protect patient privacy and maintain institutional integrity. An efficient privacy-preserving model for Internet of Medical Things (IoMT) was proposed to enable secure data sharing between devices [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>Healthcare organizations are increasingly vulnerable to SPAs, which exploit personal information to create targeted, deceptive messages [<xref ref-type="bibr" rid="ref-16">16</xref>]. These SPAs pose significant threats to patient data and healthcare systems [<xref ref-type="bibr" rid="ref-17">17</xref>]. Research has identified several factors that increase susceptibility to phishing, including personality traits like conscientiousness and gender, with women being more likely to respond. The availability of personal information about targets can significantly increase vulnerability, with high information availability making users nearly three times more susceptible [<xref ref-type="bibr" rid="ref-16">16</xref>]. Healthcare professionals often have limited awareness of these threats, emphasizing the need for robust cybersecurity infrastructure and mandatory staff training. While many employees are aware of phishing risks, ongoing education across the spectrum of cybersecurity is crucial, particularly regarding information leakage on social media platforms [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Machine Learning as a D&#x00E9;fense Mechanism</title>
<p>ML offers promising solutions to enhance defence against SPAs in healthcare. According to [<xref ref-type="bibr" rid="ref-7">7</xref>], these technologies can effectively identify and mitigate threats by learning from the vast amounts of data generated in healthcare settings. ML algorithms can detect anomalies in email communications that may indicate SPAs. However, the often-opaque nature of these algorithms can be a significant barrier in environments that require high levels of trust and regulatory compliance. The inability to understand or interrogate the decision-making process of these models can hinder their acceptance and deployment in sensitive environments like healthcare.</p>
<p>ML has emerged as a powerful tool in both offensive and defensive cybersecurity strategies, particularly against SPAs. On the offensive side, ML algorithms can automate data extraction from open-source intelligence to create personalized phishing emails, achieving up to 99.69% accuracy in predicting attack success [<xref ref-type="bibr" rid="ref-19">19</xref>]. Defensively, ML techniques are employed in threat detection, malware classification, and network risk scoring [<xref ref-type="bibr" rid="ref-20">20</xref>]. To combat SPAs specifically, various ML algorithms have been evaluated, including Support Vector Machine (SVM), Logistic Regression, and Ensemble methods. The eXtreme Gradient Boosting (XGBoost) model has shown exceptional performance, achieving 99.2% accuracy in phishing detection [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Traditional Machine Learning Models in Spear Phishing Attacks Detection</title>
<p>This literature review begins by exploring the foundational role of traditional ML models in detecting SPAs. Techniques such as LR, DT, and SVM have been widely used due to their effectiveness in classifying emails based on features derived from content and metadata [<xref ref-type="bibr" rid="ref-22">22</xref>]. Research by [<xref ref-type="bibr" rid="ref-23">23</xref>] highlights how these models apply pattern recognition to differentiate malicious from benign communications effectively. However, while these methods provide a solid base, they often struggle with the dynamic nature of SPAs, which continuously evolve to bypass static filters and detection rules.</p>
<p>Traditional machine learning models have shown promising results in detecting spear phishing attacks. Various classifiers have been employed, with Na&#x00EF;ve Bayes reaching 95.15% accuracy for phishing email detection, and Random Forest (RF) attaining 96.80% accuracy for phishing website detection [<xref ref-type="bibr" rid="ref-24">24</xref>]. A combination of stylometric, forwarding, and reputation features, along with an improved SMOTE algorithm, yielded high performance in distinguishing spear phishing emails, with a maximum recall of 95.56% and precision of 98.85% [<xref ref-type="bibr" rid="ref-25">25</xref>]. Another study utilized a hybrid approach combining Na&#x00EF;ve Bayes (NB) and DT algorithms, validated against RF and LR [<xref ref-type="bibr" rid="ref-26">26</xref>]. A more recent study proposed a hybrid algorithm using SVM and LR, achieving 99.69% accuracy in predicting phishing attack success [<xref ref-type="bibr" rid="ref-19">19</xref>]. These findings demonstrate the effectiveness of ML in combating SPAs.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Advancements with Neural Network-Based Sequential Models</title>
<p>The review progresses to examine how neural network-based sequential models, like Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (BiLSTM), and Gated Recurrent Unit (GRU) offer advancements in handling the sequential nature of text data in emails. Studies such as [<xref ref-type="bibr" rid="ref-27">27</xref>] demonstrate that these models capture temporal dependencies and nuances in email communication that traditional models might overlook. This ability makes them particularly suited for detecting sophisticated SPAs that employ subtle cues and context manipulation.</p>
<p>Recent advancements in neural network-based sequential models have significantly improved the detection of spear phishing attacks via email. Reference [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed a dynamic evolving neural network using reinforcement learning, achieving high accuracy 98.63% and adaptability to new phishing behaviours. Reference [<xref ref-type="bibr" rid="ref-29">29</xref>] developed a model that learns character and word embeddings directly from email texts, attaining 99.81% accuracy on common datasets. Reference [<xref ref-type="bibr" rid="ref-30">30</xref>] demonstrated the effectiveness of neural networks for phishing email detection and classification. To address personalized filtering, Reference [<xref ref-type="bibr" rid="ref-31">31</xref>] introduced a Stackelberg game model for calculating optimal thresholds in sequential attack scenarios, outperforming existing approaches.</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Leveraging Hierarchal Models for Enhanced Detection</title>
<p>Further, the review assesses the integration of hierarchal techniques, which have significantly improved SPAs detection&#x2019;s accuracy and adaptability. For instance, Convolutional Neural Network (CNN) have been adapted for text classification by extracting spatial hierarchies of features from textual data, as discussed by [<xref ref-type="bibr" rid="ref-32">32</xref>]. These models are noted for their ability to discern complex patterns in data, offering a more nuanced understanding of the content, which is crucial for identifying highly targeted SPAs.</p>
<p>Recent research has focused on leveraging CNN for enhanced phishing email detection. Studies have demonstrated the effectiveness of CNNs in analysing email text content, achieving high accuracy, precision, and recall rates [<xref ref-type="bibr" rid="ref-33">33</xref>]. CNNs have shown promise in extracting meaningful features from email headers, text, and attachments, enabling the detection of both known and emerging phishing attacks. Further improvements have been achieved by augmenting one-dimensional CNN models with recurrent layers such as LSTM, Bi-LSTM, GRU, and Bi-GRU [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
</sec>
<sec id="s2_6">
<label>2.6</label>
<title>Adapting to Evolving Threats</title>
<p>Adaptability is another critical aspect of ML in combating SPAs. The studies [<xref ref-type="bibr" rid="ref-35">35</xref>&#x2013;<xref ref-type="bibr" rid="ref-37">37</xref>] highlighted the importance of ML systems being capable of evolving in response to new and emerging SPAs. As attackers continuously refine their strategies, ML models must also adapt to identify and counteract these evolving threats effectively. This continuous learning approach helps maintain the relevance and efficacy of cybersecurity measures in a landscape where threat vectors swiftly change.</p>
</sec>
<sec id="s2_7">
<label>2.7</label>
<title>Integration into Healthcare Cybersecurity Protocols</title>
<p>The integration of ML into healthcare cybersecurity protocols offers a proactive approach to managing SPAs. Another research [<xref ref-type="bibr" rid="ref-38">38</xref>] emphasized the importance of not just reacting to threats as they occur but anticipating and preventing them through advanced threat detection systems. These systems, powered by ML, can significantly enhance the security posture of healthcare organizations by providing timely and accurate detection of SPAs, thereby reducing the risk of data breaches and ensuring the protection of sensitive patient information.</p>
</sec>
<sec id="s2_8">
<label>2.8</label>
<title>Challenges and Future Directions</title>
<p>Despite these advancements, the review identifies key challenges that persist in the field. The primary concern is the black-box nature of many advanced ML models, which limits their interpretability&#x2014;a critical aspect in healthcare and other sensitive sectors, where understanding decision-making processes is vital. As per the findings of [<xref ref-type="bibr" rid="ref-39">39</xref>], there is a growing need for models that not only predict accurately but also provide insights into their predictions to ensure trust and compliance with regulatory standards.</p>
<p>In conclusion, this literature review synthesizes current research on ML&#x2019;s role in SPAs detection, highlighting significant progress and outlining persistent challenges. The evolution from traditional models to more sophisticated neural network approaches marks a substantial advancement in the field. However, as SPAs become more refined, future research must focus on developing models that balance predictive power with transparency and adaptability to maintain efficacy in an ever-evolving threat landscape.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>This section details the methodology used to enhance SPAs detection in healthcare through the application of interpretable ML techniques. The methodology is structured to address the key research objectives and fill identified gaps, focusing on dataset collection, feature engineering, model development, and testing into healthcare systems as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed model for Machine Learning-Spear Phishing Attacks (ML-SPAs) detection</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset Collection and Pre-processing</title>
<p>This study compiled a comprehensive dataset of email communications, both legitimate and malicious, specifically targeting healthcare organizations. This dataset was curated from various sources and augmented with simulated SPAs. Pre-processing involved cleaning the data, standardizing email formats, and anonymizing sensitive information to comply with privacy regulations.</p>
<p>For the mathematical modelling of the dataset collection and pre-processing phase in this study, considers the following equations to formally represent the data handling processes:</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Dataset Collection</title>
<p>The entire dataset <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi mathvariant="bold-italic">D</mml:mi></mml:math></inline-formula> is represented as the union of all individual email communications <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi mathvariant="bold-italic">E</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> collected from various sources in a dataset as shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x22C3;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Dataset Augmentation</title>
<p>The augmented dataset <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msup><mml:mi mathvariant="bold-italic">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> includes the original dataset <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi mathvariant="bold-italic">D</mml:mi></mml:math></inline-formula> combined with additional simulated SPAs <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi mathvariant="bold-italic">A</mml:mi></mml:math></inline-formula> shown in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi>D</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>A</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Dataset Cleaning and Anonymization Process</title>
<p>The pre-processed dataset <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi mathvariant="bold-italic">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2033;</mml:mi></mml:mrow></mml:math></inline-formula> is obtained by applying a cleaning and anonymization function <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi mathvariant="bold-italic">f</mml:mi></mml:math></inline-formula> to the augmented dataset <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msup><mml:mi mathvariant="bold-italic">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi>D</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>D</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s3_1_4">
<label>3.1.4</label>
<title>Feature Standardization</title>
<p>The standardized feature <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is calculated by subtracting the mean <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi></mml:math></inline-formula> from the original feature <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> and dividing by the standard deviation <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi mathvariant="bold-italic">&#x03C3;</mml:mi></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi></mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>These equations collectively model the transformation of raw email data into a structured, standardized, and secure format suitable for further analysis with ML techniques. This mathematical framework ensures clarity in the methodological steps involved in preparing this dataset for effective SPAs detection.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Feature Selection and Engineering</title>
<p>Features were carefully selected and engineered to capture the nuances of SPAs. This included analysing email metadata, the linguistic style of the body text, and patterns of user interaction with previous emails. Advanced natural language processing techniques were employed to extract and quantify these features, which were then scaled and normalized to prepare them for model training.</p>
<p>For this phase study can model the process mathematically on SPAs detection using ML by defining equations that encapsulate feature extraction, scaling, and normalization. These processes are crucial for preparing data inputs for efficient and effective ML model training.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Feature Extraction</title>
<p>Extract features <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>F</mml:mi></mml:math></inline-formula> from the raw data <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>D</mml:mi></mml:math></inline-formula> using Natural Language Processing (NLP) techniques and metadata analysis to capture the nuances of SPAs shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>F</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>E</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>E</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>D</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>D</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the feature set for the <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>i</mml:mi></mml:math></inline-formula>-th email, and extract is a function that applies NLP and metadata analysis to extract features such as linguistic style, email metadata, and interaction patterns.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Feature Quantification</title>
<p>Quantify the extracted features <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>F</mml:mi></mml:math></inline-formula> into numerical values <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>Q</mml:mi></mml:math></inline-formula> to make them suitable for ML algorithms as shown in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>Q</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>F</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>E</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>F</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the quantified features for the <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>i</mml:mi></mml:math></inline-formula>-th email, and quantify is a function that converts linguistic and categorical descriptors into numerical or categorical values, often using techniques like tokenization, vectorization, and encoding.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Feature Scaling</title>
<p>Scale the quantified features <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>Q</mml:mi></mml:math></inline-formula> to have a uniform range, typically <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> or <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, 1], using a scaling function <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>S</mml:mi></mml:math></inline-formula> as shown in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>M</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>M</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mi>F</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>E</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>Q</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>F</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the scaled features for the <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>i</mml:mi></mml:math></inline-formula>-th email. This step ensures that all input features contribute equally to the model training, preventing any single feature with a large range from dominating the learning process.</p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>Feature Normalization</title>
<p>Normalize the scaled features <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>S</mml:mi></mml:math></inline-formula> to ensure they follow a standard format, facilitating smoother and more stable convergence during model training as shown in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mi>F</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>E</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>S</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>F</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the normalized features for the <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>i</mml:mi></mml:math></inline-formula>-th email, <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>&#x03BC;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the mean of the scaled features, and <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is their standard deviation. Normalization is particularly important when features have different units or variances.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Model Development and Interpretability</title>
<p>Several traditional ML models (Na&#x00EF;ve Bayes (NB), Logistic Regression (LR), Gradient Boosting Machines (GBM), XGBoost, Decision Tree (DT), Random Forest (RF), MLP, and AdaBoost), sequential models (RNN, LSTM, BiLSTM, and GRU), and hierarchal model (CNN) were deployed and trained. Each ML model was evaluated for its accuracy, precision, recall, and F1-score and sequential and hierarchal models were evaluated with Epochs, Time, Loss, Accuracy, Validation Loss, Validation Accuracy. Special emphasis was placed on interpretation of models to provide insights into the decision-making processes of the models.</p>
<p>In this phase study, formalized the approach and evaluation criteria for the various ML and sequential and hierarchal models using mathematical equations. These equations can encapsulate the training, evaluation, and interpretability of the models.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Model Training</title>
<p>The training of various ML models is represented mathematically by <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>T</mml:mi><mml:mi>y</mml:mi><mml:mi>p</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the model trained using the <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>i</mml:mi></mml:math></inline-formula>-th algorithm (such as NB, LR, GBM, XGBoost, DT, RF, MLP, and AdaBoost), <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>X</mml:mi></mml:math></inline-formula> is the feature set, and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>Y</mml:mi></mml:math></inline-formula> is the target variable.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Model Evaluation for Machine Learning Models</title>
<p>The evaluation metrics for ML models are defined as follows in <xref ref-type="disp-formula" rid="eqn-10">Eqs. (10)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-13">(13)</xref>:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x2211;</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mi>T</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>S</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x2211;</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x2211;</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>S</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Model Evaluation for Sequential and Hierarchal Models</title>
<p>For sequential and hierarchal models, the following metrics shown in <xref ref-type="disp-formula" rid="eqn-14">Eqs. (14)</xref> and <xref ref-type="disp-formula" rid="eqn-15">(15)</xref> are utilized to evaluate the models&#x2019; training and validation performance:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>V</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>V</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>V</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Continuous Learning and Adaptation</title>
<p>The deployed models were equipped with mechanisms for continuous learning, allowing them to adapt to new and evolving SPAs. Feedback loops were established to refine the models based on the latest threat intelligence and real-world detection outcomes.</p>
<p>Now SPAs detection using ML can be mathematically modelled using the adaptation processes that allow deployed models to dynamically learn and evolve over time. This process involves updating the models periodically with new data, refining their parameters based on feedback, and ensuring they remain effective against the latest SPAs.</p>
<sec id="s3_4_1">
<label>3.4.1</label>
<title>Model Update</title>
<p>The model shown in <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref> is updated iteratively based on new data and feedback to adapt to new and evolving SPAs:</p>
<p><disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the model at time <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is new data incorporating recent SPAs, and <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is feedback derived from the real-world application of the model.</p>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>Feedback Loop Incorporation</title>
<p>Feedback from the model&#x2019;s performance is processed to refine and improve its parameters as shown in <xref ref-type="disp-formula" rid="eqn-17">Eq. (17)</xref>:</p>
<p><disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents observed outcomes such as detection accuracy and false positives, and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are real-world detection outcomes that may highlight new SPAs.</p>
</sec>
<sec id="s3_4_3">
<label>3.4.3</label>
<title>Continuous Learning</title>
<p>Adjustments to the model&#x2019;s parameters or structure are made using a continuous learning function shown in <xref ref-type="disp-formula" rid="eqn-18">Eq. (18)</xref>, enhancing predictive accuracy over time:</p>
<p><disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the parameters at time <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>t</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the updated parameters for time <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, modified based on the feedback <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_4_4">
<label>3.4.4</label>
<title>Adaptation to New Threats</title>
<p>A decision rule evaluates whether the model needs further adaptation to address the dynamically changing threat landscape as shown in <xref ref-type="disp-formula" rid="eqn-19">Eq. (19)</xref>:</p>
<p><disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mi>A</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the changes in the threat landscape, and <italic>Adapt</italic> is a boolean value determining whether adaptation is necessary to maintain the model&#x2019;s effectiveness.</p>
<p>This methodology ensures a robust, adaptable, and transparent approach to SPAs detection, leveraging the latest advancements in ML while addressing the unique challenges faced by the healthcare industry. The whole process is summarized concisely in Algorithm 1:</p>
<fig id="fig-9">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-9.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Performance Evaluation</title>
<p>This research evaluates the effectiveness of proposed models in detecting SPAs, using a rigorous testing strategy that compares them against state-of-the-art models. The primary data source for this evaluation is the publicly available &#x201C;Phishing Email Detection&#x201D; dataset from Kaggle, specifically chosen for its relevance to SPAs detection tasks. This choice ensures that the testing environment is both broad and challenging, accurately reflecting the real-world complexities involved in processing various email types and contexts.</p>
<p>In this research, the proposed model undergoes rigorous testing through systematic comparisons against established benchmarks and cutting-edge developments in neural network models. The evaluations extend beyond basic ML metrics like precision, recall, F1-score, and support scores, delving into the models&#x2019; capabilities to assimilate and interpret linguistic and thematic nuances across various email categories. These categories, such as 0 and 1, help in assessing the accuracy, macro average, and weighted average of each model. This comprehensive approach allows for a detailed understanding of how each model processes and responds to the complex data characteristics typical of SPAs. Such in-depth analysis ensures the proposed model&#x2019;s efficacy and adaptability in real-world scenarios, providing a robust framework for advancing cybersecurity measures.</p>
<p>For sequential and hierarchal models, performance metrics include Epochs, Time, Loss, Accuracy, Validation Loss, and Validation Accuracy, further measuring how effectively these models can assimilate and analyse information. This comprehensive analysis extends to different email categories, providing insights into the models&#x2019; operational effectiveness under varied conditions. Moreover, this study involves a detailed examination of the models&#x2019; configurations, including layers, output shapes, parameters, total parameters, trainable parameters, and non-trainable parameters for DL models. This review highlights the bespoke nature of the proposed models, tailored specifically to tackle SPAs amidst contemporary advanced threats. The meticulous evaluation framework employed ensures that the findings are robust, offering definitive evidence of the proposed models&#x2019; superiority over existing methods. This approach not only confirms the models&#x2019; efficacy in real-world applications but also showcases their innovative use of advanced neural network techniques and adaptability to complex cybersecurity challenges in the healthcare sector.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset Description and Selection Rationale</title>
<p>The dataset titled &#x201C;Phishing Email Detection&#x201D; on Kaggle [<xref ref-type="bibr" rid="ref-40">40</xref>] provides a robust framework for training ML models to detect SPAs, a significant threat leading to data breaches and financial losses in various sectors, including healthcare. This dataset comprises two main features: &#x2018;Email Text&#x2019; and &#x2018;Email Type.&#x2019; The &#x2018;Email Text&#x2019; contains the body of the emails, which is crucial for identifying and analysing the linguistic and thematic elements associated with SPAs. &#x2018;Email Type&#x2019; classifies each email as either &#x2018;Phishing&#x2019; or &#x2018;Safe,&#x2019; facilitating the training of models through supervised learning techniques.</p>
<p>The dataset statistics reveal that it includes a total of 18,600 email entries, with 3% of the emails marked as &#x2018;empty&#x2019; under the &#x2018;Email Text&#x2019; feature, indicating missing content. This aspect underscores the importance of robust pre-processing steps to handle missing or incomplete data effectively. Moreover, the distribution between &#x2018;Safe&#x2019; and &#x2018;Phishing&#x2019; emails shows that 61% of the emails are safe, while 39% are SPAs. This balance provides a realistic scenario for training detection models, reflecting the frequent exposure to both legitimate and malicious emails in real-world settings.</p>
<p>The use of this dataset enables the deployment of various ML techniques, including NB, LR, and more complex models like RF and Neural Networks, to discern and predict SPAs accurately. By leveraging such a detailed and representative dataset, researchers and cybersecurity professionals can enhance detection algorithms, improving their ability to safeguard sensitive information against sophisticated cyber threats. The relevant statistics for the benchmark dataset are outlined in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Identified dataset&#x2019;s statistics</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Feature</th>
<th>Description</th>
<th>Statistics</th>
</tr>
</thead>
<tbody>
<tr>
<td></td>
<td>Number of entries</td>
<td>18,600</td>
</tr>
<tr>
<td>Email text</td>
<td>Contains the body of the email</td>
<td>Empty: 3% <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mrow><mml:mtext>Br</mml:mtext></mml:mrow></mml:mrow><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula> Valid: 97%</td>
</tr>
<tr>
<td></td>
<td></td>
<td>Unique entries: 17,510</td>
</tr>
<tr>
<td>Email type</td>
<td>Indicates if the email is &#x2018;Phishing&#x2019; or &#x2018;Safe&#x2019;</td>
<td>Phishing email: 39% <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mrow><mml:mtext>Br</mml:mtext></mml:mrow></mml:mrow><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula> Safe email: 61%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The dataset summary highlights key aspects crucial for effective ML model development. Specifically, the &#x2018;Email Text&#x2019; field reveals that 3% of the entries are empty, necessitating pre-processing steps such as filling in or removing these entries to ensure optimal model training. Meanwhile, the &#x2018;Email Type&#x2019; field shows a balanced distribution between safe and phishing emails, providing an ideal setup for training classification models. This balance is essential for accurately learning the distinguishing features that differentiate phishing from non-phishing emails, enhancing the effectiveness of the predictive models.</p>
<p>The dataset was specifically selected for its relevance and challenge in SPAs detection, as well as its established use in prior research, which facilitates direct comparisons with state-of-the-art models. This methodical evaluation framework ensures that the performance of the proposed model is benchmarked against the highest industry standards. It emphasizes the model&#x2019;s innovative capabilities and effectiveness in accurately identifying and responding to SPAs, highlighting its potential to advance current cybersecurity measures in this critical area.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>System Configuration and Implementation Settings</title>
<p>The proposed system configuration employs a high-performance Lenovo Mobile Workstation, ideal for processing complex datasets involved in SPAs detection. It features a 12th Generation Intel Core i9 processor, 128 GB of DDR4 memory, a 4 TB solid-state drive (SSD), and an NVIDIA RTX A4090 graphics card. Running on Windows 11 with Python version 3.12.0, this setup enables rigorous testing of the proposed model across various challenging datasets, ensuring robust analysis and performance benchmarking.</p>
<p>For data pre-processing, the system implements routines to clean the dataset by dropping duplicates and null values, essential for maintaining data integrity and model accuracy. The dataset predominantly comprises emails classified as &#x2018;safe&#x2019; over those tagged as &#x2018;Phishing.&#x2019; This imbalance informs a pre-processing strategy where suspected SPAs, potentially mislabelled or underrepresented, are carefully scrutinized and, if necessary, removed to enhance the training process. This approach helps in minimizing noise and focusing the model&#x2019;s learning on truly representative features of SPAs.</p>
<p>This sophisticated system configuration combined with meticulous pre-processing practices supports advanced ML operations. It ensures that the models developed are not only accurate but also capable of handling real-world data efficiently. This setup underscores a commitment to leveraging cutting-edge technology and methodological rigor to enhance the detection capabilities against SPAs, thereby significantly boosting cybersecurity measures in sensitive environments like healthcare and finance.</p>
<p>The pie chart titled &#x201C;Categorical Distribution&#x201D; provides a visual breakdown of email categories within a SPAs dataset. The chart shows that 62.6% of the emails are classified as &#x201C;Safe Email&#x201D;, represented by the blue segment, while 37.4% are labelled as &#x201C;Phishing Email&#x201D;, depicted in orange. This visualization highlights the distribution of emails, emphasizing a higher prevalence of &#x201C;Safe Emails&#x201D; compared to &#x201C;Phishing Emails&#x201D;. This proportionate representation in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> is crucial for understanding the dataset&#x2019;s composition and assists in evaluating the effectiveness of ML models trained on this data. The chart effectively communicates the ratio of Safe to Phishing emails, providing essential insights for further analysis and model training.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Proportional distribution of email categories in phishing detection dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-2.tif"/>
</fig>
<p>In the realm of text pre-processing for ML, particularly for natural language processing tasks like SPAs detection, several crucial steps are undertaken to refine the dataset for optimal model performance. The first step, integer encoding, involves converting textual data into numerical format so that ML algorithms can process the information. This is essential because models inherently understand numbers, not text.</p>
<p>Further pre-processing includes the removal of hyperlinks, punctuations, and extra spaces from the email texts. This cleansing process helps in reducing the noise within the data, ensuring that the algorithms focus solely on meaningful content. Hyperlinks and punctuation can introduce biases or irrelevant features that might mislead the model, while extra spaces could affect the structure of the data fed into the model.</p>
<p>Another vital pre-processing step is the creation of a word cloud of available stop words. Stop words are common words like &#x201C;and&#x201D;, &#x201C;the&#x201D;, &#x201C;a&#x201D;, which typically do not contain important significance and are removed from the text. Visualizing these stop words in a word cloud can help in understanding their frequency and distribution within the dataset. Removing these stop words further cleans the data, allowing the focus to remain on the crucial elements of the texts that contribute more significantly to the understanding and detection of SPAs. Together, these pre-processing steps refine the dataset, preparing it for effective and efficient analysis and classification by ML models.</p>
<p>This word cloud visually represents the most frequently occurring words in a dataset of email communications. The prominence of words like &#x201C;the&#x201D;, &#x201C;of&#x201D;, &#x201C;to&#x201D;, &#x201C;and&#x201D;, &#x201C;in&#x201D;, &#x201C;for&#x201D;, &#x201C;you&#x201D;, &#x201C;with&#x201D;, and &#x201C;be&#x201D; indicates their high usage in typical email texts. Such word clouds are instrumental in identifying common stop words&#x2014;words that are usually filtered out before processing text data due to their minimal contribution to the overall meaning. This graphic illustrates not only the commonality of these words but also emphasizes the linguistic patterns that could be crucial for tasks like sentiment analysis, topic modelling, or spam detection where understanding text content is essential. By analysing this distribution in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, researchers can better tailor their algorithms to focus on more meaningful and less frequent terms that might indicate specific behaviours or intentions within the emails.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Analysis of common words in word cloud of email communications</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-3.tif"/>
</fig>
<p>This word cloud shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> visualizes the most frequently occurring words in a dataset of email communications, providing insights into common language usage within emails. Words like &#x201C;time&#x201D;, &#x201C;information&#x201D;, &#x201C;people&#x201D;, &#x201C;use&#x201D;, &#x201C;work&#x201D;, &#x201C;need&#x201D;, and &#x201C;make&#x201D; are prominently displayed, highlighting their prevalence in daily email interactions. The size of each word indicates its frequency, with larger words appearing more often in the dataset. This visualization helps identify key themes and terms that are typically used in emails, which can be crucial for tasks such as email categorization, sentiment analysis, and spam detection. By understanding these patterns, organizations can better tailor their communication strategies and improve email filtering algorithms.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Linguistic patterns analysis in email communication using word cloud of unique words</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-4.tif"/>
</fig>
<p>Pre-processing text data for ML applications typically involves converting text into a numerical format that algorithms can interpret. This process starts with using techniques like the Term Frequency-Inverse Document Frequency (TF-IDF) vectorizer. The TF-IDF vectorizer quantifies the importance of a word in a document relative to a collection of documents or corpus, thereby transforming text into a set of vectors. This vectorization reflects how important a term is in the context of a document, which is pivotal for models to understand textual data.</p>
<p>Once text data is converted into vectors, it is essential to split the dataset into training and testing sets. This division allows the model to learn patterns from the training set and then validate its performance on the unseen test set. This method helps in assessing the generalizability of the model when faced with new data, which is crucial for its deployment in real-world applications.</p>
<p>After splitting the data, various algorithms can be applied to train on these vectors. Each algorithm, whether it be a traditional ML model like NB or LR, or more complex models like RF or Neural Networks, has its strengths and weaknesses in processing and learning from textual data. By applying different algorithms, one can evaluate which model performs best for the specific task of text classification or prediction, ensuring the most effective approach is used in practical scenarios.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Results and Discussion</title>
<p>This section of the study thoroughly analyzes the performance of various ML models on the &#x201C;Phishing Email Detection&#x201D; dataset, particularly for detecting SPAs. The evaluation is segmented into four parts. The initial segment assesses traditional baseline models including NB, LR, GBM, XGBoost, DT, RF, MLP, and AdaBoost. Subsequently, the study explores the efficacy of advanced sequential models like RNN, LSTM, BiLSTM, GRU, and CNN, which are specifically tailored for enhanced detection of advanced SPAs. This structured analysis helps in comparing the strengths and weaknesses of each model in a realistic setting.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Overview of Spear Phishing Attacks Detection Model Performance</title>
<p>The overview provided in <xref ref-type="table" rid="table-2">Tables 2</xref> and <xref ref-type="table" rid="table-3">3</xref> of the study details the performance metrics of various ML models applied to the &#x201C;Phishing Email Detection&#x201D; dataset for SPAs detection. These metrics include precision, recall, F1-score, and support scores, all crucial for evaluating the efficacy of the models in differentiating between SPAs and legitimate emails. Additionally, the study examines the DL models&#x2019; performance in terms of Epochs, Time, Loss, Accuracy, Validation Loss, and Validation Accuracy, offering a deeper insight into how well these models process and analyse data under varied conditions.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of performance metrices for identified machine learning models</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Performance measures</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-score</th>
<th>Support</th>
</tr>
</thead>
<tbody>
<tr>
<td>Na&#x00EF;ve Bayes</td>
<td>0</td>
<td>0.9723</td>
<td>0.9630</td>
<td>0.9676</td>
<td>1351</td>
</tr>
<tr>
<td/>
<td>1</td>
<td>0.9770</td>
<td>0.9828</td>
<td>0.9799</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9752</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9747</td>
<td>0.9729</td>
<td>0.9738</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9752</td>
<td>0.9752</td>
<td>0.9752</td>
<td>3508</td>
</tr>
<tr>
<td>Logistic regression</td>
<td>0</td>
<td>0.9826</td>
<td>0.9637</td>
<td>0.9731</td>
<td>1351</td>
</tr>
<tr>
<td/>
<td>1</td>
<td>0.9776</td>
<td>0.9893</td>
<td>0.9834</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9795</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9801</td>
<td>0.9765</td>
<td>0.9783</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9795</td>
<td>0.9795</td>
<td>0.9794</td>
<td>3508</td>
</tr>
<tr>
<td>Gradient boosting</td>
<td>0</td>
<td>0.9829</td>
<td>0.9793</td>
<td>0.9811</td>
<td>1351</td>
</tr>
<tr>
<td>machines</td>
<td>1</td>
<td>0.9870</td>
<td>0.9893</td>
<td>0.9882</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9823</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9850</td>
<td>0.9843</td>
<td>0.9846</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9823</td>
<td>0.9823</td>
<td>0.9823</td>
<td>3508</td>
</tr>
<tr>
<td>XGBoost</td>
<td>0</td>
<td>0.9610</td>
<td>0.9667</td>
<td>0.9638</td>
<td>1351</td>
</tr>
<tr>
<td/>
<td>1</td>
<td>0.9791</td>
<td>0.9754</td>
<td>0.9772</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9721</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9700</td>
<td>0.9711</td>
<td>0.9705</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9721</td>
<td>0.9721</td>
<td>0.9721</td>
<td>3508</td>
</tr>
<tr>
<td>Decision tree</td>
<td>0</td>
<td>0.9610</td>
<td>0.9667</td>
<td>0.9638</td>
<td>1351</td>
</tr>
<tr>
<td/>
<td>1</td>
<td>0.9791</td>
<td>0.9754</td>
<td>0.9772</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9721</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9700</td>
<td>0.9711</td>
<td>0.9705</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9721</td>
<td>0.9721</td>
<td>0.9721</td>
<td>3508</td>
</tr>
<tr>
<td>Random forest</td>
<td>0</td>
<td>0.9650</td>
<td>0.9800</td>
<td>0.9725</td>
<td>1351</td>
</tr>
<tr>
<td/>
<td>1</td>
<td>0.9874</td>
<td>0.9777</td>
<td>0.9825</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9786</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9762</td>
<td>0.9789</td>
<td>0.9775</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9788</td>
<td>0.9786</td>
<td>0.9787</td>
<td>3508</td>
</tr>
<tr>
<td>Multi-layer</td>
<td>0</td>
<td>0.9807</td>
<td>0.9793</td>
<td>0.9800</td>
<td>1351</td>
</tr>
<tr>
<td>perceptron</td>
<td>1</td>
<td>0.9870</td>
<td>0.9879</td>
<td>0.9875</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9846</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9839</td>
<td>0.9836</td>
<td>0.9837</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9846</td>
<td>0.9846</td>
<td>0.9846</td>
<td>3508</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>0</td>
<td>0.9489</td>
<td>0.8113</td>
<td>0.8747</td>
<td>1351</td>
</tr>
<tr>
<td/>
<td>1</td>
<td>0.8916</td>
<td>0.9726</td>
<td>0.9304</td>
<td>2157</td>
</tr>
<tr>
<td/>
<td>Accuracy</td>
<td></td>
<td></td>
<td>0.9105</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Macro avg</td>
<td>0.9203</td>
<td>0.8919</td>
<td>0.9025</td>
<td>3508</td>
</tr>
<tr>
<td/>
<td>Weighted avg</td>
<td>0.9137</td>
<td>0.9105</td>
<td>0.9089</td>
<td>3508</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Sequential models&#x2019; comparisons in terms of performance measures for spear phishing attacks detection</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Epochs</th>
<th>Time (s)</th>
<th>Loss</th>
<th>Accuracy</th>
<th>Validation loss</th>
<th>Validation accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td>Recurrent neural network</td>
<td>7</td>
<td>241</td>
<td>0.4051</td>
<td>0.7676</td>
<td>0.5325</td>
<td>0.7147</td>
</tr>
<tr>
<td>Long short-term memory</td>
<td>5</td>
<td>284</td>
<td>0.1278</td>
<td>0.9617</td>
<td>0.1307</td>
<td>0.9629</td>
</tr>
<tr>
<td>Bidirectional long short-term memory</td>
<td>10</td>
<td>156</td>
<td>0.0020</td>
<td>0.9994</td>
<td>0.0937</td>
<td>0.9781</td>
</tr>
<tr>
<td>Gated recurrent unit</td>
<td>10</td>
<td>123</td>
<td>0.0571</td>
<td>0.9868</td>
<td>0.2191</td>
<td>0.9369</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Furthermore, the study provides a detailed analysis of the models&#x2019; configurations, focusing on aspects such as layers, output shapes, parameters, total parameters, trainable parameters, and non-trainable parameters. This level of detail helps in understanding the structural and operational nuances of each model, particularly the DL models, and their adaptability to the complex requirements of SPAs detection. This comprehensive evaluation not only benchmarks the models against each other but also highlights their strengths and limitations in practical scenarios.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Performance Insights on Spear Phishing Email Detection Dataset Using Machine Learning</title>
<p>The detailed performance evaluation shown in <xref ref-type="table" rid="table-2">Table 2</xref> from the &#x201C;Phishing Email Detection&#x201D; dataset reveals significant insights into various ML models&#x2019; effectiveness in SPAs detection. The analysis categorizes the performance across two main types of email: &#x2018;0&#x2019; and &#x2018;1&#x2019;, utilizing metrics like precision, recall, F1-score, and support for detailed assessment.</p>

<p>For category &#x2018;1&#x2019;, the MLP exhibits standout performance with the highest precision of 0.9870 and an F1-score of 0.9875, demonstrating its capability in handling complex patterns associated with SPAs effectively. This makes MLP the best-performing model in this category, optimized for accuracy and reliability. Conversely, the AdaBoost model registers lower metrics, with a precision of 0.8916 and an F1-score of 0.9304, highlighting areas for improvement and making it the least effective model for this category.</p>
<p>In category &#x2018;0&#x2019;, the GBM shows superior performance with impressive scores&#x2014;particularly a precision of 0.9829 and an F1-score of 0.9811&#x2014;indicating its high accuracy and balanced detection capability. On the other end, the AdaBoost, with a lower precision of 0.9489 and an F1-score of 0.8747, falls short compared to other models, underscoring the need for further tuning.</p>
<p>This granular evaluation not only identifies the strengths and weaknesses of each model but also provides a clear benchmarking framework that can guide future improvements and selections in ML deployments for SPAs detection.</p>
<p>The provided <xref ref-type="fig" rid="fig-5">Fig. 5</xref> showcases the confusion matrices for various ML models applied to the &#x201C;Phishing Email Detection&#x201D; dataset. Each matrix represents the performance of models including NB, LR, GBM, XGBoost, DT, RF, MLP, and AdaBoost in classifying emails as either &#x2018;Phishing&#x2019; or &#x2018;Safe&#x2019;.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Performance comparison using confusion matrices for identified machine learning models</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-5.tif"/>
</fig>
<p>In these matrices, the <italic>x</italic>-axis typically represents the predicted categories and the <italic>y</italic>-axis the actual categories. The values within the matrix show the count of predictions falling into each category (True Positives, True Negatives, False Positives, False Negatives). These metrics are crucial for assessing the effectiveness of each model in correctly identifying and categorizing emails, providing insights into both their strengths and weaknesses.</p>
<p>From the arrangement and results depicted, it can be discerned that some models may exhibit higher false positives or false negatives, which could significantly impact their usability in real-world scenarios. For example, a model with a high number of false positives might frequently misclassify safe emails as phishing, leading to unnecessary alerts. Conversely, models with low false positives and negatives would be ideal as they maintain a balance, reducing the risk of overlooking actual phishing attempts and not overburdening the system with false alerts.</p>
<p>By analysing these matrices, stakeholders can make informed decisions on which models might require further tuning or optimization and which models are performing well under the current testing conditions. This form of evaluation is essential for continuous improvement in SPAs detection systems.</p>
<p>The confusion matrices provided for various ML models clearly show their performance in classifying emails within the &#x201C;Phishing Email Detection&#x201D; dataset. Notably, the AdaBoost model exhibited one of the weakest performances with 1096 true positives and 59 false negatives for SPAs, but significantly, 255 false positives and 2098 true negatives for safe emails, indicating a higher misclassification rate of safe emails as SPAs. Conversely, one of the best performances was observed in the model represented in the MLP confusion matrix, which achieved 1321 true positives and only 26 false negatives for phishing emails, along with 30 false positives and 2131 true negatives for safe emails, demonstrating a high accuracy and a balanced approach in detecting both categories effectively. These insights are crucial for refining the models to enhance their detection capabilities.</p>
<p>The <xref ref-type="fig" rid="fig-6">Fig. 6</xref> shows a bar chart titled &#x201C;Performance of the models&#x201D;, which visually represents the accuracy of various ML models applied to the &#x201C;Phishing Email Detection&#x201D; dataset. The chart displays the accuracy rates for each model, including NB, LR, GBM, XGBoost, DT, RF, MLP Classifier, and AdaBoost. Each model&#x2019;s performance is indicated by a green bar, with the height of the bar corresponding to its accuracy percentage. Notably, the MLP Classifier shows the highest accuracy at 98.29%, while AdaBoost shows the lowest at 91.05%. This graphical representation provides a clear and immediate comparison of the effectiveness of each model in detecting SPAs.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparative accuracy analysis of machine learning models in spear phishing attacks detection</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-6.tif"/>
</fig>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Performance Insights on Spear Phishing Email Detection Dataset Using Sequential and Hierarchal Models</title>
<p>The performance insights from the sequential and hierarchal models as shown in <xref ref-type="table" rid="table-3">Tables 3</xref> and <xref ref-type="table" rid="table-4">4</xref> on the SPAs detection dataset reveal significant variations in efficacy across different types of neural networks. The dataset results, presented in a comparative table, include measurements across Epochs, Training Time, Loss, Accuracy, Validation Loss, and Validation Accuracy.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Hierarchal models&#x2019; comparisons in terms of performance measures for spear phishing attacks detection</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Epochs</th>
<th>Time (s)</th>
<th>Loss</th>
<th>Accuracy</th>
<th>Validation loss</th>
<th>Validation accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td>Convolutional neural network</td>
<td>10</td>
<td>115</td>
<td>0.0001</td>
<td>0.9999</td>
<td>0.0880</td>
<td>0.9795</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The LSTM network shows excellent performance with an accuracy of 96.17% and a remarkable Validation Accuracy of 96.29%, coupled with low loss values (0.1278 training and 0.1307 validation), making it one of the best performers for this dataset. On the other hand, the basic RNN trails with notably lower performance metrics&#x2014;76.76% accuracy and 71.47% Validation Accuracy, alongside higher loss values (0.4051 training and 0.5325 validation), indicating its relative inadequacy in handling this specific task.</p>
<p>Other models like the BiLSTM, GRU, and CNN also showcase strong performances with accuracies exceeding 98%. Specifically, the BiLSTM model achieves nearly perfect accuracy at 99.94% with a Validation Accuracy of 97.81%, and the CNN impresses with a Validation Accuracy of 97.95%.</p>
<p>These insights highlight the effectiveness of advanced neural network architectures in accurately detecting SPAs, with LSTM, BiLSTM, and CNN models providing robust solutions thanks to their ability to efficiently process and learn from complex data patterns inherent in email communications.</p>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Comprehensive Evaluation and Implications</title>
<p>The analysis in <xref ref-type="fig" rid="fig-7">Figs. 7</xref> and <xref ref-type="fig" rid="fig-8">8</xref> encompasses a range of sequential and hierarchal models tailored for SPAs detection, illustrating their performance through detailed graphs and confusion matrices. The models evaluated include RNN, LSTM, BiLSTM, GRU, and CNN. Each model&#x2019;s performance is depicted through training accuracy trends, loss measurements, and confusion matrices that display true positives, false positives, false negatives, and true negatives.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Sequential models&#x2019; comparisons for spear phishing attacks detection using diverse graphs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Hierarchal models&#x2019; comparisons for spear phishing attacks detection using diverse graphs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57211-fig-8.tif"/>
</fig>
<p>Starting with the RNN, it shows a steady improvement in training accuracy, reaching around 72%, yet it struggles with relatively high numbers of false negatives and positives, indicating a need for further refinement in terms of model precision and recall. The LSTM model, however, excels in its training, with accuracy soaring to 99%. It also records low training and Validation Loss, demonstrating robust learning capabilities. Its confusion matrix confirms its effectiveness, showing a significant number of true positives and very few false negatives, marking it as a top performer.</p>
<p>The BiLSTM maintains stable and high training and Validation Accuracy, with minimized losses after initial fluctuations. Its confusion matrix indicates a strong true positive rate, with very few false negatives, underscoring its efficiency. Similarly, the GRU model showcases high accuracy, above 98%, and low losses, with a confusion matrix that highlights its ability to accurately identify SPAs.</p>
<p>The CNN emerges as the most accurate model, achieving nearly perfect training accuracy and maintaining low loss levels. This is mirrored in its high Validation Accuracy and its confusion matrix, which shows an impressive count of true positives with minimal false negatives. This model&#x2019;s performance suggests it has the best capability in the line-up for accurately detecting SPAs.</p>
<p>The overview highlights the distinct capabilities of each model, with the CNN and LSTM models showing particularly high effectiveness in accurately detecting SPAs. This is characterized by their high true positive rates and low false negatives. In contrast, the RNN, despite being useful, shows lagging performance due to higher misclassifications. These insights are vital for ongoing adjustments and optimization, aiming to enhance the reliability and efficiency of SPAs detection systems in cybersecurity applications. This comparative analysis underscores the importance of selecting the right model architecture to address the specific challenges of SPAs detection.</p>
<p>The <xref ref-type="table" rid="table-5">Table 5</xref> details the architecture of various sequential and hierarchal models used in SPAs detection, emphasizing the configuration and complexity of each. The table includes models like RNN, LSTM, BiLSTM LSTM, GRU, and CNN, each outlined with layer types, output shapes, parameter counts, and total parameters, highlighting the scale and intricacies of their design.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Sequential and hierarchal models&#x2019; comparisons for spear phishing attacks detection in terms of resources consumption</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Layer (Type)</th>
<th>Output shape</th>
<th>Number of parameters in<break/> each layer</th>
<th>Number of parameters in each model</th>
</tr>
</thead>
<tbody>
<tr>
<td>Recurrent neural network</td>
<td>Embedding</td>
<td>(None, 150, 50)</td>
<td>9,115,850</td>
<td>Total params: 9,131,051</td>
</tr>
<tr>
<td/>
<td>RNN</td>
<td>(None, 100)</td>
<td>15,100</td>
<td>Trainable params: 9,131,051</td>
</tr>
<tr>
<td/>
<td>Dropout</td>
<td>(None, 100)</td>
<td>0</td>
<td>Non-trainable params: 0</td>
</tr>
<tr>
<td/>
<td>Dense</td>
<td>(None, 1)</td>
<td>101</td>
<td></td>
</tr>
<tr>
<td>Long short-term memory</td>
<td>Embedding</td>
<td>(None, 150, 50)</td>
<td>9,115,850</td>
<td>Total params: 9,176,351</td>
</tr>
<tr>
<td></td>
<td>LSTM</td>
<td>(None, 100)</td>
<td>60,400</td>
<td>Trainable params: 9,176,351</td>
</tr>
<tr>
<td/>
<td>Dropout</td>
<td>(None, 100)</td>
<td>0</td>
<td>Non-trainable params: 0</td>
</tr>
<tr>
<td/>
<td>Dense</td>
<td>(None, 1)</td>
<td>101</td>
<td></td>
</tr>
<tr>
<td>Bidirectional long short-term memory</td>
<td>Embedding</td>
<td>(None, 150, 50)</td>
<td>9,115,850</td>
<td>Total params: 9,236,851</td>
</tr>
<tr>
<td></td>
<td>Bidirectional</td>
<td>(None, 200)</td>
<td>120,800</td>
<td>Trainable params: 9,236,851</td>
</tr>
<tr>
<td/>
<td>Dropout</td>
<td>(None, 200)</td>
<td>0</td>
<td>Non-trainable params: 0</td>
</tr>
<tr>
<td/>
<td>Dense</td>
<td>(None, 1)</td>
<td>201</td>
<td></td>
</tr>
<tr>
<td>Gated recurrent unit</td>
<td>Embedding</td>
<td>(None, 150, 50)</td>
<td>9,115,850</td>
<td>Total params: 9,161,551</td>
</tr>
<tr>
<td></td>
<td>GRU</td>
<td>(None, 100)</td>
<td>45,600</td>
<td>Trainable params: 9,161,551</td>
</tr>
<tr>
<td/>
<td>Dropout</td>
<td>(None, 100)</td>
<td>0</td>
<td>Non-trainable params: 0</td>
</tr>
<tr>
<td/>
<td>Dense</td>
<td>(None, 1)</td>
<td>101</td>
<td></td>
</tr>
<tr>
<td>Convolutional neural network</td>
<td>Embedding</td>
<td>(None, 150, 50)</td>
<td>9,115,850</td>
<td>Total params: 9,133,963</td>
</tr>
<tr>
<td></td>
<td>Conv1D</td>
<td>(None, 148, 64)</td>
<td>9664</td>
<td>Trainable params: 9,133,963</td>
</tr>
<tr>
<td/>
<td>Global max pooling one dimensional</td>
<td>(None, 64)</td>
<td>0</td>
<td>Non-trainable params: 0</td>
</tr>
<tr>
<td/>
<td>Dense</td>
<td>(None, 128)</td>
<td>8320</td>
<td></td>
</tr>
<tr>
<td/>
<td>Dropout</td>
<td>(None, 128)</td>
<td>0</td>
<td></td>
</tr>
<tr>
<td/>
<td>Dense</td>
<td>(None, 1)</td>
<td>129</td>
<td></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The RNN is relatively straightforward with fewer trainable parameters compared to others, indicating a simpler model structure that might impact its ability to capture complex patterns in data. In contrast, the LSTM model includes multiple layers such as embedding, LSTM layers, and dense layers, culminating in a significant number of trainable parameters which enhance its ability to learn from large and complex datasets effectively.</p>
<p>The BiLSTM doubles the parameter count of the LSTM by processing data in both forward and backward directions, offering a more nuanced understanding of input sequences. This is advantageous for tasks like email classification where contextual relationships in text can be pivotal.</p>
<p>The GRU model simplifies the gating mechanisms found in LSTMs while still maintaining a considerable number of parameters, allowing it to perform efficiently with less computational overhead. Meanwhile, the CNN uses one dimensional convolutional layer to capture spatial dependencies and patterns in data, which can be crucial for identifying textual features in email data. The detail configurations of one dimensional CNN are given below those produces optimized results:</p>
<p><bold>Input Layer: Embedding</bold>
<list list-type="bullet">
<list-item>
<p><bold>Input Dimension:</bold> The input to the network is a sequence of integers (tokens), with each sequence having a maximum length of 150 (defined by max_len).</p></list-item>
<list-item>
<p><bold>Output Dimension:</bold> The embedding layer transforms each token into a 50-dimensional vector. Hence, the output dimension from the Embedding layer is (150, 50) for each sample, where 150 is the sequence length and 50 is the embedding size.</p></list-item>
<list-item>
<p><bold>Parameters:</bold> The number of parameters in the Embedding layer is the product of the number of unique tokens (len(tk.word_index) &#x002B; 1) and the output dimension (50).</p></list-item>
</list></p>
<p><bold>Convolutional Layer: Conv1D</bold>
<list list-type="bullet">
<list-item>
<p><bold>Input Dimension:</bold> Accepts the output from the Embedding layer, which is (150, 50).</p></list-item>
<list-item>
<p><bold>Kernel Size:</bold> The convolution operates using a kernel of size 3. This means it looks at 3 consecutive elements in the input data at a time.</p></list-item>
<list-item>
<p><bold>Filters:</bold> The layer uses 64 filters, meaning it will produce 64 different feature maps.</p></list-item>
<list-item>
<p><bold>Output Dimension:</bold> Each filter produces an output of size 148 (assuming &#x2018;valid&#x2019; padding where no padding is applied). Thus, the output dimension of this layer is (148, 64).</p></list-item>
<list-item>
<p><bold>Parameters:</bold> Each filter has parameters for each element in the kernel for each input channel (depth). Here, each filter has 3 &#x00D7; 50 parameters, and there are 64 such filters, resulting in 3 &#x00D7; 50 &#x00D7; 64 &#x003D; 9600 parameters.</p></list-item>
</list></p>
<p><bold>Pooling Layer: GlobalMaxPooling1D</bold>
<list list-type="bullet">
<list-item>
<p><bold>Input Dimension:</bold> Accepts the output from the Conv1D layer, which is (148, 64).</p></list-item>
<list-item>
<p><bold>Output Dimension:</bold> This layer performs global max pooling over the entire length of each feature map, reducing the dimension to just the number of feature maps, i.e., (64).</p></list-item>
</list></p>
<p><bold>Dense Layer</bold>
<list list-type="bullet">
<list-item>
<p><bold>Input Dimension:</bold> Accepts the output from the GlobalMaxPooling1D layer, which is (64).</p></list-item>
<list-item>
<p><bold>Units:</bold> 128 neurons in this layer.</p></list-item>
<list-item>
<p><bold>Output Dimension:</bold> (128).</p></list-item>
<list-item>
<p><bold>Parameters</bold><bold>:</bold> Each neuron in this layer is connected to every input. Hence, the total parameters are 64 &#x00D7; 128 &#x002B; 128 (for biases) &#x003D; 8320.</p></list-item>
</list></p>
<p><bold>Dropout Layer</bold>
<list list-type="bullet">
<list-item>
<p><bold>Purpose:</bold> Randomly sets input units to 0 at each step during training time, which helps to prevent overfitting. Dropout rate is 0.5.</p></list-item>
<list-item>
<p><bold>Input/Output Dimension:</bold> Does not alter the dimension, so it remains (128).</p></list-item>
</list></p>
<p><bold>Output Dense Layer</bold>
<list list-type="bullet">
<list-item>
<p><bold>Input Dimension:</bold> (128).</p></list-item>
<list-item>
<p><bold>Units:</bold> 1 neuron (for binary classification).</p></list-item>
<list-item>
<p><bold>Activation:</bold> sigmoid to output probabilities.</p></list-item>
<list-item>
<p><bold>Output Dimension:</bold> (1).</p></list-item>
<list-item>
<p><bold>Parameters</bold><bold>:</bold> 128 &#x00D7; 1 &#x002B; 1 (for bias) &#x003D; 129.</p></list-item>
</list></p>
<p>These models&#x2019; configurations suggest a deliberate design choice to optimize learning capabilities and computational efficiency. The detailed breakdown of each model&#x2019;s architecture provides insights into how sequential and hierarchal models can be effectively applied to the problem of SPAs detection, leveraging complex structures to achieve high accuracy and robust performance in real-world applications. Each model&#x2019;s setup is tailored to balance between depth of learning and operational demands, ensuring they can adapt and respond to the evolving tactics employed in SPAs.</p>
</sec>
<sec id="s4_3_5">
<label>4.3.5</label>
<title>Highlighting the Innovation and Contribution of the Proposed Model</title>
<p>The CNN model represents a significant innovation in the field of SPAs detection. This model stands out due to its complex structure that includes multiple layers such as convolutional layers, pooling layers, and dense layers, each contributing to a highly refined processing capability. The convolutional layers effectively capture spatial and temporal dependencies in email text data, allowing for the detection of nuanced patterns that simpler models might miss. The inclusion of global max pooling and multiple dense layers further enhances the model&#x2019;s ability to consolidate learned features into precise predictions. This sophisticated architecture not only improves the accuracy of SPAs detection but also showcases the model&#x2019;s ability to handle large-scale data, adapting to new threats as they evolve. This makes the CNN model a pivotal development in cybersecurity measures against SPAs, highlighting its potential to significantly reduce the risk of email-based security breaches.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion</title>
<p><xref ref-type="table" rid="table-6">Table 6</xref> provides an overview of comparative analysis of various ML models for detecting SPAs reveals significant performance disparities among the classifiers. The MLP stands out with the highest accuracy at 98.29%, proving exceptionally capable in complex phishing scenarios. In contrast, AdaBoost, despite its robustness, shows the lowest accuracy at 91.05%. This study aligns with findings from recent research, such as Tesfom et al. [<xref ref-type="bibr" rid="ref-41">41</xref>], where NB markedly underperformed with a 66.0% accuracy rate. Meanwhile, LR and XGBoost demonstrated strong capabilities with accuracies of 96.3% and 97.27%, respectively. Notably, RF topped other models with a 97.98% accuracy, underscoring its effectiveness in phishing detection. This variance underscores the critical importance of selecting the right model based on the specific requirements and complexities of phishing detection tasks.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparison with traditional machine learning models</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>References</th>
<th>Model name</th>
<th>Description</th>
<th>Comparative performance</th>
</tr>
</thead>
<tbody>
<tr>
<td>Tesfom et al., 2023 [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>Na&#x00EF;ve Bayes</td>
<td>A NB model typically used for classification, reported to have the lowest accuracy of 66.0% in phishing detection among the models tested.</td>
<td>Significantly underperforms in phishing detection.</td>
</tr>
<tr>
<td>Mittal et al., 2023 [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>Logistic regression</td>
<td>LR demonstrated superior accuracy, achieving 96.3% in detecting phishing websites, highlighting its effectiveness.</td>
<td>Highly accurate, with superior performance in phishing detection.</td>
</tr>
<tr>
<td>Abdul Samad et al., 2023 [<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
<td>Gradient boosting machine</td>
<td>GBM can achieve high accuracy (over 97%) in detecting phishing URLs when fine-tuned with data balancing, hyperparameter optimization, and feature selection.</td>
<td>Data balancing leads to minor improvements in performance, while hyperparameter tuning and feature selection significantly improve accuracy.</td>
</tr>
<tr>
<td>Musa et al., 2019 [<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>XGBoost</td>
<td>XGBoost achieved high accuracy (97.27%) in phishing detection, outperforming other models like PNN and RF.</td>
<td>Excellent performance with top-tier accuracy in phishing detection.</td>
</tr>
<tr>
<td>Fazal et al., 2023 [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>Decision tree</td>
<td>DT model reported an accuracy of 95.97% in detecting phishing websites, showcasing high efficacy.</td>
<td>Highly effective, with strong performance in detecting phishing.</td>
</tr>
<tr>
<td>Ab Razak et al., 2022 [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>Random forest</td>
<td>RF achieved the highest reported accuracy among classifiers at 97.98% in phishing detection.</td>
<td>Top performer with the highest accuracy in phishing detection.</td>
</tr>
<tr>
<td>Subasi et al., 2020 [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>AdaBoost</td>
<td>AdaBoost combined with SVM achieved an accuracy of 97.61%, making it highly effective in phishing website detection.</td>
<td>Outstanding performance, one of the highest accuracies in phishing detection.</td>
</tr>
<tr>
<td>Akinwale et al., 2022 [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>Logistic regression and decision tree</td>
<td>Hybrid ML approach using LR and DTs classifiers can detect spear-phishing emails with 99.8% accuracy.</td>
<td>It presents a hybrid ML approach to detect and classify spear-phishing emails in organizations with high accuracy.</td>
</tr>
<tr>
<td>Hegde et al., 2023 [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Support vector machines and logistic regression</td>
<td>Hybrid algorithm combining SVM and LR to predict the success rate of phishing attacks, achieving a peak accuracy of 99.69%.</td>
<td>To increase the effectiveness of phishing attacks by automating the data extraction process and analyzing the success rate of attacks using ML before launching them.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-7">Table 7</xref> presents the analysis of various sequential and hierarchical models for SPAs detection demonstrates considerable variability in performance, highlighting the specialized capabilities of each model. The BiLSTM model excels, achieving a nearly flawless accuracy of 99.94% and a validation accuracy of 97.81%, showcasing its profound efficiency in handling dynamic and complex phishing scenarios. Similarly, the CNN model also performs impressively, registering a validation accuracy of 97.95%. This analysis underscores the effectiveness of sophisticated neural architectures in SPA detection, with both BiLSTM and CNN providing highly robust solutions. Additional models like RNN, GRU, and the hybrid CNN-LSTM further reinforce the potential of hierarchal models in enhancing cybersecurity measures, with their respective high Accuracy rates demonstrating strong suitability for SPAs detection tasks, as evidenced in the performance metrics reported across recent studies.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparison with sequential and hierarchal machine learning models</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>References</th>
<th>Model name</th>
<th>Description</th>
<th>Comparative performance</th>
</tr>
</thead>
<tbody>
<tr>
<td>Bahnsen et al.,<break/>2017 [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>Recurrent neural network</td>
<td>RNNs used to classify phishing URLs, demonstrating a high accuracy rate of 98.7%, surpassing random forest methods.</td>
<td>Top performance in URL classification with 98.7% accuracy.</td>
</tr>
<tr>
<td>Adebowale et al.,<break/>2019 [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>Convolutional neural network, and long short-term memory</td>
<td>A hybrid model combining CNN and LSTM for phishing detection, achieving an accuracy of 93.28%.</td>
<td>Effective for complex phishing detection with good accuracy.</td>
</tr>
<tr>
<td>Roy et al., 2022 [<xref ref-type="bibr" rid="ref-51">51</xref>]</td>
<td>Long short-term memory, Bi-LSTM, and gated recurrent unit</td>
<td>Utilizes LSTM, Bi-LSTM, and GRU models for phishing URL detection, reaching up to 99% accuracy.</td>
<td>Highest reported accuracy among LSTM variants, excellent at 99%.</td>
</tr>
<tr>
<td>Jafar et al., 2022 [<xref ref-type="bibr" rid="ref-52">52</xref>]</td>
<td>Gated recurrent unit</td>
<td>GRU model specifically aimed at detecting phishing URLs with 98.30% accuracy, outperforming other classifiers.</td>
<td>Highly effective in URL detection, nearly perfect accuracy.</td>
</tr>
<tr>
<td>McGinley et al.,<break/>2021 [<xref ref-type="bibr" rid="ref-53">53</xref>]</td>
<td>Convolutional neural network</td>
<td>CNN optimized for phishing email classification, achieving 98% accuracy, recall, and precision.</td>
<td>Superior performance in phishing email classification.</td>
</tr>
<tr>
<td>Hasan et al., [<xref ref-type="bibr" rid="ref-54">54</xref>]</td>
<td>Deep convolutional neural network</td>
<td>Developed a Deep Convolutional Neural Network (DCNN) model that can accurately classify phishing websites from legitimate websites, achieving an overall accuracy of 99%.</td>
<td>Other ML algorithms were not as effective as the DCNN model in classifying phishing websites, likely due to the limited dataset.</td>
</tr>
<tr>
<td>Proposed ML-SPAs</td>
<td>Convolutional neural network, BiLSTM</td>
<td>Achieved accuracy of 99.99% and 99.94%, respectively.</td>
<td>Superior performance in SPAs detection and classification.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The comparisons drawn with both single and multi-document models, including HDSG, GRETEL, and SgSum, underscore proposed model&#x2019;s enhanced capabilities in both thematic depth and structural coherence. The proposed model&#x2019;s architecture leverages advanced neural network techniques to dynamically adapt to the intricacies of the text, setting new standards in extractive summarization. This place proposed research at the forefront, pioneering next-generation summarization solutions that effectively address both the granularity of content and the coherence of summaries.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion, Limitations and Future Work</title>
<p>This study has demonstrated the effectiveness of various traditional, sequential, and hierarchical ML models in detecting SPAs with high accuracy. Among the models evaluated, BiLSTM and CNN exhibited outstanding performance, achieving near-perfect accuracy rates 99.99%. The application of these advanced neural network architectures substantially improves the detection and mitigation of phishing threats in healthcare environments, proving vital for protecting sensitive healthcare data from sophisticated cyber-attacks. This highlights the importance of adopting advanced ML techniques to enhance cybersecurity in critical sectors like healthcare.</p>
<p>Despite the promising results, this study has limitations. The primary constraint is the dependency on large and diverse datasets for training the models, which might not be readily available or could be biased towards specific types of phishing attacks.</p>
<p>Future research will focus on addressing this limitation by exploring methods to improve model transparency. Efforts will also be made to augment datasets with more varied and complex phishing scenarios to ensure robustness across different attack vectors. Furthermore, integrating these models into real-time detection systems and assessing their performance in live environments will be crucial to validate their practical applicability and efficiency in real-world settings.</p>
</sec>
</body>
<back>
<ack>
<p>Thanks to my institute who supported me throughout this work.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was funded by the Deanship of Graduate Studies and Scientific Research at Jouf University under Grant Number (DGSSR-2023-02-02513).</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The dataset used in this study is publicly available.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The author declares no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. S.</given-names> <surname>Bhuyan</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Transforming healthcare cybersecurity from reactive to proactive: Current status and future recommendations</article-title>,&#x201D; <source>J. Med. Syst.</source>, vol. <volume>44</volume>, no. <issue>5</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1007/s10916-019-1507-y</pub-id>; <pub-id pub-id-type="pmid">32239357</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Bera</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Ogbanufe</surname></string-name>, and <string-name><given-names>D. J.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Towards a thematic dimensional framework of online fraud: An exploration of fraudulent email attack tactics and intentions</article-title>,&#x201D; <source>Decis. Support Syst.</source>, vol. <volume>171</volume>, <year>2023</year>, Art. no. <comment>113977</comment>. doi: <pub-id pub-id-type="doi">10.1016/j.dss.2023.113977</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Ahmad</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Kanta</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Shiaeles</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Naeem</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Khalid</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Mahboob</surname></string-name></person-group>, &#x201C;<article-title>Enhancing ATM security management in the post-quantum era with quantum key distribution</article-title>,&#x201D; in <conf-name>2024 IEEE Int. Conf. Cyber Secur. Resil.</conf-name>, <year>2024</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. I.</given-names> <surname>Malik</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Ibrahim</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Hannay</surname></string-name>, and <string-name><given-names>L. F.</given-names> <surname>Sikos</surname></string-name></person-group>, &#x201C;<article-title>Developing resilient cyber-physical systems: A review of state-of-the-art malware detection approaches, gaps, and future directions</article-title>,&#x201D; <source>Computers</source>, vol. <volume>12</volume>, no. <issue>4</issue>, <year>2023</year>, Art. no. <comment>79</comment>. doi: <pub-id pub-id-type="doi">10.3390/computers12040079</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Shahzadi</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Machine learning empowered security management and quality of service provision in SDN-NFV environment</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>66</volume>, no. <issue>3</issue>, pp. <fpage>2723</fpage>&#x2013;<lpage>2749</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2021.014594</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. M. Mohi</given-names> <surname>Uddin</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Biswas</surname></string-name>, <string-name><given-names>S. T.</given-names> <surname>Rikta</surname></string-name>, <string-name><given-names>S. K.</given-names> <surname>Dey</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Qazi</surname></string-name></person-group>, &#x201C;<article-title>XML-LightGBMDroid: A self-driven interactive mobile application utilizing explainable machine learning for breast cancer diagnosis</article-title>,&#x201D; <source>Eng. Rep.</source>, vol. <volume>5</volume>, no. <issue>11</issue>, <year>2023</year>, Art. no. <comment>e12666</comment>. doi: <pub-id pub-id-type="doi">10.1002/eng2.12666</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Catal</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Giray</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Tekinerdogan</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Kumar</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Shukla</surname></string-name></person-group>, &#x201C;<article-title>Applications of deep learning for phishing detection: A systematic literature review</article-title>,&#x201D; <source>Knowl. Inf. Syst.</source>, vol. <volume>64</volume>, no. <issue>6</issue>, pp. <fpage>1457</fpage>&#x2013;<lpage>1500</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s10115-022-01672-x</pub-id>; <pub-id pub-id-type="pmid">35645443</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Riggs</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Impact, vulnerabilities, and mitigation strategies for cyber-secure critical infrastructure</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>23</volume>, no. <issue>8</issue>, <year>2023</year>, Art. no. <comment>4060</comment>. doi: <pub-id pub-id-type="doi">10.3390/s23084060</pub-id>; <pub-id pub-id-type="pmid">37112400</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Hasal</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Nowakov&#x00E1;</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Ahmed Saghair</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Abdulla</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Sn&#x00E1;&#x0161;el</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Ogiela</surname></string-name></person-group>, &#x201C;<article-title>Chatbots: Security, privacy, data protection, and social aspects</article-title>,&#x201D; <source>Concurr. Comput.</source>, vol. <volume>33</volume>, no. <issue>19</issue>, <year>2021</year>, Art. no. <comment>e6426</comment>. doi: <pub-id pub-id-type="doi">10.1002/cpe.6426</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Samtani</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Abate</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Benjamin</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Cybersecurity as an industry: A cyber threat intelligence perspective</article-title>,&#x201D; in <source>Palgr. Handbook Int. Cybercrim. Cyberdevian.</source>, <person-group person-group-type="editor"><string-name><given-names>T.</given-names> <surname>Holt</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Bossler</surname></string-name></person-group>, Eds, <publisher-loc>Cham</publisher-loc>: <publisher-name>Palgrave Macmillan</publisher-name>, <year>2020</year>, pp. <fpage>135</fpage>&#x2013;<lpage>154</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-319-78440-3_8</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Moor</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Foundation models for generalist medical artificial intelligence</article-title>,&#x201D; <source>Nature</source>, vol. <volume>616</volume>, no. <issue>7956</issue>, pp. <fpage>259</fpage>&#x2013;<lpage>265</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1038/s41586-023-05881-4</pub-id>; <pub-id pub-id-type="pmid">37045921</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Priestman</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Anstis</surname></string-name>, <string-name><given-names>I. G.</given-names> <surname>Sebire</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Sridharan</surname></string-name>, and <string-name><given-names>N. J.</given-names> <surname>Sebire</surname></string-name></person-group>, &#x201C;<article-title>Phishing in healthcare organisations: Threats, mitigation and approaches</article-title>,&#x201D; <source>BMJ Health Care Inform.</source>, vol. <volume>26</volume>, no. <issue>1</issue>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1136/bmjhci-2019-100031</pub-id>; <pub-id pub-id-type="pmid">31488498</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. E.</given-names> <surname>Henry</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Human-machine teaming is key to AI adoption: Clinicians&#x2019; experiences with a deployed machine learning system</article-title>,&#x201D; <source>npj Dig. Med.</source>, vol. <volume>5</volume>, no. <issue>1</issue>, <year>2022</year>, Art. no. <comment>97</comment>. doi: <pub-id pub-id-type="doi">10.1038/s41746-022-00597-7</pub-id>; <pub-id pub-id-type="pmid">35864312</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Alabdan</surname></string-name></person-group>, &#x201C;<article-title>Phishing attacks survey: Types, vectors, and technical approaches</article-title>,&#x201D; <source>Future Internet</source>, vol. <volume>12</volume>, no. <issue>10</issue>, <year>2020</year>, Art. no. <comment>168</comment>. doi: <pub-id pub-id-type="doi">10.3390/fi12100168</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Xin</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X. -B.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Ota</surname></string-name></person-group>, &#x201C;<article-title>Efficient privacy-preserving in IoMT with blockchain and lightweight secret sharing</article-title>,&#x201D; <source>IEEE Internet Things J.</source>, vol. <volume>10</volume>, no. <issue>24</issue>, pp. <fpage>22051</fpage>&#x2013;<lpage>22064</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1109/JIOT.2023.3296595</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Singh</surname></string-name>, and <string-name><given-names>P.</given-names> <surname>Rajivan</surname></string-name></person-group>, &#x201C;<article-title>Personalized persuasion: Quantifying susceptibility to information exploitation in spear-phishing attacks</article-title>,&#x201D; <source>Appl. Ergon.</source>, vol. <volume>108</volume>, <year>2023</year>, Art. no. <comment>103908</comment>. doi: <pub-id pub-id-type="doi">10.1016/j.apergo.2022.103908</pub-id>; <pub-id pub-id-type="pmid">36403509</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sushma</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Viji</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Rajkumar</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Ravi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Stalin</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Najmusher</surname></string-name></person-group>, &#x201C;<article-title>Healthcare 4.0: A review of phishing attacks in cyber security</article-title>,&#x201D; <source>Procedia Comput. Sci.</source>, vol. <volume>230</volume>, no. <issue>10</issue>, pp. <fpage>874</fpage>&#x2013;<lpage>878</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.procs.2023.12.045</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Rizzoni</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Magalini</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Casaroli</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Mari</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Dixon</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Coventry</surname></string-name></person-group>, &#x201C;<article-title>Phishing simulation exercise in a large hospital: A case study</article-title>,&#x201D; <source>Digit. Health</source>, vol. <volume>8</volume>, <year>2022</year>, Art. no. <comment>20552076221081716</comment>. doi: <pub-id pub-id-type="doi">10.1177/20552076221081716</pub-id>; <pub-id pub-id-type="pmid">35321019</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A. M.</given-names> <surname>Hegde</surname></string-name>, <string-name><given-names>S. B.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Bhuvantej</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Vyshak</surname></string-name>, and <string-name><given-names>V.</given-names> <surname>Sarasvathi</surname></string-name></person-group>, &#x201C;<article-title>Spear phishing using machine learning</article-title>,&#x201D; in <conf-name>Int. Conf. Adv. Comput. Data Sci.</conf-name>, <year>2023</year>, pp. <fpage>529</fpage>&#x2013;<lpage>542</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Rege</surname></string-name> and <string-name><given-names>R. B. K.</given-names> <surname>Mbah</surname></string-name></person-group>, &#x201C;<article-title>Machine learning for cyber defense and attack</article-title>,&#x201D; in <conf-name>The Seventh Int. Conf. Data Anal.</conf-name>, <year>2018</year>, pp. <fpage>73</fpage>&#x2013;<lpage>78</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. -H.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>J. -L.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Machine learning mechanisms for cyber-phishing attack</article-title>,&#x201D; <source>IEICE Trans. Inf. Syst.</source>, vol. <volume>102</volume>, pp. <fpage>878</fpage>&#x2013;<lpage>887</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Alshammari</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Security monitoring and management for the network services in the orchestration of SDN-NFV environment using machine learning techniques</article-title>,&#x201D; <source>Comput. Syst. Sci. Eng.</source>, vol. <volume>48</volume>, no. <issue>2</issue>, pp. <fpage>363</fpage>&#x2013;<lpage>394</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.32604/csse.2023.040721</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Du</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Lu</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Phishing behavior detection on different blockchains via adversarial domain adaptation</article-title>,&#x201D; <source>Cybersecurity</source>, vol. <volume>7</volume>, no. <issue>1</issue>, <year>2024</year>, Art. no. <comment>45</comment>. doi: <pub-id pub-id-type="doi">10.1186/s42400-024-00237-5</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. P.</given-names> <surname>Ripa</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Islam</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Arifuzzaman</surname></string-name></person-group>, &#x201C;<article-title>The emergence threat of phishing attack and the detection techniques using machine learning models</article-title>,&#x201D; in <conf-name>2021 Int. Conf. Automat., Control Mech. Indus. 4.0 (ACMI)</conf-name>, <year>2021</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Xin</surname></string-name></person-group>, &#x201C;<article-title>Spear phishing emails detection based on machine learning</article-title>,&#x201D; in <conf-name>2021 IEEE 24th Int. Conf. Comput. Support. Cooperat. Work Des. (CSCWD)</conf-name>, <publisher-loc>Dalian, China</publisher-loc>, <year>2021</year>, pp. <fpage>354</fpage>&#x2013;<lpage>359</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CSCWD49262.2021.9437758</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Espinoza</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Simba</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Fuertes</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Benavides</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Andrade</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Toulkeridis</surname></string-name></person-group>, &#x201C;<article-title>Phishing attack detection: A solution based on the typical machine learning modeling cycle</article-title>,&#x201D; in <conf-name>2019 Int. Conf. Computat. Sci. Comput. Intell. (CSCI)</conf-name>, <year>2019</year>, pp. <fpage>202</fpage>&#x2013;<lpage>207</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. N.</given-names> <surname>Crothers</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Japkowicz</surname></string-name>, and <string-name><given-names>H. L.</given-names> <surname>Viktor</surname></string-name></person-group>, &#x201C;<article-title>Machine-generated text: A comprehensive survey of threat models and detection methods</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>11</volume>, pp. <fpage>70977</fpage>&#x2013;<lpage>71002</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3294090</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Smadi</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Aslam</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Detection of online phishing email using dynamic evolving neural network based on reinforcement learning</article-title>,&#x201D; <source>Decis. Support Syst.</source>, vol. <volume>107</volume>, pp. <fpage>88</fpage>&#x2013;<lpage>102</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.dss.2018.01.001</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Stevanovi&#x0107;</surname></string-name></person-group>, &#x201C;<article-title>Character and word embeddings for phishing email detection</article-title>,&#x201D; <source>Comput. Inform.</source>, vol. <volume>41</volume>, no. <issue>5</issue>, pp. <fpage>1337</fpage>&#x2013;<lpage>1357</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.31577/cai_2022_5_1337</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Moradpoor</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Clavie</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Buchanan</surname></string-name></person-group>, &#x201C;<article-title>Employing machine learning techniques for detection and classification of phishing emails</article-title>,&#x201D; in <conf-name>2017 Comput. Conf.</conf-name>, <year>2017</year>, pp. <fpage>149</fpage>&#x2013;<lpage>156</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>B.</given-names> <surname>An</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Kiekintveld</surname></string-name></person-group>, &#x201C;<article-title>An initial study on personalized filtering thresholds in defending sequential spear phishing attacks</article-title>,&#x201D; in <conf-name>Proc. 2015 IJCAI Workshop Behav., Econ. Computat. Intell. Secur.</conf-name>, <year>2015</year>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. A.</given-names> <surname>Aldakheel</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zakariah</surname></string-name>, <string-name><given-names>G. A.</given-names> <surname>Gashgari</surname></string-name>, <string-name><given-names>F. A.</given-names> <surname>Almarshad</surname></string-name>, and <string-name><given-names>A. I.</given-names> <surname>Alzahrani</surname></string-name></person-group>, &#x201C;<article-title>A deep learning-based innovative technique for phishing detection in modern security with uniform resource locators</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>23</volume>, no. <issue>9</issue>, <year>2023</year>, Art. no. <comment>4403</comment>. doi: <pub-id pub-id-type="doi">10.3390/s23094403</pub-id>; <pub-id pub-id-type="pmid">37177607</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Alotaibi</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Al-Turaiki</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Alakeel</surname></string-name></person-group>, &#x201C;<article-title>Mitigating email phishing attacks using convolutional neural networks</article-title>,&#x201D; in <conf-name>2020 3rd Int. Conf. Comput. App. Inf. Secur. (ICCAIS)</conf-name>, <year>2020</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Altwaijry</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Al-Turaiki</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Alotaibi</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Alakeel</surname></string-name></person-group>, &#x201C;<article-title>Advancing phishing email detection: A comparative study of deep learning models</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>24</volume>, no. <issue>7</issue>, <year>2024</year>, Art. no. <comment>2077</comment>. doi: <pub-id pub-id-type="doi">10.3390/s24072077</pub-id>; <pub-id pub-id-type="pmid">38610289</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="bool"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Samanta</surname></string-name>, <string-name><given-names>S. H.</given-names> <surname>Islam</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Chilamkurti</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Hammoudeh</surname></string-name></person-group>, <source>Data Analytics, Computational Statistics, and Operations Research for Engineers: Methodologies and Applications</source>, <edition>1st ed</edition>. <publisher-loc>Boca Raton, FL, USA</publisher-loc>: <publisher-name>CRC Press</publisher-name>, <year>2022</year>, pp. <fpage>203</fpage>&#x2013;<lpage>234</lpage>. doi: <pub-id pub-id-type="doi">10.1201/9781003152392</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. M. Ud</given-names> <surname>Din</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>InteliRank: A four-pronged agent for the intelligent ranking of cloud services based on end-users&#x2019; feedback</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>22</volume>, <year>2022</year>, Art. no. <comment>4627</comment>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Shabbir</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Ahmad</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Shabbir</surname></string-name>, and <string-name><given-names>S. A.</given-names> <surname>Alanazi</surname></string-name></person-group>, &#x201C;<article-title>Cognitively managed multi-level authentication for security using fuzzy logic based quantum key distribution</article-title>,&#x201D; <source>J. King Saud Univ.-Comput. Inf. Sci.</source>, vol. <volume>34</volume>, pp. <fpage>1468</fpage>&#x2013;<lpage>1485</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Frumento</surname></string-name></person-group>, &#x201C;<article-title>Cybersecurity and the evolutions of healthcare: Challenges and threats behind its evolution</article-title>,&#x201D; <source>M_Health Current Future App.</source>, pp. <fpage>35</fpage>&#x2013;<lpage>69</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Zhuo</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Biddle</surname></string-name>, <string-name><given-names>Y. S.</given-names> <surname>Koh</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Lottridge</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Russello</surname></string-name></person-group>, &#x201C;<article-title>SoK: Human-centered phishing susceptibility</article-title>,&#x201D; <source>ACM Trans. Priv. Secur.</source>, vol. <volume>26</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>27</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Subhadeep</surname></string-name></person-group>, &#x201C;<article-title>Phishing email detection</article-title>,&#x201D; <comment>2023. Accesed: Jul. 15, 2024</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/subhajournal/phishingemails/data">https://www.kaggle.com/datasets/subhajournal/phishingemails/data</ext-link></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Tesfom</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Belay</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Daniel</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Salem</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Otoum</surname></string-name></person-group>, &#x201C;<article-title>Phishing detection using deep learning and machine learning algorithms: Comparative analysis</article-title>,&#x201D; in <conf-name>2023 IEEE Int. Conf. Depend., Autonomic Secur. Comput., Int. Conf. Pervas. Intell. Comput., Int. Conf. Cloud Big Data Comput., Int. Conf. Cyber Sci. Technol. Congress (DASC/PiCom/CBDCom/CyberSciTech)</conf-name>, <year>2023</year>, pp. <fpage>684</fpage>&#x2013;<lpage>689</lpage>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Mittal</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Agarwal</surname></string-name>, <string-name><given-names>M. L.</given-names> <surname>Saini</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>A logistic regression approach for detecting phishing websites</article-title>,&#x201D; in <conf-name>2023 Int. Conf. Adv. Comput., Commun. Inf. Technol. (ICAICCIT)</conf-name>, <year>2023</year>, pp. <fpage>76</fpage>&#x2013;<lpage>81</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. R.</given-names> <surname>Abdul Samad</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Analysis of the performance impact of fine-tuned machine learning model for phishing URL detection</article-title>,&#x201D; <source>Electronics</source>, vol. <volume>12</volume>, <year>2023</year>, Art. no. <comment>1642</comment>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Musa</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Gital</surname></string-name>, <string-name><given-names>F. U.</given-names> <surname>Zambuk</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Umar</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Umar</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Waziri</surname></string-name></person-group>, &#x201C;<article-title>A comparative analysis of phishing website detection using XGBOOST algorithm</article-title>,&#x201D; <source>J. Theoret. Appl. Informat. Technol.</source>, vol. <volume>97</volume>, pp. <fpage>1434</fpage>&#x2013;<lpage>1443</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. A.</given-names> <surname>Fazal</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Daud</surname></string-name></person-group>, &#x201C;<article-title>Detecting phishing websites using Decision Trees: A machine learning approach</article-title>,&#x201D; <source>Int. J. Electron. Crime Invest.</source>, vol. <volume>7</volume>, no. <issue>2</issue>, pp. <fpage>73</fpage>&#x2013;<lpage>79</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M. F. Ab</given-names> <surname>Razak</surname></string-name>, <string-name><given-names>M. I.</given-names> <surname>Jaya</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Ernawan</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Firdaus</surname></string-name>, and <string-name><given-names>F. A.</given-names> <surname>Nugroho</surname></string-name></person-group>, &#x201C;<article-title>Comparative analysis of machine learning classifiers for phishing detection</article-title>,&#x201D; in <conf-name>2022 6th Int. Conf. Inf. Computat. Sci. (ICICoS)</conf-name>, <year>2022</year>, pp. <fpage>84</fpage>&#x2013;<lpage>88</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Subasi</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Kremic</surname></string-name></person-group>, &#x201C;<article-title>Comparison of adaboost with multiboosting for phishing website detection</article-title>,&#x201D; <source>Procedia Comput. Sci.</source>, vol. <volume>168</volume>, no. <issue>2</issue>, pp. <fpage>272</fpage>&#x2013;<lpage>278</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1016/j.procs.2020.02.251</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>P. F.</given-names> <surname>Akinwale</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Jahankhani</surname></string-name></person-group>, &#x201C;<article-title>Detection and binary classification of spear-phishing emails in organizations using a hybrid machine learning approach</article-title>,&#x201D; in <source>Artificial Intelligence in Cyber Security: Impact and Implications. Advanced Sciences and Technologies for Security Applications</source>, <person-group person-group-type="editor"><string-name><given-names>T.</given-names> <surname>Holt</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Bossler</surname></string-name></person-group>, Eds, <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>, <year>2022</year>, pp. <fpage>215</fpage>&#x2013;<lpage>252</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-030-88040-8_9</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A. C.</given-names> <surname>Bahnsen</surname></string-name>, <string-name><given-names>E. C.</given-names> <surname>Bohorquez</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Villegas</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Vargas</surname></string-name>, and <string-name><given-names>F. A.</given-names> <surname>Gonz&#x00E1;lez</surname></string-name></person-group>, &#x201C;<article-title>Classifying phishing URLs using recurrent neural networks</article-title>,&#x201D; in <conf-name>2017 APWG Symp. Electron. Crime Res. (eCrime)</conf-name>, <year>2017</year>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. A.</given-names> <surname>Adebowale</surname></string-name>, <string-name><given-names>K. T.</given-names> <surname>Lwin</surname></string-name>, and <string-name><given-names>M. A.</given-names> <surname>Hossain</surname></string-name></person-group>, &#x201C;<article-title>Deep learning with convolutional neural network and long short-term memory for phishing detection</article-title>,&#x201D; in <source>2019 13th Int. Conf. Softw., Knowl., Inform. Manag. Appl. (SKIMA)</source>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/SKIMA47702.2019</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. S.</given-names> <surname>Roy</surname></string-name>, <string-name><given-names>A. I.</given-names> <surname>Awad</surname></string-name>, <string-name><given-names>L. A.</given-names> <surname>Amare</surname></string-name>, <string-name><given-names>M. T.</given-names> <surname>Erkihun</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Anas</surname></string-name></person-group>, &#x201C;<article-title>Multimodel phishing URL detection using LSTM, bidirectional LSTM, and GRU models</article-title>,&#x201D; <source>Future Internet</source>, vol. <volume>14</volume>, no. <issue>11</issue>, <year>2022</year>, Art. no. <comment>340</comment>. doi: <pub-id pub-id-type="doi">10.3390/fi14110340</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. T.</given-names> <surname>Jafar</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Al-Fawa&#x2019;reh</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Barhoush</surname></string-name>, and <string-name><given-names>M. H.</given-names> <surname>Alshira&#x2019;H</surname></string-name></person-group>, &#x201C;<article-title>Enhanc&#x0435;d analysis approach to detect phishing attacks during COVID-19 crisis</article-title>,&#x201D; <source>Cybern. Inf. Technol.</source>, vol. <volume>22</volume>, no. <issue>1</issue>, pp. <fpage>60</fpage>&#x2013;<lpage>76</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.2478/cait-2022-0004</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>McGinley</surname></string-name> and <string-name><given-names>S. A. S.</given-names> <surname>Monroy</surname></string-name></person-group>, &#x201C;<article-title>Convolutional neural network optimization for phishing email classification</article-title>,&#x201D; in <conf-name>2021 IEEE Int. Conf. Big Data (Big Data)</conf-name>, <year>2021</year>, pp. <fpage>5609</fpage>&#x2013;<lpage>5613</lpage>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K. Z.</given-names> <surname>Hasan</surname></string-name>, <string-name><given-names>M. Z.</given-names> <surname>Hasan</surname></string-name>, and <string-name><given-names>N.</given-names> <surname>Zahan</surname></string-name></person-group>, &#x201C;<article-title>Automated prediction of phishing websites using deep convolutional neural network</article-title>,&#x201D; in <conf-name>2019 Int. Conf. Comput., Commun., Chem., Mater. Electron. Eng. (IC4ME2)</conf-name>, <year>2019</year>, pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>