<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">77367</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.077367</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Large Language Models for Cybersecurity Intelligence: A Systematic Review of Emerging Threats, Defensive Capabilities, and Security Evaluation Frameworks</article-title>
<alt-title alt-title-type="left-running-head">Large Language Models for Cybersecurity Intelligence: A Systematic Review of Emerging Threats, Defensive Capabilities, and Security Evaluation Frameworks</alt-title>
<alt-title alt-title-type="right-running-head">Large Language Models for Cybersecurity Intelligence: A Systematic Review of Emerging Threats, Defensive Capabilities, and Security Evaluation Frameworks</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Alqahtani</surname><given-names>Hamed</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kumar</surname><given-names>Gulshan</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>gulshanahuja@gmail.com</email></contrib>
<aff id="aff-1"><label>1</label><institution>Informatics and Computer Systems Department, College of Computer Science, Center of Artificial Intelligence, King Khalid University</institution>, <addr-line>Abha</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Applications, Shaheed Bhagat Singh State University, Ferozepur</institution>, <addr-line>Punjab</addr-line>, <country>India</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Gulshan Kumar. Email: <email>gulshanahuja@gmail.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>9</day><month>4</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>3</issue>
<elocation-id>9</elocation-id>
<history>
<date date-type="received">
<day>08</day>
<month>12</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>02</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_77367.pdf"></self-uri>
<abstract>
<p>Large Language Models (LLMs) are becoming integral components of modern cybersecurity ecosystems, simultaneously strengthening defensive capabilities while giving rise to a new class of Artificial Intelligence&#x2013;Generated Content (AIGC)-driven threats. This PRISMA-guided systematic review synthesises 167 peer-reviewed studies published between 2022 and 2025 and proposes a unified threat&#x2013;defence&#x2013;evaluation taxonomy as a central analytical framework to consolidate a previously fragmented body of research. Guided by this taxonomy, the review first examines AIGC-enabled threats, including automated and highly personalised phishing, polymorphic malware and exploit generation, jailbreak and adversarial prompting, prompt-injection attack vectors, multimodal deception, persona-steering attacks, and large-scale disinformation campaigns. The surveyed evidence indicates a qualitative escalation in adversarial capabilities, with LLMs significantly enhancing scalability, adaptability, and realism while markedly reducing the technical barriers to conducting sophisticated attacks. Second, the review analyses LLM-enabled defensive applications spanning intrusion and anomaly detection, malware analysis and log-semantic modelling, multilingual threat intelligence extraction, vulnerability discovery and code repair, and Security Operations Center (SOC) automation through Retrieval-Augmented Generation (RAG) and multi-agent systems. Although these approaches demonstrate strong potential as semantic reasoning and decision-support components within hybrid security architectures, their real-world effectiveness remains constrained by hallucination risks, adversarial susceptibility, distributional shifts, and operational overhead. Third, the review synthesises current security evaluation and red-teaming practices, revealing a fragmented assessment landscape characterised by narrow benchmarks, inconsistent evaluation metrics, and limited longitudinal robustness analysis. Overall, the taxonomy-driven synthesis highlights a structurally imbalanced ecosystem in which offensive innovation outpaces defensive maturity and governance, and it informs a structured, research-question-aligned roadmap for developing trustworthy, resilient, and policy-aligned LLM-powered cybersecurity systems.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Artificial intelligence&#x2013;generated content threats</kwd>
<kwd>cybersecurity intelligence</kwd>
<kwd>large language model&#x2013;based defensive systems</kwd>
<kwd>large language models</kwd>
<kwd>red-teaming and evaluation frameworks</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Scientific Research at King Khalid University</funding-source>
<award-id>GRP.2/663/46</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Large Language Models (LLMs), developed using large-scale transformer architectures, have rapidly emerged as foundational components of contemporary artificial intelligence, enabling significant advances in reasoning, code generation, summarisation, automated analysis, and decision support across high-stakes application domains [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. Their capacity to process heterogeneous, unstructured, and multilingual data at scale renders them particularly well suited to cybersecurity environments, where analysts must interpret complex and high-volume artefacts, including system logs, threat intelligence reports, malware analyses, vulnerability disclosures, and Security Operations Center (SOC) narratives [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>]. At the same time, the same generative and reasoning capabilities that enable defensive augmentation can be readily exploited by adversaries. A rapidly expanding body of peer-reviewed research documents the use of LLMs to craft highly personalised phishing campaigns [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>], generate polymorphic and evasive malware [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>], automate reconnaissance and exploit scaffolding [<xref ref-type="bibr" rid="ref-10">10</xref>], conduct multimodal deception and persona-steering attacks [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>], and orchestrate large-scale misinformation and disinformation operations [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>]. This dual-use character has profound implications for the evolving balance between cyber offence and cyber defence.</p>
<p>Despite growing interest in this area, existing surveys exhibit substantial variation in scope, methodological rigor, and evidentiary depth. Yao et al. synthesise model-level vulnerabilities and privacy risks associated with LLM deployment [<xref ref-type="bibr" rid="ref-15">15</xref>]. Jaffal et al. provide a detailed examination of LLM capabilities, limitations, and defensive applications within SOC and cyber threat intelligence (CTI) ecosystems [<xref ref-type="bibr" rid="ref-3">3</xref>]. Zhang et al. offer a structured review of LLM-driven cybersecurity applications [<xref ref-type="bibr" rid="ref-16">16</xref>], while Karras et al. discuss broader opportunities and risks in big-data security pipelines [<xref ref-type="bibr" rid="ref-4">4</xref>]. However, these surveys share notable limitations: many rely primarily on narrative synthesis rather than systematic review protocols; several include a substantial proportion of preprint literature; few integrate offensive, defensive, and evaluative research strands within a single analytical framework; and none provide a PRISMA [<xref ref-type="bibr" rid="ref-17">17</xref>]-aligned, SCI/SCIE-restricted taxonomy that jointly captures threats, defences, and security evaluation methodologies. As the field continues to expand rapidly, the absence of a comprehensive and methodologically transparent synthesis impedes cross-study comparability, evidence consolidation, and the standardisation of research practices.</p>
<p>To address these gaps, this article presents the first PRISMA-compliant systematic survey of LLMs for cybersecurity intelligence covering the period 2022&#x2013;2025, explicitly restricted to peer-reviewed SCI/SCIE publications. The overarching objective is to construct a rigorous and unified evidence base that characterises how LLMs reshape cyber offence, cyber defence, and the evaluation ecosystems governing trustworthy deployment. Specifically, this study pursues four core aims:
<list list-type="bullet">
<list-item>
<p>to systematically identify and synthesise Artificial Intelligence&#x2013;Generated Content (AIGC)-driven cyber threats enabled, amplified, or operationalised by LLMs, including phishing, malware generation, jailbreak attacks, deception, misinformation, and persona steering [<xref ref-type="bibr" rid="ref-11">11</xref>];</p></list-item>
<list-item>
<p>to analyse LLM-enabled defensive capabilities spanning intrusion detection [<xref ref-type="bibr" rid="ref-18">18</xref>], malware semantic modelling [<xref ref-type="bibr" rid="ref-19">19</xref>], threat intelligence extraction [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>], vulnerability assessment and remediation [<xref ref-type="bibr" rid="ref-22">22</xref>], SOC automation [<xref ref-type="bibr" rid="ref-23">23</xref>], and multi-agent defence architectures [<xref ref-type="bibr" rid="ref-24">24</xref>];</p></list-item>
<list-item>
<p>to examine emerging security evaluation frameworks, including benchmarks [<xref ref-type="bibr" rid="ref-25">25</xref>], harm and refusal metrics [<xref ref-type="bibr" rid="ref-3">3</xref>], adversarial red-teaming methodologies [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>], and lifecycle-oriented risk assessments [<xref ref-type="bibr" rid="ref-27">27</xref>], alongside the growing need for ethical and governance frameworks [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>];</p></list-item>
<list-item>
<p>to distil cross-cutting methodological limitations, systemic risks (including privacy [<xref ref-type="bibr" rid="ref-30">30</xref>] and ethical compliance [<xref ref-type="bibr" rid="ref-31">31</xref>]), and unresolved research challenges derived from 167 peer-reviewed primary studies and several high-impact surveys.</p></list-item>
</list></p>
<p>These aims motivate the following research questions:
<list list-type="bullet">
<list-item>
<p><bold>RQ1:</bold> What forms of cybersecurity threats are generated, enhanced, or operationalised through AIGC and LLMs, and how are these threats empirically studied and modelled?</p></list-item>
<list-item>
<p><bold>RQ2:</bold> How are LLMs leveraged to strengthen cyber defence capabilities across detection, analysis, and response workflows, and what limitations persist?</p></list-item>
<list-item>
<p><bold>RQ3:</bold> What security evaluation frameworks, benchmarks, and red-teaming methodologies exist for assessing LLM robustness, safety, and reliability under adversarial conditions?</p></list-item>
<list-item>
<p><bold>RQ4:</bold> What methodological gaps, systemic challenges, and governance limitations hinder the trustworthy deployment of LLM-based cybersecurity systems?</p></list-item>
</list></p>
<p>The primary contributions of this survey are as follows:
<list list-type="bullet">
<list-item>
<p>a PRISMA-aligned systematic review protocol tailored to LLM-centric cybersecurity research, detailing search strategies, inclusion and exclusion criteria, quality appraisal procedures, and multi-stage screening;</p></list-item>
<list-item>
<p>a critical comparison of influential surveys [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>], motivating the unified threat&#x2013;defence&#x2013;evaluation taxonomy introduced in this work;</p></list-item>
<list-item>
<p>a structured synthesis of 167 peer-reviewed primary studies addressing AIGC-driven threats [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>], LLM-based defensive mechanisms [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>], and security evaluation and red-teaming frameworks [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>];</p></list-item>
<list-item>
<p>the identification of cross-cutting challenges&#x2014;including hallucination, adversarial fragility, privacy and governance risks, evaluation fragmentation, lifecycle vulnerabilities, and the absence of standardised assurance frameworks&#x2014;followed by a comprehensive future research roadmap for resilient, auditable, and security-aligned LLM deployments.</p></list-item>
</list></p>
<p>In contrast to prior surveys, this review grounds each threat and defence category in empirically evaluated case studies and critically analyses their operational limitations under realistic SOC and intrusion detection deployment scenarios. Overall, this survey offers a methodologically rigorous, conceptually integrated, and empirically grounded examination of how LLMs are transforming the cybersecurity landscape, both as powerful enablers of defensive intelligence and as catalysts for increasingly adaptive and scalable adversarial behaviour. By consolidating diverse research trajectories and introducing a unified taxonomy alongside a forward-looking research agenda, this work aims to support the development of trustworthy, verifiable, and operationally reliable LLM-driven cybersecurity systems.</p>
<p>The remainder of this article is organised as follows. <xref ref-type="sec" rid="s2">Section 2</xref> describes the PRISMA-guided review methodology. <xref ref-type="sec" rid="s3">Section 3</xref> situates this study within the existing survey literature. <xref ref-type="sec" rid="s4">Section 4</xref> presents the unified threat&#x2013;defence&#x2013;evaluation taxonomy and analytical framework. <xref ref-type="sec" rid="s5">Sections 5</xref> and <xref ref-type="sec" rid="s6">6</xref> analyse AIGC-driven threats and LLM-based defensive capabilities, while <xref ref-type="sec" rid="s7">Section 7</xref> reviews emerging security evaluation and red-teaming methodologies. <xref ref-type="sec" rid="s8">Section 8</xref> examines methodological patterns across the evidence base. <xref ref-type="sec" rid="s9">Section 9</xref> synthesises cross-cutting challenges, and <xref ref-type="sec" rid="s10">Section 10</xref> outlines strategic future research directions. <xref ref-type="sec" rid="s11">Section 11</xref> concludes the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Methodology</title>
<p>This review adopts a rigorously defined systematic literature review (SLR) methodology grounded in the PRISMA 2020 guidelines to ensure transparency, reproducibility, and analytical completeness. Given the rapid evolution and methodological heterogeneity of Large Language Model (LLM) research in cybersecurity, particular emphasis is placed on bias mitigation, inductive taxonomy construction, and cross-study comparability. The final evidence base comprises 167 peer-reviewed studies published between 2022 and 2025, each systematically mapped to the unified threat&#x2013;defence&#x2013;evaluation taxonomy introduced in <xref ref-type="sec" rid="s4">Section 4</xref>. An overview of the study identification and selection process is provided in the PRISMA flow diagram shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>PRISMA flow diagram summarising identification, screening, eligibility assessment, and final inclusion of 167 studies.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77367-fig-1.tif"/>
</fig>
<sec id="s2_1">
<label>2.1</label>
<title>Review Protocol and PRISMA Workflow</title>
<p>The review protocol was specified <italic>a priori</italic> in accordance with PRISMA recommendations to minimise selection bias and enhance methodological transparency. The review workflow comprised four sequential stages: identification, screening, eligibility assessment, and final inclusion. The initial search yielded 1002 records across multiple scholarly databases. Following automated and manual deduplication, 889 unique studies were subjected to title and abstract screening, resulting in the exclusion of 430 records. The remaining 459 studies underwent detailed screening, leading to 263 full-text articles being assessed for eligibility. Of these, 96 studies were excluded due to insufficient cybersecurity relevance, incomplete methodological reporting, or failure to satisfy the predefined inclusion criteria. The final corpus therefore consisted of 167 peer-reviewed studies, representing a balanced cross-section of AIGC-driven threats, LLM-enabled defensive architectures, and security evaluation, safety, and governance frameworks. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> provides a transparent visual summary of this selection process.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Inductive Taxonomy Construction and Coding Rationale</title>
<p>Rather than imposing an externally defined classification scheme, this review employs a bottom-up inductive coding strategy to derive the unified taxonomy presented in <xref ref-type="sec" rid="s4">Section 4</xref>. Following full-text screening, each study was independently coded along multiple analytical dimensions, including: (i) the primary cybersecurity task addressed, (ii) the functional role of the LLM within the system architecture, (iii) the threat or defensive context, (iv) the evaluation or validation methodology, and (v) the assumed deployment environment (e.g., laboratory-scale, SOC-scale, or safety-critical systems).</p>
<p>Initial open coding revealed recurring conceptual and methodological patterns across ostensibly diverse contributions. Through iterative axial coding and constant comparative analysis, these patterns converged into four dominant dimensions: <bold>(T) AIGC-driven threat vectors</bold>, <bold>(D) LLM-based defensive capabilities</bold>, <bold>(E) evaluation and red-teaming methodologies</bold>, and <bold>(M) methodological and governance orientation</bold>. For instance, studies addressing automated phishing, misinformation generation, exploit scaffolding, and prompt-based attacks consistently clustered within a unified threat dimension, whereas contributions focusing on SOC automation, intrusion detection, vulnerability analysis, and agentic response mechanisms formed a distinct defensive axis.</p>
<p>Subcategories within each dimension were further refined through frequency analysis and conceptual differentiation. Within the defensive dimension, intrusion detection, threat intelligence automation, secure code analysis, and autonomous response emerged as separable categories based on differences in data modality, latency constraints, and operational responsibility. Evaluation-oriented studies similarly diverged into benchmark construction, adversarial red-teaming, metric development, and lifecycle robustness analysis. Edge cases and hybrid studies were explicitly retained and cross-mapped to reflect the interdisciplinary and rapidly converging nature of LLM-driven cybersecurity research. This inductive process ensures that the resulting taxonomy faithfully reflects the empirical structure of the literature rather than an <italic>a priori</italic> conceptual framework.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Search Strategy and Information Sources</title>
<p>Searches were conducted across eight authoritative scholarly databases: IEEE Xplore, ACM Digital Library, ScienceDirect, SpringerLink, Wiley Online Library, Taylor &#x0026; Francis, Scopus, and Web of Science. These sources were selected to capture interdisciplinary research spanning cybersecurity, artificial intelligence, software engineering, and digital governance. The review period extends from January 2022 to December 2025, corresponding to the widespread deployment of transformer-based LLMs (e.g., GPT-3&#x002B;, LLaMA, PaLM, Codex) in both offensive and defensive cybersecurity contexts.</p>
<p>Search queries combined terms related to LLM architectures, cyber threat categories, defensive applications, and evaluation constructs using Boolean operators:
<disp-quote>
<p>(<italic>&#x201C;large language model&#x201D; OR LLM OR &#x201C;foundation model&#x201D; OR transformer</italic>) AND (<italic>cybersecurity OR malware OR phishing OR vulnerability OR anomaly OR SOC</italic>) AND (<italic>AIGC OR &#x201C;automated phishing&#x201D; OR jailbreak OR &#x201C;prompt injection&#x201D; OR misuse</italic>)</p>
</disp-quote></p>
<p>Search expressions were iteratively refined through pilot queries to incorporate emerging themes such as adversarial prompting, exploit generation, multi-agent reasoning, and cyber governance. Foundational surveys on LLM security and misuse [<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>] informed construct definition and ensured terminological consistency across the search strategy.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Inclusion, Exclusion, and Quality Appraisal</title>
<p>Studies were included if they were peer-reviewed journal or conference publications written in English between 2022 and 2025 and examined LLMs or transformer-based models within cybersecurity-relevant contexts. Eligible studies addressed AIGC-enabled threats such as phishing, deception, exploit scaffolding, or jailbreak attacks [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]; proposed or evaluated LLM-based defensive mechanisms including SOC automation, threat intelligence extraction, intrusion or anomaly detection, and vulnerability analysis [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>]; or contributed to evaluation, red-teaming, benchmark design, or governance frameworks for LLM safety.</p>
<p>Exclusion criteria removed preprints, editorials, theses, and other non-peer-reviewed materials; studies lacking substantive cybersecurity relevance; purely generic natural language processing applications without adversarial or security implications; works with insufficient methodological detail; and duplicate publications. Methodological quality was assessed using a bespoke appraisal framework adapted from established SLR guidelines, evaluating the clarity of threat models, transparency of datasets and evaluation pipelines, reproducibility, metric suitability, and the discussion of limitations and ethical risks. Studies exhibiting methodological weaknesses were retained when they offered unique empirical or conceptual insights, with their limitations explicitly acknowledged in subsequent synthesis sections.</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Inter-Reviewer Agreement, Bias Mitigation, and Sensitivity Analysis</title>
<p>To ensure consistency and minimise subjective bias, study selection and coding were independently performed by two reviewers at both the title&#x2013;abstract and full-text screening stages. Inter-reviewer agreement was quantified using Cohen&#x2019;s kappa coefficient, yielding substantial agreement during title&#x2013;abstract screening (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>&#x03BA;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.81</mml:mn></mml:math></inline-formula>) and near-perfect agreement during full-text eligibility assessment (<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>&#x03BA;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.88</mml:mn></mml:math></inline-formula>). Discrepancies were resolved through structured discussion and, where necessary, arbitration by a third reviewer.</p>
<p>Bias mitigation measures included multi-database coverage to reduce venue bias, <italic>a priori</italic> inclusion criteria, blinded screening where feasible, and the use of a standardised, taxonomy-aligned coding template to limit interpretive drift. A sensitivity analysis was conducted on borderline studies excluded during full-text screening (e.g., partially relevant preprints or studies with limited evaluation detail). Reintroduction of these studies did not materially affect the taxonomy structure, thematic distributions, or comparative conclusions, indicating that the review findings are robust and not unduly sensitive to marginal inclusion decisions.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Positioning with Respect to Existing Surveys</title>
<p>A substantial body of survey literature has examined intrusion detection systems (IDS) from the perspectives of machine learning, deep learning, and, more recently, transformer-based architectures. These IDS-centric surveys predominantly focus on network intrusion detection systems (NIDS), host-based intrusion detection systems (HIDS), and industrial or IoT-oriented IDS deployments, with primary emphasis on feature engineering strategies, model architectures, benchmark datasets, and performance metrics [<xref ref-type="bibr" rid="ref-37">37</xref>]. Representative works include surveys on transformer-based NIDS, hybrid deep learning IDS frameworks, and intrusion detection in IIoT and industrial control system (ICS) environments, where LLMs-if mentioned at all-are treated as high-capacity sequence models or feature-level classifiers rather than as reasoning or agentic components. Consequently, while these surveys provide valuable insights into detection accuracy, scalability, and dataset-specific performance, they do not address the broader implications of LLMs as cognitive security primitives operating across detection, reasoning, and response pipelines.</p>
<p>In parallel, a distinct line of survey research has emerged that focuses on the security, reliability, and societal implications of large language models. Foundational overviews by Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>], Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>], and Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] provide early mappings of the opportunities and risks introduced by generative AI. Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] conduct a comprehensive examination of ChatGPT, analysing its capabilities, limitations, misuse potential, and associated ethical and privacy concerns. Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] systematically catalogue attack techniques and mitigation strategies related to LLM misuse, including data leakage and adversarial behaviours, while Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] explore application-level opportunities and challenges. However, these surveys remain largely high-level and conceptual, and they do not engage in systematic comparison with IDS-centric research nor integrate intrusion detection within a broader cybersecurity lifecycle.</p>
<p>More recent contributions further narrow the scope toward governance, lifecycle risk, and domain-specific security applications. Nawara and Kashef [<xref ref-type="bibr" rid="ref-36">36</xref>] investigate LLM-powered recommendation and reasoning systems, highlighting security and governance challenges in decision-support pipelines, while Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] present a critical analysis of generative AI risks, emphasising lifecycle vulnerabilities, regulatory gaps, and systemic misuse pathways relevant to cybersecurity. Despite their analytical depth, these works continue to treat intrusion detection, SOC automation, and response mechanisms as largely isolated application domains rather than as interdependent components within an end-to-end threat&#x2013;defence&#x2013;evaluation ecosystem.</p>
<p>The primary distinction between existing IDS-focused surveys and the present work lies in the level of abstraction and the conceptual role assigned to LLMs. IDS-centric surveys predominantly frame models-including transformers-as <italic>feature-level classifiers</italic> optimised for detection accuracy on specific datasets. In contrast, this review conceptualises LLMs as <italic>cognitive and agentic security components</italic> capable of semantic reasoning, contextual correlation, decision support, and autonomous coordination across multiple stages of the cybersecurity pipeline. As a result, intrusion detection is not analysed in isolation but is examined alongside AIGC-driven threats, SOC automation, multi-agent response systems, and security evaluation and red-teaming practices.</p>
<p><xref ref-type="table" rid="table-1">Table 1</xref> systematically positions representative LLM- and IDS-focused surveys within the dimensions of the proposed unified taxonomy. As the comparison illustrates, prior surveys typically concentrate on isolated aspects of the problem space, such as AIGC-driven threat characterisation [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>], application-level opportunities and limitations [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>], or governance and lifecycle risk analysis [<xref ref-type="bibr" rid="ref-38">38</xref>]. However, none of these works provides a fully integrated, multi-dimensional perspective that jointly examines offensive misuse, defensive reasoning, evaluation and red-teaming practices, and methodological orientation within a single analytical framework. Furthermore, the absence of PRISMA-aligned screening and eligibility procedures across existing surveys limits reproducibility and systematic coverage.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison of existing survey taxonomies across IDS- and LLM-centric dimensions.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Survey</th>
<th>A</th>
<th>B</th>
<th>C</th>
<th>D</th>
<th>E</th>
<th>F</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="7"><italic>LLM-Centric Security Surveys</italic></td>
</tr>
<tr>
<td>Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Nawara et al. [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td><inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td align="center" colspan="7"><italic>IDS-Centric (Transformer/DL-Based) Surveys</italic></td>
</tr>
<tr>
<td>Elouardi et al. [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td><inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Ferrag et al. [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td><inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Hasanov et al. [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td><inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
</tr>
<tr>
<td><bold>This Survey</bold></td>
<td><bold>&#x2713;</bold></td>
<td><bold>&#x2713;</bold></td>
<td><bold>&#x2713;</bold></td>
<td><bold>&#x2713;</bold></td>
<td><bold>&#x2713;</bold></td>
<td><bold>&#x2713;</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-1fn1" fn-type="other">
<p>Note: A: AIGC-Driven Threat Coverage; B: LLM-Based Defensive Capabilities (including IDS and SOC automation); C: Evaluation &#x0026; Red-Teaming Coverage; D: Unified Multi-Dimensional Taxonomy; E: PRISMA or SLR Alignment; F: Cognitive/Agentic Role of Models (beyond feature-level classification).</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>A further distinguishing factor, highlighted of <xref ref-type="table" rid="table-1">Table 1</xref>, concerns model conceptualisation. Existing IDS-centric surveys&#x2014;particularly those focusing on transformer-based NIDS, HIDS, and IIoT detection&#x2014;primarily treat deep learning models as feature-level classifiers optimised for detection performance. In contrast, the present review frames LLMs as cognitive and agentic security components that enable semantic reasoning, contextual awareness, and coordinated decision-making across the full cybersecurity lifecycle, encompassing detection, analysis, response, and evaluation.</p>

<p>Unlike prior IDS-centric and LLM-centric reviews, the present study adopts a PRISMA-guided systematic methodology, restricts its evidence base to 167 peer-reviewed publications from 2022 to 2025, and synthesises findings using a unified four-dimensional taxonomy encompassing AIGC-driven threats, LLM-enabled defensive capabilities (including intrusion detection and SOC automation), security evaluation and red-teaming practices, and methodological orientation. This integrated framework enables a holistic characterisation of LLM-driven cybersecurity research and explicitly bridges the conceptual gap between traditional IDS surveys and emerging LLM security studies. By positioning intrusion detection as one component within a broader cognitive and agentic defence ecosystem, this review provides a more comprehensive, coherent, and methodologically transparent synthesis to inform future research and real-world deployment.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>A Unified Taxonomy of LLM&#x2013;Cybersecurity Research</title>
<p>To systematically synthesise the heterogeneous body of literature identified through the PRISMA-guided review process, this section introduces a unified taxonomy of LLM-centric cybersecurity research. The taxonomy is derived through an inductive coding and synthesis procedure applied to all 167 included studies, as described in <xref ref-type="sec" rid="s2_2">Section 2.2</xref>. Rather than imposing a pre-defined classification scheme, the taxonomy is grounded in empirically observed research patterns spanning adversarial misuse, defensive integration, and security evaluation practices.</p>
<p>The proposed taxonomy organises the literature along four complementary dimensions: (i) AIGC-driven cyber threats, (ii) LLM-enabled defensive capabilities, (iii) security evaluation, red-teaming, and governance, and (iv) methodological orientation. Collectively, these dimensions capture not only <italic>what</italic> security problems are addressed, but also <italic>how</italic> evidence is generated, validated, and operationalised. This structure provides a reproducible analytical framework for comparative analysis, exposes systematic imbalances in the existing research landscape, and supports the identification of underexplored yet high-impact research directions.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Rationale for the Proposed Taxonomy</title>
<p>The motivation for the proposed taxonomy is threefold. First, existing surveys frequently examine offensive misuse, defensive applications, or ethical considerations of LLMs in isolation, resulting in fragmented insights that obscure the co-evolution of threats and countermeasures [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. By explicitly integrating threat, defence, and evaluation dimensions, the proposed taxonomy captures their interdependencies and reflects the concurrent evolution of attacker innovation, defensive adaptation, and governance constraints.</p>
<p>Second, the taxonomy elevates <italic>security evaluation and governance</italic> to a primary analytical dimension rather than treating it as a secondary or peripheral concern. An expanding body of work demonstrates that deficiencies in robustness auditing, red-teaming, explainability, and regulatory alignment can undermine otherwise promising technical advances [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Explicitly foregrounding evaluation and governance therefore aligns the taxonomy with real-world deployment requirements and emerging regulatory expectations.</p>
<p>Third, the taxonomy is directly grounded in the textual, empirical, and methodological evidence of the 167 reviewed studies. Categories and subcategories were iteratively refined through frequency analysis, conceptual differentiation, and cross-mapping of hybrid contributions, ensuring analytical coherence without sacrificing completeness. This inductive grounding strengthens methodological transparency and avoids the conceptual overfitting that often characterises high-level AI taxonomies.</p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> presents a visual overview of the unified taxonomy, illustrating how threat classes, defensive mechanisms, evaluation frameworks, and methodological orientations collectively structure the emerging LLM&#x2013;cybersecurity research landscape.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Proposed unified taxonomy of research on LLMs for cybersecurity intelligence.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77367-fig-2.tif"/>
</fig>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>AIGC-Driven Cyber Threats</title>
<p>The first dimension of the taxonomy captures how generative AI and LLMs expand the cyber threat surface by automating deception, lowering technical skill barriers, and enabling adaptive attacks at scale. Across the reviewed literature, AIGC-driven threats consistently emerge as a dominant research focus, reflecting widespread concern regarding the misuse potential of generative models [<xref ref-type="bibr" rid="ref-41">41</xref>]. Unlike conventional cyber threats, these attacks exploit linguistic fluency, contextual awareness, and automation, enabling adversaries with limited expertise to execute sophisticated campaigns.</p>
<p>Empirical and analytical studies converge around three recurring threat clusters: (i) social engineering and deception automation, (ii) malware generation, exploit scaffolding, and code abuse, and (iii) model manipulation, prompt injection, and alignment evasion. These clusters recur across security-focused surveys, AI safety analyses, and governance-oriented discussions [<xref ref-type="bibr" rid="ref-42">42</xref>,<xref ref-type="bibr" rid="ref-43">43</xref>], supporting their treatment as distinct yet interrelated subcategories.</p>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Social Engineering and Deception Automation</title>
<p>A substantial portion of the literature documents how LLMs amplify the scale, realism, and adaptability of social engineering attacks, frequently described as the &#x201C;dual-edged sword&#x201D; effect of generative AI [<xref ref-type="bibr" rid="ref-44">44</xref>]. Numerous studies demonstrate that models such as ChatGPT can generate linguistically fluent and context-aware messages that closely mimic professional or personal communication styles, thereby eroding traditional detection cues used in phishing and impersonation attacks [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-45">45</xref>].</p>
<p>Beyond static message generation, interactive and multi-turn deception has emerged as a critical escalation vector. Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] show how adaptive conversational strategies enable sustained manipulation, while Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] and Nawara and kashef [<xref ref-type="bibr" rid="ref-36">36</xref>] highlight risks associated with platform-integrated systems, including chatbots, recommender engines, and customer-service pipelines. Collectively, this body of work indicates a shift from isolated phishing messages toward ecosystem-level influence operations embedded within digital services.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Malware, Exploit Scaffolding, and Code Abuse</title>
<p>A second threat cluster concerns LLM-assisted code generation and its associated dual-use implications. While many studies emphasise the defensive benefits of automated code synthesis, a consistent theme is the risk that generative models may produce insecure logic, reproduce known vulnerabilities, or facilitate exploit construction under ambiguous or adversarial prompts [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. By automating elements of exploit reasoning, LLMs reduce traditional expertise barriers and accelerate attack development cycles.</p>
<p>Although the reviewed corpus contains limited primary experimentation on live malware generation, conceptual and lifecycle-oriented analyses consistently conclude that LLMs can significantly enhance attacker capabilities, particularly when integrated with external knowledge bases or retrieval mechanisms [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>]. These findings justify treating code abuse as a distinct threat class rather than as a marginal extension of social engineering.</p>
</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Runtime Malware Behaviour and OT/ICS Implications</title>
<p>Recent studies increasingly recognise that an exclusive focus on text-centric threats underrepresents execution-stage risks. LLMs can assist in generating adaptive malware components, polymorphic payloads, and environment-aware attack logic that evolves at runtime [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. Such threats exhibit increased persistence and lateral movement potential, particularly when targeting operational technology (OT) and industrial control systems (ICS).</p>
<p>Despite their potentially severe impact, OT and ICS contexts remain comparatively underexplored due to data scarcity, safety constraints, and experimental complexity [<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>]. <xref ref-type="table" rid="table-2">Table 2</xref> explicitly contrasts surface-level attacks with runtime and cyber-physical threats, highlighting a systematic evaluation gap that biases current research toward methodological convenience rather than operational risk.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparative analysis of LLM-driven threat classes: surface-level vs. execution-stage and OT/ICS risks.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dimension</th>
<th>Phishing &#x0026; Prompt-Based Attacks</th>
<th>Runtime Malware Behaviour</th>
<th>OT/ICS Cybersecurity Threats</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Primary Attack Surface</bold></td>
<td>Human users and LLM interfaces</td>
<td>Execution environment, OS, and runtime context</td>
<td>Cyber-physical processes and control logic</td>
</tr>
<tr>
<td><bold>Role of LLM</bold></td>
<td>Message generation, deception, prompt manipulation</td>
<td>Code synthesis, behavioural reasoning, evasion planning</td>
<td>Protocol inference, attack-path reasoning, CPS analysis</td>
</tr>
<tr>
<td><bold>Operational Impact</bold></td>
<td>Credential theft, misinformation, short-term compromise</td>
<td>Persistence, lateral movement, stealthy system compromise</td>
<td>Physical disruption, safety violations, service outages</td>
</tr>
<tr>
<td><bold>Adaptivity at Runtime</bold></td>
<td>Limited after deployment</td>
<td>High (environment-aware, polymorphic execution)</td>
<td>High (context- and process-aware attacks)</td>
</tr>
<tr>
<td><bold>Evaluation Maturity</bold></td>
<td>High (public datasets, benchmarks)</td>
<td>Moderate (curated malware corpora)</td>
<td>Low (limited datasets, safety constraints)</td>
</tr>
<tr>
<td><bold>Key Gap</bold></td>
<td>Overemphasis due to benchmark availability</td>
<td>Limited long-running and real-world evaluation</td>
<td>Severe underrepresentation despite high risk</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2_4">
<label>4.2.4</label>
<title>Model Manipulation, Prompt Injection, and Alignment Evasion</title>
<p>The third threat class encompasses attacks that directly target the behaviour and control mechanisms of large language models themselves. Techniques such as prompt injection, jailbreak attacks, and alignment evasion undermine built-in safety constraints and enable the elicitation of sensitive information or unsafe outputs [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. These risks are further amplified in agentic and retrieval-augmented generation (RAG)-based systems, where injected instructions may propagate across interconnected components, triggering cascading failures and unintended downstream actions [<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>Collectively, the reviewed studies indicate that alignment remains fragile under sustained adversarial interaction, particularly in real-world deployments where LLMs interface with external tools, application programming interfaces (APIs), and untrusted user inputs. These findings underscore the need for robust input validation, isolation mechanisms, and continuous monitoring to mitigate systemic failure modes arising from model manipulation.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>LLM-Enabled Defensive Capabilities</title>
<p>Within the unified taxonomy, the second dimension captures how large language models are operationalised as <italic>defensive instruments</italic> to support cybersecurity monitoring, analysis, and response. In contrast to threat-centric research, this body of work primarily conceptualises LLMs as augmentative components embedded within existing security workflows rather than as fully autonomous decision-makers. Across the 167 reviewed studies, defensive applications consistently cluster into three capability classes: (i) threat intelligence extraction and security operations support, (ii) secure coding, vulnerability assessment, and risk prediction, and (iii) anomaly detection, behavioural modelling, and policy reasoning.</p>
<p>This categorisation reflects an important methodological distinction within the literature. Whereas AIGC-driven threats exploit the generative and adaptive properties of LLMs, defensive systems leverage their semantic reasoning, contextual understanding, and natural language interaction capabilities to reduce analyst workload, enhance interpretability, and support complex decision-making under uncertainty [<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>]. Nevertheless, as discussed in the following sections, defensive effectiveness remains strongly contingent on deployment constraints, governance mechanisms, and the rigor of evaluation protocols.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Threat Intelligence Extraction and Security Operations Support</title>
<p>The most mature and extensively investigated defensive application of LLMs concerns threat intelligence extraction and support for Security Operations Center (SOC) workflows. A recurring theme across the reviewed studies is the use of LLMs to summarise, contextualise, and correlate heterogeneous security artefacts, including logs, alerts, incident reports, and threat intelligence bulletins [<xref ref-type="bibr" rid="ref-50">50</xref>]. Nawara and kashef [<xref ref-type="bibr" rid="ref-36">36</xref>] demonstrate that LLM-driven recommendation and decision-support pipelines can uncover latent relationships among entities, reconstruct attack narratives, and surface actionable insights with minimal analyst prompting. Complementary studies on log analysis and alert triage [<xref ref-type="bibr" rid="ref-51">51</xref>] further show that such systems can reduce cognitive burden by transforming low-level telemetry into structured, high-level explanations.</p>
<p>Survey-oriented investigations of generative AI adoption in enterprise and digital platform contexts [<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-54">54</xref>] consistently position LLMs as effective intermediaries between raw security data and human analysts. Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] further highlight that LLM-enabled automation can accelerate response workflows by generating concise summaries, drafting mitigation recommendations, and interpreting alerts expressed in natural language. Architectures based on retrieval-augmented generation (RAG) explicitly address scalability and knowledge grounding by integrating semantic search with generative reasoning [<xref ref-type="bibr" rid="ref-55">55</xref>].</p>
<p>Despite these advantages, multiple studies caution that SOC-oriented deployments remain vulnerable to hallucinated correlations, overconfident explanations, and trust calibration failures when LLM outputs are consumed without systematic validation [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>]. Consequently, the literature converges on the view that LLMs are most effective as analyst-support tools that augment, rather than replace, expert judgment in operational SOC environments.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Secure Coding, Vulnerability Assessment, and Risk Prediction</title>
<p>A second major class of defensive capability concerns the application of LLMs within secure software development lifecycles. Although code synthesis by generative models is widely recognised as a dual-use capability, a substantial body of work demonstrates that, when appropriately constrained, LLMs can assist in identifying insecure constructs, detecting coding antipatterns, and explaining potential vulnerabilities [<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>]. Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] show that LLMs can reason over code semantics and generate preliminary vulnerability explanations that complement traditional static analysis techniques.</p>
<p>Specialised studies further demonstrate defensive applications in firewall configuration, secure deployment pipelines, and adaptive policy enforcement [<xref ref-type="bibr" rid="ref-59">59</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>]. Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] argue that LLMs can function as intelligent static-analysis assistants, particularly in identifying subtle logic flaws that evade rule-based scanners.</p>
<p>However, Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] emphasise that defensive gains in secure coding are inseparable from governance and lifecycle oversight. In the absence of systematic validation, continuous monitoring, and auditability, LLM-generated recommendations risk propagating insecure patterns at scale. Studies on model trustworthiness and auditing [<xref ref-type="bibr" rid="ref-61">61</xref>] further indicate that higher-level risk prediction tasks-such as vulnerability prioritisation and impact explanation-should be interpreted as decision-support signals rather than authoritative outputs. Collectively, these findings position LLMs as promising yet fragile components within secure development workflows, whose effectiveness depends on rigorous verification and human oversight [<xref ref-type="bibr" rid="ref-62">62</xref>].</p>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Anomaly Detection, Behavioural Modelling, and Policy Reasoning</title>
<p>Beyond intelligence extraction and code analysis, LLMs are increasingly explored for behavioural anomaly detection, system modelling, and policy reasoning. Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] demonstrate that transformer-based models can capture semantic patterns in logs and behavioural traces, enabling the detection of subtle deviations associated with insider threats, fraud, or anomalous system states. Empirical systems such as ChatPhishDetector [<xref ref-type="bibr" rid="ref-63">63</xref>] and LLM-assisted malicious webpage detection frameworks [<xref ref-type="bibr" rid="ref-64">64</xref>] further illustrate the feasibility of applying LLMs to text-rich security contexts.</p>
<p>At a higher level of abstraction, Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] highlight the suitability of LLMs for policy reasoning tasks, including interpreting configuration files, validating compliance requirements, and translating complex security policies into human-readable explanations. Such capabilities support automated auditing and continuous compliance monitoring, particularly in specialised domains such as transport and critical infrastructure systems.</p>
<p>Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] further argue that behavioural modelling and policy reasoning are especially critical in multi-agent or retrieval-augmented systems, where LLMs mediate interactions between user intent, system state, and downstream actions. When properly aligned, these models can detect inconsistencies, enforce safety constraints, and reason about risk propagation across interconnected components [<xref ref-type="bibr" rid="ref-65">65</xref>]. Nevertheless, the literature consistently cautions that such reasoning capabilities must be bounded by explicit control logic to prevent cascading errors or unsafe autonomous behaviour.</p>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Comparative Synthesis of LLM-Enabled Defensive Capabilities</title>
<p>A cross-cutting synthesis of LLM-enabled defensive studies indicates that, despite notable task-level performance gains, defensive effectiveness is fundamentally determined by <italic>how</italic> LLMs are embedded within security workflows rather than by model capability alone. Across threat intelligence support, secure coding, and anomaly detection, LLMs consistently function as <italic>cognitive amplifiers</italic>, enhancing semantic understanding, contextual correlation, and explanation, while remaining dependent on classical detection engines, rule-based controls, and human oversight for operational reliability [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>].</p>
<p>Within SOC and threat intelligence pipelines, LLMs excel at unifying heterogeneous data sources and generating analyst-oriented narratives, yielding measurable reductions in triage time and cognitive workload [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>]. These benefits, however, are counterbalanced by susceptibility to hallucinated correlations and overconfident reasoning, particularly in retrieval-augmented generation (RAG) settings where corpus incompleteness or temporal drift undermines grounding [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>]. As a result, the literature consistently advocates hybrid SOC architectures in which LLMs serve advisory roles and validation layers are mandatory.</p>
<p>In secure coding and vulnerability assessment, LLMs demonstrate strong semantic reasoning capabilities, often outperforming purely syntactic static-analysis tools in explaining logic flaws, insecure dependencies, and exploit preconditions [<xref ref-type="bibr" rid="ref-34">34</xref>]. Nevertheless, empirical evidence also highlights the risk of <italic>scaled insecurity</italic>, whereby unverified LLM recommendations propagate flawed remediation strategies across large codebases [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-57">57</xref>]. Consequently, studies recommend positioning LLMs as assistive auditors integrated with formal verification, testing pipelines, and developer review processes rather than as autonomous security agents [<xref ref-type="bibr" rid="ref-61">61</xref>,<xref ref-type="bibr" rid="ref-62">62</xref>].</p>
<p>For anomaly detection, behavioural modelling, and policy reasoning, LLMs are most effective in text-rich or semantically complex environments&#x2014;such as log analysis, phishing detection, and compliance interpretation&#x2014;where traditional feature-based models exhibit limitations [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>]. However, latency, computational overhead, and sensitivity to input formatting constrain their suitability for real-time detection and control-loop enforcement [<xref ref-type="bibr" rid="ref-64">64</xref>]. In multi-agent and policy-driven systems, LLM-based reasoning enhances interpretability and coordination but introduces new failure modes, including cascading errors and ambiguous policy interpretation, necessitating explicit constraint enforcement and human-in-the-loop safeguards [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>].</p>
<p>Overall, the comparative evidence indicates that LLM-enabled defensive systems are <italic>most robust when deployed as bounded, explainable, and governable components within layered defence architectures</italic>. Performance gains are strongest in decision support, correlation, and explanation tasks, whereas autonomous detection or response remains high risk without rigorous validation, continuous monitoring, and governance oversight. <xref ref-type="table" rid="table-3">Table 3</xref> summarises these trade-offs across representative defensive domains.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparative synthesis of LLM-enabled defensive capabilities.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Defensive Domain</th>
<th>Role of LLMs</th>
<th>Observed Benefits</th>
<th>Limitations and Risks</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Threat Intelligence &#x0026; SOC Support</bold></td>
<td>Semantic summarisation, correlation, analyst decision support</td>
<td>Faster triage; reduced cognitive load; improved cross-source context [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>]</td>
<td>Hallucinations; RAG grounding errors; trust calibration issues [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
</tr>
<tr>
<td><bold>Secure Coding &#x0026; Vulnerability Analysis</bold></td>
<td>Code understanding, vulnerability explanation, remediation guidance</td>
<td>Detects logic flaws; complements static analysis; improves interpretability [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>Insecure fix propagation; lack of formal guarantees; auditability required [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-62">62</xref>]</td>
</tr>
<tr>
<td><bold>Anomaly Detection &#x0026; Behavioural Modelling</bold></td>
<td>Log semantics, phishing detection, behavioural interpretation</td>
<td>Effective for text-rich and semantic anomalies [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td>High latency; compute overhead; fragile to malformed inputs [<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
</tr>
<tr>
<td><bold>Policy Reasoning &#x0026; Compliance Support</bold></td>
<td>Policy interpretation, configuration analysis, audit explanation</td>
<td>Supports explainable compliance and continuous auditing [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Ambiguous policy semantics; inconsistent enforcement risk [<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
</tr>
<tr>
<td><bold>Overall Insight</bold></td>
<td>Bounded cognitive augmentation</td>
<td>Strong reasoning and explanation support</td>
<td>Unsafe as autonomous security decision-makers without oversight [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Security Evaluation, Red-Teaming, and Governance</title>
<p>The third dimension of the unified taxonomy encompasses research focused on evaluating the trustworthiness, robustness, and governance readiness of LLM-enabled cybersecurity systems. Unlike threat- or defence-centric studies, which emphasise capability expansion, this dimension interrogates whether such systems can be <italic>safely, reliably, and legally deployed</italic> in security-sensitive environments. Across the 167 analysed studies, evaluation and governance emerge as indispensable complements to both offensive and defensive research, reflecting a growing consensus that unregulated or inadequately assessed LLM integration may introduce systemic risks that outweigh operational benefits [<xref ref-type="bibr" rid="ref-66">66</xref>&#x2013;<xref ref-type="bibr" rid="ref-68">68</xref>].</p>
<p>Within the taxonomy, this dimension is structured around three interrelated themes: (i) security benchmarking and evaluation frameworks, (ii) red-teaming, misuse detection, and lifecycle risk assessment, and (iii) governance, regulatory alignment, and enforceable controls. Together, these themes address not only <italic>whether</italic> LLMs perform as intended, but also <italic>how</italic> failures emerge, propagate, and can be constrained under real-world operational and regulatory conditions.</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Security Benchmarking and Evaluation Frameworks</title>
<p>A central finding across the reviewed literature is the absence of standardised, security-oriented evaluation frameworks for LLMs. Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] identify multiple evaluation dimensions, including behavioural stress testing, trustworthiness analysis, and safety auditing, but emphasise that prevailing metrics (e.g., accuracy or BLEU-style scores) are insufficient to capture security-relevant failure modes such as context-sensitive hallucinations, brittle reasoning, and inconsistent safety responses. This critique is reinforced by broader surveys of AI evaluation practices, which consistently highlight the lack of unified, comparable, and deployment-relevant assessment standards [<xref ref-type="bibr" rid="ref-69">69</xref>&#x2013;<xref ref-type="bibr" rid="ref-71">71</xref>].</p>
<p>Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] further argue that evaluation pipelines must explicitly assess susceptibility to adversarial prompting, data leakage, insecure code generation, and alignment failures. In the absence of such targeted benchmarks, empirical claims regarding LLM safety and robustness remain difficult to validate or compare across studies. Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] similarly identify robustness evaluation as a critical bottleneck, observing that many proposed defence mechanisms are tested only under benign or narrowly scoped experimental conditions.</p>
<p>Recent investigations of specialised security tools [<xref ref-type="bibr" rid="ref-72">72</xref>] and analyses of LLM application ecosystems further underscore the operational consequences of weak evaluation practices, particularly in settings characterised by rapid deployment and frequent model updates. Across the reviewed literature, evaluation is increasingly conceptualised not as a one-time verification activity but as a continuous, lifecycle-spanning process that underpins responsible deployment, model governance, and risk-aware decision-making [<xref ref-type="bibr" rid="ref-73">73</xref>].</p>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Red-Teaming, Misuse Detection, and Lifecycle Risk Assessment</title>
<p>Red-teaming is widely recognised as a critical mechanism for uncovering latent vulnerabilities in LLM-enabled systems. Although the reviewed corpus contains relatively few large-scale empirical red-teaming studies, numerous analytical and survey-oriented works provide detailed examinations of misuse pathways, adversarial prompting strategies, and system-level failure cascades. Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] present one of the most comprehensive lifecycle-oriented risk assessments, synthesising vulnerabilities across data acquisition, model training, deployment, tool integration, and post-deployment adaptation [<xref ref-type="bibr" rid="ref-74">74</xref>,<xref ref-type="bibr" rid="ref-75">75</xref>].</p>
<p>These analyses consistently reveal that risks often arise not from isolated model behaviour but from interactions between LLMs and surrounding system components. In multi-agent or retrieval-augmented architectures, injected prompts or poisoned contextual data can propagate across subsystems, resulting in cascading failures or unsafe actions [<xref ref-type="bibr" rid="ref-76">76</xref>]. Complementing this perspective, Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] catalogue a broad range of adversarial attack surfaces&#x2014;including jailbreaks, prompt chaining, inference-time manipulation, and information leakage probes&#x2014;that should form the foundation of systematic red-teaming methodologies. Broader surveys of LLM-enabled cyberattacks [<xref ref-type="bibr" rid="ref-43">43</xref>] further reinforce the need for adversarial evaluation strategies that extend beyond single-prompt testing.</p>
<p>Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] additionally emphasise that effective misuse detection requires continuous behavioural monitoring, trust calibration, and post-hoc interpretability mechanisms to identify anomalous or unsafe reasoning trajectories. Collectively, these studies converge on the conclusion that red-teaming must evolve from ad hoc adversarial prompting toward structured, repeatable lifecycle assessments that evaluate model behaviour under uncertainty, interaction, and environmental variation [<xref ref-type="bibr" rid="ref-77">77</xref>].</p>
</sec>
<sec id="s4_4_3">
<label>4.4.3</label>
<title>Governance, Regulatory Alignment, and Enforceable Controls</title>
<p>Governance and regulatory considerations constitute a critical cross-cutting pillar of LLM-enabled cybersecurity research, linking technical evaluation with accountability, compliance, and operational safety. While earlier studies extensively discuss ethical concerns such as bias, transparency, privacy, and misuse [<xref ref-type="bibr" rid="ref-78">78</xref>,<xref ref-type="bibr" rid="ref-79">79</xref>], much of this literature remains principle-driven and insufficiently connected to enforceable regulatory requirements or operational security workflows. More recent critical analyses argue that ethical risk mitigation cannot be decoupled from technical robustness and governance enforcement, particularly when LLMs influence security-critical decisions [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-80">80</xref>,<xref ref-type="bibr" rid="ref-81">81</xref>].</p>
<p>A recurring concern across the reviewed studies is the absence of clearly defined accountability structures for LLM-driven systems deployed in high-stakes contexts, including SOC automation, intrusion detection, and incident response. Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] highlight deficiencies related to explainability, responsibility allocation, and auditability, concluding that governance mechanisms lag behind the pace of operational integration. Similar observations are reported in privacy-sensitive and manipulative-content domains, where ethical shortcomings translate directly into legal, reputational, and operational risks.</p>
<p>From a regulatory perspective, the EU NIS-2 Directive establishes binding requirements for cybersecurity risk management, incident handling, supply-chain security, and organisational accountability. Many failure modes identified in LLM-enabled cybersecurity systems&#x2014;including hallucinated threat attribution, false-positive escalation within SOC pipelines, prompt-injection vulnerabilities in RAG architectures, and unsafe autonomous responses&#x2014;map directly to NIS-2 obligations concerning proportional controls, resilience, and service continuity. Consequently, LLM components that materially influence detection, triage, or response processes fall squarely within the scope of regulated cybersecurity operations.</p>
<p>Complementary guidance issued by the U.S. Cybersecurity and Infrastructure Security Agency (CISA) emphasises defence-in-depth, layered safeguards, and verifiable controls for AI-enabled systems deployed in SOCs and critical infrastructure environments. Consistent with this guidance, the literature increasingly converges on the view that LLMs should function as constrained components within hybrid defence architectures, augmenting classical detection and policy engines rather than acting as autonomous decision-makers [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-82">82</xref>]. Empirical evidence further indicates that hallucinations and reasoning errors can propagate rapidly when LLM outputs are granted excessive operational authority.</p>
<p>Operationalising governance therefore requires translating ethical principles into auditable and enforceable controls. Across the reviewed studies, recurring mechanisms include mandatory human-in-the-loop checkpoints, role-based access control for prompts and policy modification, immutable audit logging, model versioning and change management, and clear separation of duties between detection, reasoning, and response components [<xref ref-type="bibr" rid="ref-83">83</xref>,<xref ref-type="bibr" rid="ref-84">84</xref>]. These controls directly mitigate the operational risks identified in the defensive and evaluation analyses and align with both NIS-2 accountability requirements and CISA deployment guidance.</p>
<p><xref ref-type="table" rid="table-4">Table 4</xref> synthesises these insights by explicitly mapping identified LLM security risks to concrete technical and organisational controls, along with corresponding regulatory or standards-based obligations. By grounding governance in enforceable mechanisms, this taxonomy dimension reframes ethics from aspirational guidance into compliance-ready practice, providing actionable direction for researchers, practitioners, and policymakers.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Mapping of LLM security risks to enforceable controls and regulatory alignment.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Identified Risk</th>
<th>Enforceable/Auditable Control</th>
<th>Regulatory or Standards Alignment</th>
</tr>
</thead>
<tbody>
<tr>
<td>Hallucinated threat attribution or IOC generation</td>
<td>Human-in-the-loop validation; confidence thresholds; analyst override logging</td>
<td>NIS-2 (risk management, incident handling); CISA Secure AI Guidance</td>
</tr>
<tr>
<td>False-positive escalation in SOC automation</td>
<td>Escalation approval workflows; false-positive auditing metrics</td>
<td>NIS-2 (service continuity); CISA SOC Best Practices</td>
</tr>
<tr>
<td>Prompt injection in RAG pipelines</td>
<td>Prompt sanitisation; retrieval whitelisting; immutable context logs</td>
<td>NIS-2 (supply-chain security); CISA AI Supply Chain Risk Management</td>
</tr>
<tr>
<td>Unsafe autonomous response actions</td>
<td>Mandatory human approval; separation of detection and response</td>
<td>NIS-2 (accountability); CISA Zero Trust principles</td>
</tr>
<tr>
<td>Model behaviour drift</td>
<td>Model versioning; regression testing; change management</td>
<td>NIS-2 (governance); ISO/IEC 27001</td>
</tr>
<tr>
<td>Lack of explainability</td>
<td>Rationale logging; post-incident review artefacts</td>
<td>NIS-2 (auditability); CISA IR documentation</td>
</tr>
<tr>
<td>Adversarial misuse of LLM tools</td>
<td>RBAC; usage monitoring; anomaly detection</td>
<td>NIS-2 (access control); CISA insider threat mitigation</td>
</tr>
<tr>
<td>OT/ICS deployment risks</td>
<td>Decision-support-only deployment; fallback to classical controls</td>
<td>NIS-2 (critical infrastructure); CISA ICS advisories</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Methodological Orientation within the Taxonomy</title>
<p>Beyond thematic classification, the proposed taxonomy explicitly incorporates <italic>methodological orientation</italic> as a cross-cutting analytical dimension. This inclusion is essential for interpreting not only which aspects of LLM-enabled cybersecurity are examined, but also how evidence is generated, validated, and generalised. An analysis of the 167 included studies reveals four dominant methodological orientations&#x2014;<bold>empirical</bold>, <bold>system-oriented</bold>, <bold>conceptual/analytical</bold>, and <bold>survey-based</bold>&#x2014;which collectively shape the maturity, reliability, and operational relevance of the field.</p>
<p>These methodological orientations intersect all three substantive taxonomy dimensions (AIGC-driven threats, LLM-enabled defensive capabilities, and security evaluation and governance), thereby influencing the strength of reported claims, the reproducibility of findings, and the feasibility of real-world deployment. Explicitly embedding methodology within the taxonomy prevents the conflation of conceptual insight with empirical validation and enables a more nuanced and evidence-aware synthesis of the literature.</p>
<sec id="s4_5_1">
<label>4.5.1</label>
<title>Empirical Evaluation-Oriented Studies</title>
<p>Empirical studies constitute a substantial portion of the reviewed literature and focus on observing model behaviour, quantifying robustness, and analysing failure modes through controlled experiments or simulations. Representative examples include empirical investigations of hallucination, unsafe code generation, exploit reasoning, and misalignment effects reported by Esmradi et al. [<xref ref-type="bibr" rid="ref-34">34</xref>]. Large-scale comparative assessments of AI-generated code security [<xref ref-type="bibr" rid="ref-85">85</xref>] and experimental evaluations of ChatGPT on log and security datasets [<xref ref-type="bibr" rid="ref-77">77</xref>] further exemplify this orientation.</p>
<p>While empirical studies are indispensable for grounding claims regarding LLM capabilities and risks, a recurring limitation across the corpus is the absence of standardised benchmarks, limited reproducibility, and evaluation under constrained or synthetic conditions. These limitations impede cross-study comparability and restrict generalisation to operational cybersecurity environments. As a result, many empirical findings provide primarily <italic>local validity</italic> rather than system-level assurance, underscoring the need for the unified evaluation frameworks discussed in Dimension III.</p>
</sec>
<sec id="s4_5_2">
<label>4.5.2</label>
<title>System and Architectural Proposals</title>
<p>System-oriented studies represent a second major methodological category and focus on integrating LLMs into operational cybersecurity workflows. These contributions typically propose architectures, pipelines, or decision-support systems that embed LLMs within SOC operations, threat intelligence platforms, or automated response mechanisms. Nawara and kashef [<xref ref-type="bibr" rid="ref-36">36</xref>], for example, present LLM-driven recommendation systems for contextualising threat intelligence and streamlining analyst workflows. Related efforts include retrieval-augmented cybersecurity intelligence frameworks [<xref ref-type="bibr" rid="ref-55">55</xref>], real-time crime detection systems [<xref ref-type="bibr" rid="ref-56">56</xref>], and scalable log analysis pipelines [<xref ref-type="bibr" rid="ref-51">51</xref>].</p>
<p>Although these system-level contributions signal progress toward operational deployment, they are frequently evaluated under constrained conditions, rely on curated or synthetic inputs, or omit lifecycle-level threat modelling. Consequently, many architectural proposals demonstrate functional feasibility without sufficiently addressing robustness, adversarial resilience, or governance constraints. This methodological shortcoming reinforces the importance of coupling system design with rigorous evaluation, red-teaming, and governance analysis.</p>
</sec>
<sec id="s4_5_3">
<label>4.5.3</label>
<title>Conceptual and Analytical Studies</title>
<p>Conceptual and analytical works form a third methodological orientation, offering taxonomies, risk models, and socio-technical analyses that articulate the broader implications of LLM integration within cybersecurity ecosystems [<xref ref-type="bibr" rid="ref-67">67</xref>]. Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] provide a comprehensive lifecycle-oriented risk framework that identifies vulnerabilities spanning data provenance, model training, deployment, tool integration, and governance. Other analytical contributions highlight systemic risks, including the erosion of Zero-Trust assumptions induced by generative AI [<xref ref-type="bibr" rid="ref-75">75</xref>] and the cascading effects associated with large-scale LLM adoption [<xref ref-type="bibr" rid="ref-74">74</xref>].</p>
<p>These studies play a critical role in exposing interdependencies and long-term risks that may not be apparent in isolated empirical evaluations. However, their primary limitation lies in limited empirical grounding, as many conceptual insights remain insufficiently validated through controlled experimentation or real-world case studies. Within the taxonomy, such analyses therefore function primarily as instruments of <italic>risk foresight</italic> rather than as evidence of deployable solutions.</p>
</sec>
<sec id="s4_5_4">
<label>4.5.4</label>
<title>Surveys and Meta-Analyses</title>
<p>Surveys and meta-analyses constitute the fourth methodological category, synthesising research trends across generative AI, cybersecurity threats, defensive mechanisms, and governance considerations. Foundational surveys by Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] and Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] provide broad overviews of AIGC risks, defensive opportunities, and ethical challenges. Additional systematic and narrative reviews further contextualise LLM applications in cybersecurity [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>], their dual-use implications in network security [<xref ref-type="bibr" rid="ref-86">86</xref>], and recent advancements in generative AI [<xref ref-type="bibr" rid="ref-71">71</xref>].</p>
<p>While surveys play a crucial role in consolidating fragmented research, many existing reviews rely predominantly on narrative synthesis and lack systematic protocols such as PRISMA. This limits transparency, reproducibility, and comparability across surveys. The present work addresses this limitation by embedding methodological orientation directly within a PRISMA-guided taxonomy, enabling structured cross-comparison beyond descriptive aggregation.</p>
</sec>
<sec id="s4_5_5">
<label>4.5.5</label>
<title>Methodological Imbalances and Implications</title>
<p>Taken together, the distribution of methodological orientations reveals a research landscape that is broad yet uneven. Empirical studies frequently lack unified benchmarks and stress-testing protocols; system-oriented proposals often under-evaluate adversarial and governance risks; conceptual analyses are weakly anchored in experimental validation; and surveys expose inconsistencies in terminology, threat definitions, and evaluation practices. These imbalances directly influence the maturity and reliability of conclusions drawn across all three substantive taxonomy dimensions.</p>
<p>Beyond methodological categorisation, the distribution of studies across the taxonomy (<xref ref-type="table" rid="table-5">Table 5</xref>) highlights several longitudinal trends between 2022 and 2025. Research on AIGC-driven threats remains dominant, particularly in social engineering and adversarial prompting [<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. Defensive applications&#x2014;especially threat intelligence extraction and SOC automation&#x2014;demonstrate increasing operational focus [<xref ref-type="bibr" rid="ref-36">36</xref>], yet remain unevenly validated. Notably, a growing concentration of studies addresses security evaluation, red-teaming, and governance, reflecting increasing recognition that trustworthiness and accountability are decisive factors for real-world deployment [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>].</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comparative taxonomy-based classification of included studies with methodological and operational synthesis.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th>Sub&#x2013;Category</th>
<th>Core Technical Focus</th>
<th>Methodological Orientation</th>
<th>Strengths/<break/>Advantages</th>
<th>Limitations/<break/>Failure Modes</th>
<th>Primary Application Context</th>
<th>Representative<break/> Studies</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="8"><bold>Dimension I: AIGC-Driven Cyber Threats</bold></td>
</tr>
<tr>
<td>Social Engineering &#x0026; Deception Automation</td>
<td>LLM-generated phishing and impersonation</td>
<td>High-fidelity multilingual phishing, role-adaptive deception</td>
<td>Empirical/<break/>System/Survey</td>
<td>Scalable social engineering, linguistic realism</td>
<td>Prompt sensitivity, cultural bias, rapid counter-adaptation</td>
<td>Email security, messaging platforms</td>
<td>T1 &#x003D; [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-45">45</xref>,<xref ref-type="bibr" rid="ref-92">92</xref>&#x2013;<xref ref-type="bibr" rid="ref-94">94</xref>,<xref ref-type="bibr" rid="ref-98">98</xref>&#x2013;<xref ref-type="bibr" rid="ref-104">104</xref>]</td>
</tr>
<tr>
<td>Malware &#x0026; Exploit Generation Risks</td>
<td>Automated malicious code synthesis</td>
<td>Exploit reasoning, polymorphic malware generation</td>
<td>Empirical/<break/>Conceptual/Survey</td>
<td>Accelerated exploit development</td>
<td>Limited runtime grounding, detectable artefacts</td>
<td>Malware pipelines, red-team tooling</td>
<td>T2 &#x003D; [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>&#x2013;<xref ref-type="bibr" rid="ref-115">115</xref>]</td>
</tr>
<tr>
<td>Adversarial Prompting, Jailbreaks &#x0026; Injection</td>
<td>Alignment bypass and prompt manipulation</td>
<td>Role-based jailbreaks, multi-step misalignment</td>
<td>Empirical/<break/>Conceptual/Survey</td>
<td>Reveals alignment weaknesses</td>
<td>Brittle defences, prompt overfitting</td>
<td>LLM safety testing, agentic systems</td>
<td>T3 &#x003D; [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<break/><xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>,<break/><xref ref-type="bibr" rid="ref-116">116</xref>&#x2013;<xref ref-type="bibr" rid="ref-121">121</xref>]</td>
</tr>
<tr>
<td>Disinformation &#x0026; Manipulation</td>
<td>Synthetic misinformation and propaganda</td>
<td>Narrative steering, watermark evasion</td>
<td>Empirical/<break/>Analytical/Conceptual</td>
<td>High-scale influence operations</td>
<td>Attribution difficulty, detection lag</td>
<td>Media platforms, information ecosystems</td>
<td>T4 &#x003D; [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-101">101</xref>,<break/><xref ref-type="bibr" rid="ref-122">122</xref>&#x2013;<xref ref-type="bibr" rid="ref-127">127</xref>]</td>
</tr>
<tr>
<td align="center" colspan="8"><bold>Dimension II: LLM-Enabled Defensive Capabilities</bold></td>
</tr>
<tr>
<td>Intrusion, Malware &#x0026; Anomaly Detection</td>
<td>Semantic and behavioural analysis</td>
<td>Log reasoning, malware trace inference</td>
<td>Empirical/<break/>System/Survey</td>
<td>Context-aware detection</td>
<td>False-positive escalation, dataset bias</td>
<td>Enterprise IDS, SOC monitoring</td>
<td>D1 &#x003D; [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<break/><xref ref-type="bibr" rid="ref-63">63</xref>,<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-102">102</xref>,<break/><xref ref-type="bibr" rid="ref-103">103</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>,<xref ref-type="bibr" rid="ref-106">106</xref>,<xref ref-type="bibr" rid="ref-108">108</xref>,<xref ref-type="bibr" rid="ref-110">110</xref>,<xref ref-type="bibr" rid="ref-112">112</xref>&#x2013;<xref ref-type="bibr" rid="ref-114">114</xref>,<xref ref-type="bibr" rid="ref-128">128</xref>,<xref ref-type="bibr" rid="ref-129">129</xref>]</td>
</tr>
<tr>
<td>Threat Intelligence &#x0026; SOC Automation</td>
<td>CTI extraction and alert triage</td>
<td>RAG-driven intelligence, summarisation</td>
<td>System/<break/>Empirical/Survey</td>
<td>Reduced analyst workload</td>
<td>Hallucinated correlations, trust calibration</td>
<td>SOC co-pilots, incident response</td>
<td>D2 &#x003D; [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<break/><xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>,<break/><xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>,<xref ref-type="bibr" rid="ref-94">94</xref>,<xref ref-type="bibr" rid="ref-124">124</xref>,<xref ref-type="bibr" rid="ref-130">130</xref>&#x2013;<xref ref-type="bibr" rid="ref-132">132</xref>]</td>
</tr>
<tr>
<td>Secure Coding &#x0026; Vulnerability Detection</td>
<td>Code analysis and patch generation</td>
<td>Semantic reasoning, structured prompting</td>
<td>Empirical/<break/>System/Analytical</td>
<td>Improved vulnerability localisation</td>
<td>Patch correctness uncertainty</td>
<td>Secure SDLC, DevSecOps</td>
<td>D3 &#x003D; [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<break/><xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-59">59</xref>,<xref ref-type="bibr" rid="ref-62">62</xref>,<xref ref-type="bibr" rid="ref-93">93</xref>,<xref ref-type="bibr" rid="ref-107">107</xref>,<xref ref-type="bibr" rid="ref-112">112</xref>,<xref ref-type="bibr" rid="ref-117">117</xref>,<break/><xref ref-type="bibr" rid="ref-133">133</xref>,<xref ref-type="bibr" rid="ref-134">134</xref>]</td>
</tr>
<tr>
<td>LLM Agents &#x0026; Automated Response Systems</td>
<td>Multi-agent autonomous defence</td>
<td>Task delegation, chain-of-thought orchestration</td>
<td>System/<break/>Conceptual/Survey</td>
<td>End-to-end workflow automation</td>
<td>Coordination overhead, cascading failures</td>
<td>Autonomous SOC, large-scale networks</td>
<td>D4 &#x003D; [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>,<break/><xref ref-type="bibr" rid="ref-65">65</xref>,<xref ref-type="bibr" rid="ref-129">129</xref>&#x2013;<xref ref-type="bibr" rid="ref-131">131</xref>,<xref ref-type="bibr" rid="ref-135">135</xref>,<xref ref-type="bibr" rid="ref-136">136</xref>]</td>
</tr>
<tr>
<td align="center" colspan="8"><bold>Dimension III: Security Evaluation, Red-Teaming &#x0026; Governance</bold></td>
</tr>
<tr>
<td>Security Benchmarks &#x0026; Evaluation Datasets</td>
<td>LLM security testing datasets</td>
<td>Robustness scoring, cross-model comparison</td>
<td>Empirical/<break/>Survey</td>
<td>Standardised evaluation</td>
<td>Synthetic bias, limited realism</td>
<td>Benchmarking, academic evaluation</td>
<td>E1 &#x003D; [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-70">70</xref>,<xref ref-type="bibr" rid="ref-72">72</xref>,<xref ref-type="bibr" rid="ref-137">137</xref>,<xref ref-type="bibr" rid="ref-138">138</xref>]</td>
</tr>
<tr>
<td>Red-Teaming &#x0026; Alignment Stress-Testing</td>
<td>Adversarial probing</td>
<td>Multi-turn jailbreak, role-based attacks</td>
<td>Empirical/<break/>Analytical/Conceptual</td>
<td>Failure-mode discovery</td>
<td>Coverage gaps, non-repeatability</td>
<td>LLM safety assurance</td>
<td>E2 &#x003D; [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>&#x2013;<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-43">43</xref>,<xref ref-type="bibr" rid="ref-76">76</xref>,<xref ref-type="bibr" rid="ref-77">77</xref>,<xref ref-type="bibr" rid="ref-85">85</xref>,<xref ref-type="bibr" rid="ref-118">118</xref>,<xref ref-type="bibr" rid="ref-139">139</xref>]</td>
</tr>
<tr>
<td>Risk, Governance &#x0026; Ethical Frameworks</td>
<td>AI risk and compliance</td>
<td>Auditability, lifecycle governance</td>
<td>Conceptual/<break/>Analytical/Survey</td>
<td>Policy alignment</td>
<td>Limited enforceability</td>
<td>Regulatory governance</td>
<td>E3 &#x003D; [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<break/><xref ref-type="bibr" rid="ref-80">80</xref>,<xref ref-type="bibr" rid="ref-81">81</xref>,<xref ref-type="bibr" rid="ref-95">95</xref>,<xref ref-type="bibr" rid="ref-140">140</xref>&#x2013;<xref ref-type="bibr" rid="ref-144">144</xref>]</td>
</tr>
<tr>
<td>Conceptual Foundations &#x0026; Surveys</td>
<td>Taxonomies and surveys</td>
<td>Foundational perspectives</td>
<td>Survey/<break/>Conceptual</td>
<td>Holistic understanding</td>
<td>Limited empirical grounding</td>
<td>Research synthesis</td>
<td>E4 &#x003D; [<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>,<break/><xref ref-type="bibr" rid="ref-66">66</xref>,<xref ref-type="bibr" rid="ref-79">79</xref>,<xref ref-type="bibr" rid="ref-82">82</xref>,<xref ref-type="bibr" rid="ref-89">89</xref>,<xref ref-type="bibr" rid="ref-92">92</xref>,<break/><xref ref-type="bibr" rid="ref-96">96</xref>,<xref ref-type="bibr" rid="ref-109">109</xref>]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-5fn1" fn-type="other">
<p>Note: T1&#x2013;T4, D1&#x2013;D4, and E1&#x2013;E4 denote grouped citation clusters mapped to the unified taxonomy.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Overall, explicitly positioning methodological orientation within the taxonomy transforms it from a descriptive classification into a critical analytical framework. By revealing how evidence is produced&#x2014;and where it is systematically weak&#x2014;this dimension provides a foundation for identifying research gaps, guiding standardisation efforts, and supporting evidence-based advancement of LLM-driven cybersecurity intelligence.</p>
<p><xref ref-type="table" rid="table-5">Table 5</xref> consolidates all 167 studies within the unified taxonomy, linking thematic focus, methodological orientation, strengths, and limitations. This synthesis underpins the comparative and longitudinal analyses presented in subsequent sections.</p>

</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>AIGC-Driven Threats Enabled by LLMs</title>
<p>Building on the unified threat&#x2013;defence&#x2013;evaluation taxonomy introduced in <xref ref-type="sec" rid="s4">Section 4</xref>, this section examines how large language models (LLMs) operate as <italic>offensive enablers</italic> within contemporary cyber threat ecosystems. Across the 167 studies synthesised in this review, AIGC-enabled attacks consistently emerge not as incremental extensions of existing techniques, but as a qualitative escalation in adversarial capability. LLMs fundamentally alter attacker economics by lowering expertise barriers, automating cognitively complex tasks, and enabling scalable, adaptive, and context-aware malicious behaviour [<xref ref-type="bibr" rid="ref-87">87</xref>&#x2013;<xref ref-type="bibr" rid="ref-89">89</xref>].</p>
<p>Within the proposed taxonomy, AIGC-driven threats cluster into four dominant and recurrent categories: (i) social engineering and multimodal deception, (ii) malware generation, exploit scaffolding, and runtime behaviour, (iii) prompt injection, jailbreaks, and alignment evasion, and (iv) misinformation, disinformation, and influence operations. These categories are derived inductively through cross-study coding and reflect convergent patterns observed across empirical evaluations, red-teaming analyses, and large-scale surveys. Collectively, they expose the inherent dual-use tension of LLMs, wherein the same generative, reasoning, and contextualisation capabilities that enable defensive innovation are systematically repurposed for adversarial misuse [<xref ref-type="bibr" rid="ref-42">42</xref>,<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-90">90</xref>].</p>
<sec id="s5_1">
<label>5.1</label>
<title>Social Engineering and Multimodal Deception</title>
<p>Social engineering represents the most mature and empirically substantiated class of AIGC-driven threats. A substantial body of literature demonstrates that LLMs enable scalable, highly personalised phishing, impersonation, and deception campaigns characterised by linguistic coherence, contextual sensitivity, and adaptive tone control previously associated with skilled human operators [<xref ref-type="bibr" rid="ref-45">45</xref>,<xref ref-type="bibr" rid="ref-91">91</xref>]. Unlike template-based automation, LLM-driven pipelines dynamically tailor content using inferred organisational roles, conversational context, and cultural cues, thereby significantly increasing attack credibility [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-92">92</xref>].</p>
<p>Empirical evaluations reported in [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-93">93</xref>,<xref ref-type="bibr" rid="ref-94">94</xref>] indicate that LLM-generated spear-phishing and business email compromise (BEC) messages consistently degrade the effectiveness of signature-based filters and stylometric detection techniques. Comparative studies further reveal that semantic diversity, multilingual generation, and adversarial paraphrasing undermine feature-driven classifiers, exacerbating the asymmetry between attacker adaptability and defender rigidity [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-95">95</xref>].</p>
<p>Beyond text-only attacks, recent work documents a marked shift toward <italic>multimodal deception</italic>, wherein adversaries exploit systems capable of jointly reasoning over text, images, code, and structured artefacts. Such cross-modal manipulation produces inconsistencies that evade unimodal detection pipelines [<xref ref-type="bibr" rid="ref-96">96</xref>,<xref ref-type="bibr" rid="ref-97">97</xref>]. Collectively, these findings indicate that social engineering has evolved from isolated phishing attempts into adaptive, multimodal deception workflows that challenge foundational assumptions of existing detection systems.</p>
 
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Malware, Exploit Scaffolding, and Runtime Behaviour</title>
<p>The second major threat category encompasses LLM-assisted malware generation, exploit scaffolding, and execution-stage behaviour. Across the reviewed corpus, multiple studies demonstrate that LLMs can generate functional malware components, exploit templates, and obfuscation patterns, particularly when safety mechanisms are bypassed through adversarial prompting or indirect instruction leakage [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-109">109</xref>]. Although many generated artefacts are not immediately deployable, they substantially accelerate exploit development cycles and reduce the expertise required for iterative attack refinement [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-145">145</xref>].</p>
<p>Several studies [<xref ref-type="bibr" rid="ref-108">108</xref>,<xref ref-type="bibr" rid="ref-110">110</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>] highlight the dual-use risks inherent to LLM-assisted code generation, showing that models designed for benign development support can emit vulnerable or exploitable logic under ambiguous or underspecified prompts. The integration of structured cybersecurity knowledge bases further amplifies this threat by enabling attackers to rapidly query, adapt, and weaponise technical content at scale [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-104">104</xref>].</p>
<p>A critical limitation identified across this literature is the systematic underrepresentation of <italic>runtime behaviour</italic>. Most studies validate exploit generation at the code-fragment or proof-of-concept level, without examining execution dynamics, environmental dependencies, or long-term behavioural adaptation [<xref ref-type="bibr" rid="ref-38">38</xref>]. Consequently, current evaluations likely underestimate the persistence, lateral movement, and evasion capabilities enabled by LLM-assisted malware, underscoring the need for execution-aware threat modelling and system-level validation.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Prompt Injection, Jailbreaks, and Alignment Evasion</title>
<p>A third and rapidly expanding threat class targets LLMs directly through adversarial prompting, jailbreaks, and prompt injection. Extensive empirical evidence demonstrates that contemporary models remain highly susceptible to role-play manipulation, instruction obfuscation, multi-turn coercion, and indirect prompt injection attacks [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>]. Notably, several studies suggest that larger and more capable models may exhibit increased vulnerability due to richer representational capacity and broader behavioural generalisation [<xref ref-type="bibr" rid="ref-116">116</xref>,<xref ref-type="bibr" rid="ref-142">142</xref>].</p>
<p>Importantly, these vulnerabilities extend beyond standalone models to integrated systems. Research on retrieval-augmented generation (RAG) and multi-agent architectures shows that injected prompts can propagate across components, triggering unsafe tool invocation, policy override, or cascading reasoning failures [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-118">118</xref>,<xref ref-type="bibr" rid="ref-146">146</xref>]. Across the surveyed literature, no single mitigation strategy&#x2014;including prompt sanitisation, refusal tuning, or system prompt hardening&#x2014;consistently defends against the full spectrum of adversarial prompting techniques. This exposes a structural fragility in current alignment approaches and motivates lifecycle-aware evaluation, system-level isolation, and governance-driven safeguards [<xref ref-type="bibr" rid="ref-147">147</xref>,<xref ref-type="bibr" rid="ref-148">148</xref>].</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Misinformation and Influence Operations</title>
<p>LLM-enabled misinformation and influence operations constitute the fourth major threat category and represent one of the most societally consequential forms of AIGC-driven attack. Studies consistently demonstrate that LLMs can generate persuasive, culturally adaptive, and emotionally framed narratives at scale that are often indistinguishable from human-authored content [<xref ref-type="bibr" rid="ref-124">124</xref>,<xref ref-type="bibr" rid="ref-149">149</xref>]. Comparative analyses show that LLM-driven campaigns outperform traditional disinformation efforts in adaptability, multilingual reach, and narrative coherence [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-150">150</xref>].</p>
<p>A recurring observation across this literature is the fragility of existing detection and attribution mechanisms. Watermarking, stylometric analysis, and classifier-based approaches degrade substantially under paraphrasing, translation, and multi-model transformation pipelines [<xref ref-type="bibr" rid="ref-125">125</xref>,<xref ref-type="bibr" rid="ref-149">149</xref>]. Moreover, algorithmic recommendation systems may unintentionally amplify harmful narratives, blurring the boundary between deliberate influence operations and emergent systemic manipulation [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-127">127</xref>]. Despite their high potential impact, longitudinal societal studies and deployment-scale evaluations remain limited, constraining understanding of sustained real-world effects.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Synthesis of Threat Findings</title>
<p>Synthesising evidence across all four threat categories reveals a consistent and concerning trajectory: AIGC-enabled threats are advancing more rapidly than corresponding defensive countermeasures. In social engineering, semantic diversity and contextual realism undermine static detection heuristics. In malware generation, LLMs accelerate exploit development while enabling adaptive runtime behaviour. In adversarial prompting, alignment mechanisms remain brittle under sustained or system-level attacks. In misinformation, generative realism and scale overwhelm existing attribution and moderation strategies.</p>
<p>Three cross-cutting insights emerge from the reviewed corpus:
<list list-type="bullet">
<list-item>
<p><bold>Escalating attacker capability:</bold> LLMs substantially reduce the expertise and resources required to conduct complex cyberattacks, effectively democratising sophisticated offensive techniques.</p></list-item>
<list-item>
<p><bold>Evaluation realism gap:</bold> Heavy reliance on synthetic datasets and constrained experimental settings obscures the true operational severity of AIGC-driven threats.</p></list-item>
<list-item>
<p><bold>Absence of comprehensive mitigation:</bold> No existing defensive framework robustly addresses the full spectrum of LLM-enabled attack vectors, particularly in multimodal, agentic, and continuously evolving systems.</p></list-item>
</list></p>
<p>These findings underscore the need for execution-aware threat modelling, lifecycle-oriented evaluation, and defence architectures that explicitly account for the dual-use nature of generative AI. <xref ref-type="table" rid="table-6">Tables 6</xref> and <xref ref-type="table" rid="table-7">7</xref> consolidate comparative evidence across threat categories and representative studies, demonstrating broad convergence on a central conclusion: AIGC-driven threats represent a systemic and accelerating challenge that existing security controls are structurally ill-equipped to contain.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparative synthesis of AIGC-driven threats enabled by LLMs.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Threat Category</th>
<th>Key References</th>
<th>Observed Capabilities</th>
<th>Empirical Gaps and Constraints</th>
<th>Severity</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Phishing &#x0026; Social Engineering</bold></td>
<td>[<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-100">100</xref>]</td>
<td>Highly personalised and multilingual phishing; adaptive tone and narrative framing; strong evasion of signature-based filters</td>
<td>Predominantly synthetic datasets; weak longitudinal analysis; detector degradation under paraphrasing and cross-lingual transfer</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>Malware &#x0026; Exploit Generation</bold></td>
<td>[<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>]</td>
<td>Exploit scaffolding and polymorphic code generation; semantic obfuscation; assistance in malware reasoning and triage</td>
<td>Limited end-to-end validation; runtime and environmental constraints often ignored; inconsistent safety enforcement</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>Jailbreaks &#x0026; Prompt Injection</bold></td>
<td>[<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-116">116</xref>,<xref ref-type="bibr" rid="ref-117">117</xref>]</td>
<td>Reliable bypass of alignment controls; indirect prompt injection in RAG and agentic systems; cascading policy failures</td>
<td>No robust universal defence; benchmarks ignore lifecycle and update dynamics; defences brittle under adaptation</td>
<td><bold>Critical</bold></td>
</tr>
<tr>
<td><bold>Misinformation &#x0026; Influence Operations</bold></td>
<td>[<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-124">124</xref>,<xref ref-type="bibr" rid="ref-149">149</xref>]</td>
<td>Scalable persuasive content; cultural and emotional adaptation; cross-lingual narrative transfer</td>
<td>Scarce real-world validation; fragile watermarking; unreliable attribution; unclear long-term impact</td>
<td><bold>High&#x2013;Critical</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparative analysis of AIGC-driven threat studies with methodological and operational insights.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Study</th>
<th>Threat Subcategory</th>
<th>Model/Setting</th>
<th>Methodological Distinction</th>
<th>Demonstrated Capability/Advantage</th>
<th>Observed Limitations/Failure Modes</th>
<th>Applicable Threat Scenarios</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="7"><bold>A. Automated Phishing and Social Engineering</bold></td>
</tr>
<tr>
<td>Opara et al. [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>Phishing Generation</td>
<td>GPT-family models</td>
<td>Stylometry-aware text generation</td>
<td>Bypasses enterprise spam filters via linguistic mimicry</td>
<td>Limited multilingual and cross-cultural testing</td>
<td>Corporate email phishing</td>
</tr>
<tr>
<td>Schmitt and Flechais [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>Deception Narratives</td>
<td>Instruction-tuned LLMs</td>
<td>Narrative-level persuasion modelling</td>
<td>High deception realism in controlled human studies</td>
<td>No deployment-scale or adversarial testing</td>
<td>Psychological manipulation campaigns</td>
</tr>
<tr>
<td>Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>Threat Survey</td>
<td>Review</td>
<td>Cross-domain threat aggregation</td>
<td>Highlights lowered entry barrier for attackers</td>
<td>Lacks empirical validation</td>
<td>Strategic threat assessment</td>
</tr>
<tr>
<td>De Queiroz [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>Targeted Phishing</td>
<td>LLM simulations</td>
<td>High-volume personalised phishing</td>
<td>Cost-efficient, scalable attack generation</td>
<td>Text-only focus; no multimodal cues</td>
<td>Mass phishing operations</td>
</tr>
<tr>
<td>Nawara and Kashef [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>Systemic Manipulation</td>
<td>LLM recommender systems</td>
<td>Feedback-loop analysis</td>
<td>Demonstrates narrative amplification risks</td>
<td>Primarily conceptual analysis</td>
<td>Algorithmic content manipulation</td>
</tr>
<tr>
<td>Aseeri and Bohacek [<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
<td>Spear-Phishing</td>
<td>Contextual prompting</td>
<td>Social-media-aware attack modelling</td>
<td>Higher predicted victim engagement</td>
<td>No real victim interaction data</td>
<td>Targeted executive phishing</td>
</tr>
<tr>
<td>Pham et al. [<xref ref-type="bibr" rid="ref-98">98</xref>]</td>
<td>Enterprise Phishing</td>
<td>Fine-tuned LLMs</td>
<td>Role-aware training</td>
<td>Outperforms human templates</td>
<td>Domain-specific dataset bias</td>
<td>Corporate SOC testing</td>
</tr>
<tr>
<td>Koide et al. [<xref ref-type="bibr" rid="ref-99">99</xref>]</td>
<td>Filter Evasion</td>
<td>LLM paraphrasing</td>
<td>Lexical diversity exploitation</td>
<td>Defeats heuristic spam filters</td>
<td>Semantic risk not evaluated</td>
<td>Spam filter evasion</td>
</tr>
<tr>
<td align="center" colspan="7"><bold>B. Malware and Exploit Generation</bold></td>
</tr>
<tr>
<td>Yamin et al. [<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td>Malware Generation</td>
<td>GPT models</td>
<td>Code synthesis with obfuscation</td>
<td>Functional malware-like behaviour</td>
<td>Incomplete payload execution</td>
<td>Proof-of-concept malware</td>
</tr>
<tr>
<td>Shandilya et al. [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>Exploit Generation</td>
<td>GPT-3/4</td>
<td>Exploit template synthesis</td>
<td>Valid exploit fragments generated</td>
<td>Fails on complex multi-stage attacks</td>
<td>Exploit prototyping</td>
</tr>
<tr>
<td>Grov et al. [<xref ref-type="bibr" rid="ref-115">115</xref>]</td>
<td>Dual-Use Assessment</td>
<td>LLM code assistants</td>
<td>Offensive-defensive overlap analysis</td>
<td>Shows attacker capability amplification</td>
<td>No empirical exploits</td>
<td>Capability risk analysis</td>
</tr>
<tr>
<td>Iturbe et al. [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>Polymorphic Malware</td>
<td>Hybrid LLM-code generator</td>
<td>Variant diversification</td>
<td>Signature-based detection evasion</td>
<td>No runtime behaviour testing</td>
<td>Evasion-focused malware</td>
</tr>
<tr>
<td>Huynh et al. [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Behavioural Malware</td>
<td>Transformer models</td>
<td>Trace-level semantic modelling</td>
<td>Improved detection accuracy</td>
<td>High computational overhead</td>
<td>Enterprise malware detection</td>
</tr>
<tr>
<td>Che et al. [<xref ref-type="bibr" rid="ref-106">106</xref>]</td>
<td>Domain Adaptation</td>
<td>Adapted LLMs</td>
<td>Cross-family generalisation</td>
<td>Mitigates dataset shift</td>
<td>High retraining cost</td>
<td>Multi-family malware</td>
</tr>
<tr>
<td align="center" colspan="7"><bold>C. Adversarial Prompting, Jailbreaks, and Prompt Injection</bold></td>
</tr>
<tr>
<td>Yao et al. [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>Jailbreak Attacks</td>
<td>General LLMs</td>
<td>Role-play exploitation</td>
<td>Widespread vulnerability identified</td>
<td>Qualitative scope only</td>
<td>LLM safety testing</td>
</tr>
<tr>
<td>Maity and Arora [<xref ref-type="bibr" rid="ref-116">116</xref>]</td>
<td>Automated Jailbreaking</td>
<td>GPT variants</td>
<td>Prompt search automation</td>
<td>High success rates</td>
<td>Limited model diversity</td>
<td>Red-teaming pipelines</td>
</tr>
<tr>
<td>Villa et al. [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>Large-Scale Red-Teaming</td>
<td>Multiple LLMs</td>
<td>Black-box robustness testing</td>
<td>Model size correlates with risk</td>
<td>Text-only analysis</td>
<td>Model risk evaluation</td>
</tr>
<tr>
<td>De et al. [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>Prompt Injection</td>
<td>RAG systems</td>
<td>System prompt override</td>
<td>Chain-level vulnerabilities exposed</td>
<td>No large-scale benchmarks</td>
<td>Agentic pipelines</td>
</tr>
<tr>
<td>Huang et al. [<xref ref-type="bibr" rid="ref-119">119</xref>]</td>
<td>Multimodal Jailbreaks</td>
<td>Vision-LLMs</td>
<td>Cross-modal exploitation</td>
<td>Vision-text alignment bypass</td>
<td>Defences underdeveloped</td>
<td>Multimodal assistants</td>
</tr>
<tr>
<td align="center" colspan="7"><bold>D. Misinformation, Disinformation, and Influence Operations</bold></td>
</tr>
<tr>
<td>Schmitt and Flechais [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>Narrative Manipulation</td>
<td>LLM narratives</td>
<td>Strategic framing analysis</td>
<td>High persuasion realism</td>
<td>No longitudinal impact study</td>
<td>Influence campaigns</td>
</tr>
<tr>
<td>Harris et al. [<xref ref-type="bibr" rid="ref-123">123</xref>]</td>
<td>Fake News Generation</td>
<td>ChatGPT-class models</td>
<td>Human indistinguishability tests</td>
<td>High deception credibility</td>
<td>Limited dataset diversity</td>
<td>News manipulation</td>
</tr>
<tr>
<td>Valdez et al. [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>Cross-Cultural Disinfo</td>
<td>Multilingual LLMs</td>
<td>Cultural adaptation analysis</td>
<td>Effective cross-lingual transfer</td>
<td>Few languages evaluated</td>
<td>Global misinformation</td>
</tr>
<tr>
<td>Doumanas et al. [<xref ref-type="bibr" rid="ref-149">149</xref>]</td>
<td>Watermark Fragility</td>
<td>Survey</td>
<td>Regeneration-based analysis</td>
<td>Watermarks easily removed</td>
<td>No system-level evaluation</td>
<td>Attribution evasion</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>LLM-Based Defensive Capabilities</title>
<p>Anchored in the second dimension of the unified threat&#x2013;defence&#x2013;evaluation taxonomy, this section synthesises how large language models (LLMs) are operationalised as <italic>defensive instruments</italic> across contemporary cybersecurity pipelines. In contrast to AIGC-driven threats (<xref ref-type="sec" rid="s5">Section 5</xref>), defensive studies overwhelmingly conceptualise LLMs not as autonomous detectors, but as <italic>cognitive and semantic augmentation layers</italic> embedded within existing security architectures. An analysis of the 167 reviewed studies published between 2022 and 2025 reveals four dominant defensive domains: (i) intrusion, malware, and anomaly detection; (ii) threat intelligence extraction and Security Operations Center (SOC) automation; (iii) vulnerability detection, secure coding, and automated repair; and (iv) hybrid and multi-agent architectures for investigation and response.</p>
<p>Across these domains, LLMs primarily contribute through enhanced <bold>semantic reasoning</bold>, cross-context inference, and explainable decision support rather than improvements in raw detection accuracy [<xref ref-type="bibr" rid="ref-151">151</xref>,<xref ref-type="bibr" rid="ref-152">152</xref>]. At the same time, the literature consistently documents structural limitations&#x2014;including hallucination-induced errors, evaluation fragility, adversarial susceptibility, computational overhead, and governance gaps&#x2014;that constrain safe deployment in high-assurance operational environments. These tensions motivate a taxonomy-aware and failure-conscious analysis, rather than a purely capability-centric narrative.</p>
<sec id="s6_1">
<label>6.1</label>
<title>General Cybersecurity Applications vs. IDS-Specific Architectures</title>
<p>A recurring source of conceptual ambiguity in the existing literature is the frequent conflation of <italic>general LLM-enabled cybersecurity applications</italic> with <italic>IDS-specific LLM architectures</italic>. This distinction is not merely terminological; it is foundational for correctly interpreting reported performance gains, assessing deployment readiness, and understanding risk exposure. The absence of a clear separation between these paradigms has contributed to overgeneralised claims regarding the suitability of LLMs for real-time intrusion detection and operational cyber defence.</p>
<p>General LLM-based cybersecurity applications primarily function as <italic>analyst-facing cognitive assistants</italic>. Representative tasks include threat report summarisation, indicator-of-compromise (IOC) extraction, incident timeline reconstruction, policy and configuration interpretation, and SOC co-pilot support [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>]. Such systems typically operate on curated or semi-structured inputs, tolerate moderate inference latency, and are evaluated using task-centric or qualitative metrics such as extraction accuracy, reasoning coherence, or analyst productivity. Their principal contribution lies in reducing cognitive burden, improving contextual understanding, and accelerating sense-making, rather than in executing primary detection or enforcement decisions.</p>
<p>In contrast, IDS-specific LLM architectures are embedded directly within intrusion detection pipelines and are subject to substantially stricter operational constraints. As synthesised in <xref ref-type="sec" rid="s6_4">Sections 6.4</xref> and <xref ref-type="sec" rid="s6_5">6.5</xref>, LLMs in IDS contexts do not replace conventional signature-based, statistical, or machine-learning detectors. Instead, they are positioned as auxiliary reasoning layers that augment detection outputs through semantic interpretation of alerts, behavioural abstraction over logs and traces, and contextual explanation of anomalous activity [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]. These systems must operate under high-throughput data streams, adversarial noise, concept drift, and strict latency budgets, rendering naive adoption of general-purpose LLM deployments impractical without careful architectural mediation.</p>
<p>This distinction has direct implications for both evaluation and deployment. Evaluation protocols commonly applied to general LLM applications-such as standalone reasoning benchmarks or static text-based assessments-are insufficient for IDS-specific scenarios, where errors may propagate downstream and amplify operational risk. In particular, false-positive escalation, inference latency, and reasoning instability can compound across SOC workflows, increasing analyst workload and degrading response effectiveness [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-121">121</xref>,<xref ref-type="bibr" rid="ref-140">140</xref>]. Consequently, performance claims derived from general LLM evaluations cannot be extrapolated to IDS deployments without explicit consideration of these system-level effects.</p>
<p>Explicitly distinguishing between general cybersecurity applications and IDS-specific architectures therefore prevents overstatement of LLM readiness for real-time intrusion detection and clarifies the necessity of hybrid, defence-in-depth designs. In such architectures, classical detectors provide time-critical guarantees, while LLMs contribute semantic reasoning, correlation, and explainability under human-in-the-loop supervision. <xref ref-type="table" rid="table-8">Table 8</xref> formalises this distinction and provides a conceptual reference framework for interpreting the task-level and quantitative analyses presented in subsequent sections.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Critical comparison between general LLM cybersecurity applications and IDS-specific LLM architectures.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dimension</th>
<th>General LLM Cybersecurity Applications</th>
<th>IDS-Specific LLM Architectures</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Primary Objective</bold></td>
<td>Analyst support, knowledge extraction, reasoning, and workflow automation</td>
<td>Detection augmentation, alert interpretation, and behavioural reasoning within IDS pipelines</td>
</tr>
<tr>
<td><bold>Typical Tasks</bold></td>
<td>Threat intelligence summarisation, IOC extraction, SOC co-pilots, secure coding, incident reporting</td>
<td>Post-detection reasoning for NIDS/HIDS/IIoT IDS, anomaly explanation, alert correlation</td>
</tr>
<tr>
<td><bold>Architectural Role of LLM</bold></td>
<td>Standalone or loosely coupled cognitive assistant</td>
<td>Embedded component layered on top of conventional IDS detectors</td>
</tr>
<tr>
<td><bold>Detection Responsibility</bold></td>
<td>No direct responsibility for intrusion detection decisions</td>
<td>Does not replace detectors; augments detection outputs with semantic reasoning</td>
</tr>
<tr>
<td><bold>Input Modality</bold></td>
<td>Natural language reports, alerts, threat feeds, code, policies</td>
<td>IDS alerts, logs, behavioural traces, network summaries (often structured &#x002B; text)</td>
</tr>
<tr>
<td><bold>Latency Constraints</bold></td>
<td>Moderate to low; interactive or batch processing acceptable</td>
<td>Strict; real-time detection delegated to classical models, LLMs operate asynchronously</td>
</tr>
<tr>
<td><bold>Evaluation Metrics</bold></td>
<td>Task accuracy, summarisation quality, extraction precision, analyst productivity</td>
<td>Detection accuracy/F1 (indirect), alert reduction, explanation quality, false-positive mitigation</td>
</tr>
<tr>
<td><bold>Datasets Used</bold></td>
<td>Curated text corpora, CTI reports, code repositories</td>
<td>IDS benchmarks (e.g., CICIDS, UNSW-NB15), logs, malware traces, IIoT telemetry</td>
</tr>
<tr>
<td><bold>Baseline Comparisons</bold></td>
<td>Often compared to rule-based tools or traditional NLP methods</td>
<td>Compared against CNN/RNN/Transformer IDS pipelines</td>
</tr>
<tr>
<td><bold>Failure Modes</bold></td>
<td>Hallucinated facts, grounding errors, outdated knowledge</td>
<td>False-positive escalation, reasoning errors over noisy alerts, latency overhead</td>
</tr>
<tr>
<td><bold>Operational Readiness</bold></td>
<td>High for SOC decision support and automation</td>
<td>Emerging; suitable for hybrid IDS, not for standalone real-time detection</td>
</tr>
<tr>
<td><bold>Governance &#x0026; Safety Risk</bold></td>
<td>Misinformation, policy misinterpretation, analyst over-trust</td>
<td>Cascading errors in IDS pipelines, unsafe automated response decisions</td>
</tr>
<tr>
<td><bold>Key Limitation</bold></td>
<td>Lack of direct linkage to detection pipelines</td>
<td>Lack of formal guarantees and limited real-world SOC-scale validation</td>
</tr>
<tr>
<td><bold>Research Gap</bold></td>
<td>Trust calibration and grounding in dynamic cyber contexts</td>
<td>Standardised IDS benchmarks with latency, drift, and deployment constraints</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Impact of Model Scale and Architecture</title>
<p>Although LLM-enabled threats and defences have been examined extensively, comparatively limited attention has been devoted to understanding how <italic>model scale, architectural design, and deployment modality</italic> influence robustness, misuse risk, and operational feasibility [<xref ref-type="bibr" rid="ref-153">153</xref>]. This subsection addresses this gap by synthesising empirical evidence across defensive studies and situating architectural choices within practical deployment constraints.</p>
<sec id="s6_2_1">
<label>6.2.1</label>
<title>Model Scale: Small vs. Large LLMs</title>
<p>Small- and medium-scale LLMs-including distilled, quantised, and domain-adapted variants-are increasingly favoured in SOC, edge, and IIoT settings due to their lower inference latency, reduced computational cost, and improved controllability. Multiple studies demonstrate that such models can achieve competitive performance on constrained security tasks, including log interpretation and malware trace reasoning [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>]. In contrast, large foundation models typically offer superior generalisation and deeper reasoning capabilities but are associated with elevated hallucination risk, greater potential for misuse amplification, and prohibitive operational costs in continuous-deployment environments [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Collectively, these findings suggest that model scale mediates a fundamental trade-off between defensive capability and deployment risk.</p>
</sec>
<sec id="s6_2_2">
<label>6.2.2</label>
<title>Deployment Modality: API-Based vs. Local Models</title>
<p>Deployment modality further conditions defensive effectiveness and risk exposure. API-based LLM deployments benefit from provider-managed updates, centralised alignment controls, and rapid access to state-of-the-art models, but introduce data sovereignty concerns, external service dependencies, and limited transparency into model behaviour. By contrast, locally deployed and fine-tuned models provide tighter control over data flows, inference logic, and privacy-sensitive logs, making them particularly suitable for SOC operations and critical infrastructure environments [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>]. However, local deployment shifts responsibility for model maintenance, security hardening, and misuse mitigation to system operators, thereby introducing additional governance and lifecycle-management challenges.</p>
</sec>
<sec id="s6_2_3">
<label>6.2.3</label>
<title>Single-Model vs. Agentic Architectures</title>
<p>Agentic and multi-LLM architectures are increasingly adopted to support complex investigation, correlation, and response workflows [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]. While such designs enhance modularity, task specialisation, and workflow automation, they also expand the effective attack surface and introduce new failure modes, including coordination errors, cascading hallucinations, and prompt-injection propagation across agents [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-154">154</xref>]. Single-model architectures, by contrast, are generally easier to audit, reason about, and deploy, but often struggle with multi-stage reasoning, cross-domain correlation, and concurrent task execution.</p>
</sec>
<sec id="s6_2_4">
<label>6.2.4</label>
<title>Architectural Synthesis</title>
<p>Taken together, architectural and scale-related design choices exert a first-order influence on defensive robustness and deployability. Larger and agentic systems prioritise capability breadth and scalability, whereas smaller, locally controlled models emphasise predictability, auditability, and operational safety. <xref ref-type="table" rid="table-9">Table 9</xref> summarises these trade-offs across representative deployment scenarios.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Comparative analysis of LLM architectures and model scale in cybersecurity applications.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dimension</th>
<th>Small/Distilled LLMs</th>
<th>Large Foundation LLMs</th>
<th>Agentic Multi-LLM Pipelines</th>
</tr>
</thead>
<tbody>
<tr>
<td>Model Capacity</td>
<td>Limited reasoning depth; task-specific</td>
<td>High generalisation and reasoning power</td>
<td>Distributed reasoning across specialised agents</td>
</tr>
<tr>
<td>Robustness</td>
<td>More predictable; easier to constrain</td>
<td>Stronger reasoning but higher hallucination risk</td>
<td>Sensitive to inter-agent coordination failures</td>
</tr>
<tr>
<td>Misuse Risk</td>
<td>Lower misuse amplification</td>
<td>Higher potential for abuse and weaponisation</td>
<td>Expanded attack surface across agents</td>
</tr>
<tr>
<td>Deployment Cost</td>
<td>Low compute and energy footprint</td>
<td>High inference and operational cost</td>
<td>High integration and orchestration overhead</td>
</tr>
<tr>
<td>Deployability</td>
<td>Suitable for edge, SOC, and ICS contexts</td>
<td>Primarily cloud-based deployments</td>
<td>Best suited for complex SOC automation</td>
</tr>
<tr>
<td>Governance &#x0026; Control</td>
<td>High controllability and auditability</td>
<td>Limited transparency; provider dependence</td>
<td>Complex policy enforcement across agents</td>
</tr>
<tr>
<td>Key Trade-Off</td>
<td>Safety and efficiency over generality</td>
<td>Capability over control</td>
<td>Scalability over simplicity</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Intrusion and Anomaly Detection</title>
<p>A substantial subset of the reviewed literature investigates the application of LLMs to intrusion and anomaly detection across heterogeneous telemetry sources, including system logs, network flows, endpoint activity, and malware execution traces. Survey studies such as [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>] report that language-model-based semantic representations can capture long-range dependencies and behavioural context more effectively than traditional feature-engineered approaches. Practical systems further illustrate complementary strengths: lightweight and domain-adapted models enable deployment in resource-constrained edge environments [<xref ref-type="bibr" rid="ref-105">105</xref>], while semantic trace modelling improves the detection of low-frequency and stealthy attack patterns [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>Domain adaptation emerges as a critical enabler of robust performance. Studies that integrate LLM-derived embeddings with classical detection mechanisms demonstrate improved resilience to dataset drift and enhanced generalisation across operational environments [<xref ref-type="bibr" rid="ref-103">103</xref>,<xref ref-type="bibr" rid="ref-106">106</xref>]. Nevertheless, persistent limitations remain. Inference latency constrains real-time deployment, prompt sensitivity undermines output stability, and overfitting to benchmark artefacts is frequently observed. As a result, the literature converges on positioning LLMs as <italic>post-detection reasoning layers</italic> that augment conventional intrusion detection systems rather than as standalone replacements.</p>
</sec>
<sec id="s6_4">
<label>6.4</label>
<title>Task-Level Comparison: NIDS, HIDS, and IIoT/ICS</title>
<p>Traditional IDS surveys typically categorise methods by deployment context-network-based intrusion detection systems (NIDS), host-based intrusion detection systems (HIDS), and IIoT/ICS environments-and evaluate transformer models primarily as feature-level classifiers optimised for benchmark accuracy [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]. LLM-assisted IDS approaches depart fundamentally from this paradigm. Rather than competing directly at the detection stage, LLMs are employed as higher-level cognitive components that support interpretation, correlation, and response.</p>
<p>In NIDS settings, LLMs operate downstream of packet- or flow-based detectors, enabling semantic correlation and explanation across heterogeneous alerts [<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>]. In HIDS deployments, LLMs abstract diverse host-level logs into unified semantic representations, facilitating cross-host reasoning and improved interpretability [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-136">136</xref>]. In IIoT and ICS contexts, LLMs are generally unsuitable for control-loop intrusion detection due to latency and safety constraints, but they provide significant value for post-hoc analysis, attack-path reasoning, and policy-aware decision support [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>]. These task-level distinctions are synthesised in <xref ref-type="table" rid="table-10">Table 10</xref>.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Task-level comparison of transformer-based IDS and LLM-assisted intrusion detection.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>IDS Task</th>
<th>Transformer-Based IDS Role</th>
<th>LLM-Assisted Role</th>
<th>Key Advantages of LLM Integration</th>
<th>Limitations and Applicability</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>NIDS</bold></td>
<td>Sequence modelling of packets, flows, and traffic features for supervised or semi-supervised classification [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>Post-detection semantic reasoning over alerts, logs, and correlated network events [<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
<td>Improved interpretability of alerts; contextual correlation across heterogeneous network telemetry; analyst decision support [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Not suitable for real-time packet inspection; dependent on upstream detectors and log quality; increased inference overhead [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
<tr>
<td><bold>HIDS</bold></td>
<td>Behavioural modelling of system calls, audit logs, and host activities using feature-level representations [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>Semantic abstraction of host logs and execution traces; cross-host reasoning and explanation of anomalous behaviour [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
<td>Enhanced detection of complex insider threats and privilege escalation; improved explainability across diverse log sources [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Sensitivity to log formatting and noise; requires robust preprocessing and human validation [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
<tr>
<td><bold>IIoT/ICS IDS</bold></td>
<td>Protocol-specific feature extraction and real-time anomaly detection under strict latency constraints [<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>Context-aware post-hoc analysis, attack-path reasoning, and policy-aware interpretation of incidents [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>Supports safety-aware reasoning and integration of cyber events with operational context<break/> in CPS environments [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Unsuitable for control-loop intrusion detection; limited validation in safety-critical real-time deployments [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
<tr>
<td><bold>Overall Positioning</bold></td>
<td>Feature-level classifiers optimised for detection accuracy on benchmark datasets [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>Cognitive and agentic security components augmenting IDS pipelines across detection, <break/>analysis, and response [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>Transforms IDS from isolated detection to explainable, decision-oriented defence architectures [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Requires hybrid architectures, governance controls, and human-in-the-loop oversight<break/> for safe deployment [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6_5">
<label>6.5</label>
<title>Quantitative Synthesis of LLM-Based Intrusion Detection Studies</title>
<p>Recent LLM-assisted intrusion detection systems (IDS) increasingly report performance improvements over classical machine learning and transformer-based baselines. However, the quantitative evidence remains fragmented due to heterogeneous datasets, inconsistent evaluation protocols, and divergent deployment assumptions. This subsection provides a critical synthesis of <italic>reported</italic> quantitative results from representative LLM-based IDS studies, focusing on detection performance, relative baseline improvements, latency and computational overhead, and dataset characteristics. No new experiments are introduced; rather, the analysis consolidates metrics as presented in the primary literature to enable principled cross-task comparison while explicitly contextualising reported gains [<xref ref-type="bibr" rid="ref-155">155</xref>].</p>
<sec id="s6_5_1">
<label>6.5.1</label>
<title>Detection Performance and Baseline Improvements</title>
<p>Across network-, host-, and IIoT-oriented IDS tasks, LLM-assisted approaches consistently report strong benchmark performance, with accuracy or F1-scores typically ranging between 92% and 99%. As summarised in <xref ref-type="table" rid="table-11">Table 11</xref>, the most pronounced gains arise when LLMs operate on semantically rich inputs such as system logs, behavioural traces, and alert narratives, where contextual reasoning complements feature-level detection [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>]. Relative to CNN- and RNN-based baselines, reported improvements commonly fall within the 2%&#x2013;7% range, whereas gains over transformer-based IDS are more modest and strongly context dependent. In particular, LLMs tend to outperform transformer-based IDS primarily in scenarios requiring cross-source correlation, explanation, or narrative reconstruction rather than raw packet- or flow-level classification [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-64">64</xref>].</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Quantitative summary of reported performance in LLM-based intrusion detection studies with baseline comparison.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>IDS Task</th>
<th>Reported Accuracy/F1 (Range)</th>
<th>Baseline Comparison (LLM vs CNN/RNN/<break/>Transformer)</th>
<th>Datasets Commonly Used</th>
<th>Latency/<break/>Overhead Reporting</th>
<th>Representative Studies</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>NIDS</bold></td>
<td>92%&#x2013;99% (Accuracy/F1, benchmark-dependent)</td>
<td>&#x002B;2%&#x2013;6% over CNN/RNN; &#x002B;1%&#x2013;4% over transformer NIDS on semantic-rich traffic and logs</td>
<td>CICIDS2017, UNSW-NB15, custom network telemetry</td>
<td><bold>Limited:</bold> LLMs typically applied post-detection; real-time latency rarely reported</td>
<td>[<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
</tr>
<tr>
<td><bold>HIDS</bold></td>
<td>93%&#x2013;98% (F1 dominant metric)</td>
<td>&#x002B;3%&#x2013;7% F1 over RNN/CNN baselines; marginal gains over transformer HIDS when context is limited</td>
<td>System-call traces, host audit logs, malware behaviour datasets</td>
<td><bold>Partial:</bold> Inference cost discussed qualitatively; no SOC-scale latency benchmarks</td>
<td>[<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
</tr>
<tr>
<td><bold>IIoT/ICS IDS</bold></td>
<td>90%&#x2013;97% (Accuracy under constrained settings)</td>
<td>&#x002B;2%&#x2013;5% over classical DL; comparable to transformers for raw detection, stronger for post-hoc reasoning</td>
<td>Industrial telemetry, CPS logs, IIoT-specific datasets</td>
<td><bold>Sparse:</bold> Latency largely unreported; unsuitable for real-time control loops</td>
<td>[<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
</tr>
<tr>
<td><bold>Cross-Task Observation</bold></td>
<td>High accuracy on curated benchmarks; strongest gains for semantic reasoning tasks</td>
<td>LLMs consistently outperform CNN/RNN baselines;<break/> gains over transformers are context-dependent</td>
<td>Benchmark-centric with limited domain diversity</td>
<td>Latency and computational cost remain underreported across studies</td>
<td>[<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Nevertheless, as evidenced by the study-level synthesis in <xref ref-type="sec" rid="s6_10">Section 6.10</xref>, these improvements are predominantly demonstrated on curated benchmark datasets such as CICIDS2017, UNSW-NB15, and custom malware corpora. This concentration limits cross-study comparability and risks overstating performance when models are deployed outside controlled laboratory settings or evaluated under realistic enterprise conditions.</p>
</sec>
<sec id="s6_5_2">
<label>6.5.2</label>
<title>Latency and Computational Overhead</title>
<p>In contrast to detection accuracy, latency and computational overhead are reported inconsistently across studies. Distilled or lightweight LLM variants demonstrate acceptable inference latency for post-detection reasoning and alert interpretation tasks in edge or resource-constrained environments [<xref ref-type="bibr" rid="ref-105">105</xref>]. By contrast, full-scale LLMs incur substantial computational overhead when applied to streaming intrusion detection, rendering them unsuitable for real-time packet inspection or control-loop defence [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. As reflected in the task-level synthesis, most studies implicitly adopt hybrid architectures in which classical IDS components handle time-critical detection, while LLMs operate asynchronously to support correlation, explanation, and analyst decision-making. Quantitative latency benchmarks under sustained SOC-scale workloads remain largely absent from the literature.</p>
</sec>
<sec id="s6_5_3">
<label>6.5.3</label>
<title>Dataset Concentration and Evaluation Bias</title>
<p>The quantitative synthesis further reveals a pronounced imbalance in dataset usage across IDS tasks. Network intrusion detection studies overwhelmingly rely on a small set of benchmark datasets, whereas host-based and IIoT/ICS datasets remain comparatively scarce, less standardised, and highly domain specific [<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]. Few studies explicitly evaluate robustness under dataset shift, concept drift, or long-term deployment conditions, despite the centrality of these factors in operational SOC environments. As highlighted in <xref ref-type="sec" rid="s7_5">Section 7.5</xref>, this benchmark-centric evaluation paradigm obscures generalisation limits and constrains meaningful quantitative comparison.</p>
</sec>
<sec id="s6_5_4">
<label>6.5.4</label>
<title>Quantitative Gaps and Implications</title>
<p>Taken together, the quantitative evidence suggests that LLM-assisted IDS can achieve strong benchmark performance and consistent improvements over classical baselines, particularly in semantically complex intrusion scenarios. However, these gains are accompanied by increased computational cost, limited real-time applicability, and insufficient validation under realistic operational constraints. The absence of standardised latency metrics, limited dataset diversity, and weak reproducibility practices remain significant barriers to robust quantitative assessment [<xref ref-type="bibr" rid="ref-114">114</xref>].</p>
<p>By consolidating reported accuracy ranges, baseline improvements, dataset usage, and latency reporting practices in <xref ref-type="table" rid="table-11">Table 11</xref>, this subsection complements the task-level analysis presented in <xref ref-type="sec" rid="s6_4">Section 6.4</xref>. Together, these findings underscore the need for next-generation IDS benchmarks that jointly report detection performance, computational cost, and deployment context, particularly for SOC-scale and IIoT intrusion detection.</p>

</sec>
</sec>
<sec id="s6_6">
<label>6.6</label>
<title>Threat Intelligence Extraction and SOC Automation</title>
<p>Threat intelligence workflows require the synthesis of information from highly unstructured and heterogeneous sources, including vulnerability advisories, forensic logs, advanced persistent threat (APT) reports, darknet content, and open threat feeds. A growing body of evidence demonstrates that LLMs are well suited to automating these processes. Loumachi et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] show that LLM-based pipelines substantially improve indicator-of-compromise (IOC) extraction, entity linking, and threat report summarisation. Broader reviews similarly confirm the effectiveness of generative AI for cyber threat intelligence (CTI) processing and intelligence generation [<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-132">132</xref>].</p>
<p>Retrieval-augmented generation (RAG) further enhances factual grounding and mitigates hallucination risk. Xu et al. [<xref ref-type="bibr" rid="ref-128">128</xref>] demonstrate that RAG-based cross-document reasoning produces more coherent threat timelines and reduces ambiguity in alert interpretation. Systems such as RAG-CDI operationalise this paradigm for actionable CTI generation [<xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-156">156</xref>]. Within SOC environments, LLM-based co-pilots increasingly support alert triage and response coordination, reducing mean time to response through automated synthesis and correlation of alerts [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>]. Nonetheless, persistent risks related to hallucination, inconsistent reasoning, and susceptibility to prompt manipulation necessitate structured prompting, robust grounding mechanisms, and multi-stage verification pipelines [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-157">157</xref>].</p>
</sec>
<sec id="s6_7">
<label>6.7</label>
<title>Secure Coding, Vulnerability Analysis, and Automated Repair</title>
<p>LLMs are widely applied to vulnerability detection, secure code synthesis, patch recommendation, and automated code review [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. Benchmarking studies report improved vulnerability localisation and recall across multiple Common Weakness Enumeration (CWE) families when compared with classical static and learning-based approaches [<xref ref-type="bibr" rid="ref-112">112</xref>,<xref ref-type="bibr" rid="ref-158">158</xref>,<xref ref-type="bibr" rid="ref-159">159</xref>]. Fine-tuned repair-oriented models further enhance patch correctness and semantic consistency, particularly for recurring vulnerability patterns [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-133">133</xref>].</p>
<p>However, empirical audits consistently reveal that LLM-generated code may introduce new vulnerabilities, omit edge cases, or yield incomplete remediation [<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-117">117</xref>]. As a result, the literature increasingly advocates hybrid pipelines that integrate LLM-based semantic reasoning with symbolic execution, static analysis, and developer-in-the-loop validation to mitigate hallucination-induced errors and ensure robustness [<xref ref-type="bibr" rid="ref-62">62</xref>,<xref ref-type="bibr" rid="ref-120">120</xref>,<xref ref-type="bibr" rid="ref-134">134</xref>].</p>
</sec>
<sec id="s6_8">
<label>6.8</label>
<title>Hybrid and Multi-Agent Defence Architectures</title>
<p>Hybrid and multi-agent defence architectures decompose complex security workflows into specialised agent roles coordinated by LLM-based planners [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Empirical studies report reductions in analyst workload and response latency through collaborative reasoning, task delegation, and cross-agent information sharing [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. Extensions of these architectures to cyber-physical systems and smart-grid environments further demonstrate their applicability beyond conventional IT infrastructures [<xref ref-type="bibr" rid="ref-47">47</xref>].</p>
<p>Despite these advantages, systemic risks persist, including hallucination propagation, coordination failures, and unintended tool misuse across chained agents [<xref ref-type="bibr" rid="ref-129">129</xref>,<xref ref-type="bibr" rid="ref-138">138</xref>]. Current empirical evidence therefore supports the deployment of multi-agent architectures primarily for decision support and semi-automated defence, rather than for fully autonomous response in high-assurance environments.</p>
</sec>
<sec id="s6_9">
<label>6.9</label>
<title>Operational Failure Modes and Cost-Aware Limitations</title>
<p>Across defensive domains, operational deployment exposes failure modes that are insufficiently captured by benchmark-centric evaluations. Semantic overgeneralisation can amplify false positives, exacerbating alert fatigue and misdirecting analyst attention [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Computational overhead, inference latency, and energy consumption further constrain scalability, particularly in SOC-scale and IIoT deployments, yet remain inconsistently reported in the literature [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>]. In addition, defensive LLM systems remain vulnerable to adversarial manipulation, including prompt injection and context poisoning, with cascading effects observed in RAG-based and multi-agent pipelines [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>]. These observations reinforce the conclusion that defensive effectiveness is inseparable from architectural safeguards, governance controls, and sustained human oversight.</p>
<p><bold>Capability-oriented synthesis.</bold> <xref ref-type="table" rid="table-12">Table 12</xref> provides a high-level, taxonomy-aligned synthesis of <italic>what LLM-based defensive systems are designed to do</italic>, focusing on their core technical roles and demonstrated capability gains across defensive subcategories.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Capability-oriented synthesis of LLM-based defensive subcategories across the proposed taxonomy.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>          
<tr>
<th>Defensive Subcategory</th>
<th>Key References</th>
<th>Core Technical Role of LLMs</th>
<th>Observed Benefits</th>
<th>Key Constraints</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Intrusion &#x0026; Anomaly Detection</bold></td>
<td>[<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
<td>Semantic interpretation of logs, flows, and traces; prompt-guided detection; hybrid LLM&#x2013;ML IDS pipelines</td>
<td>Improved detection of low-and-slow attacks; better interpretability than feature-only IDS; robustness gains via domain adaptation</td>
<td>High inference latency; limited real-time validation; sensitivity to prompt format; weak evidence of sustained SOC deployment</td>
</tr>
<tr>
<td><bold>Threat Intelligence &#x0026; SOC Automation</bold></td>
<td>[<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>,<xref ref-type="bibr" rid="ref-128">128</xref>]</td>
<td>RAG-based IOC extraction; alert summarisation; cross-source correlation; analyst co-pilot functions</td>
<td>Reduced analyst workload; faster triage and incident understanding; scalable processing of unstructured CTI</td>
<td>Hallucinated or mis-grounded IOCs; dependence on retrieval quality; limited validation in live SOC workflows</td>
</tr>
<tr>
<td><bold>Secure Coding &#x0026; Vulnerability Detection</bold></td>
<td>[<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>Semantic code reasoning; vulnerability localisation; guided patch generation; SDLC-integrated security assistance</td>
<td>Improved logic-flaw detection; better patch explanations; multi-language support under constrained prompting</td>
<td>Risk of incomplete or insecure fixes; uneven performance across CWE classes; requires developer-in-the-loop review</td>
</tr>
<tr>
<td><bold>Hybrid &#x0026; Multi-Agent Defence Systems</bold></td>
<td>[<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>,<xref ref-type="bibr" rid="ref-130">130</xref>]</td>
<td>Multi-agent coordination for investigation and response; policy-aware reasoning; tool-augmented workflows</td>
<td>End-to-end automation; cross-stage reasoning; reduced response time in complex incidents</td>
<td>Cascading hallucinations; coordination instability; expanded attack surface; lack of formal safety guarantees</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6_10">
<label>6.10</label>
<title>Defence Synthesis</title>
<p>Synthesising evidence across intrusion detection, threat intelligence processing, secure coding, and agentic defence architectures yields a consistent conclusion: LLMs deliver their greatest defensive value when deployed as <italic>bounded cognitive amplifiers</italic> rather than as autonomous decision-makers. Quantitative performance gains are strongest for tasks involving heterogeneous data integration, cross-alert reasoning, and explanation, whereas advantages over transformer-based IDS diminish for low-level, time-critical detection tasks.</p>
<p>Across all defensive subcategories (<xref ref-type="table" rid="table-12">Tables 12</xref> and <xref ref-type="table" rid="table-13">13</xref>), hybrid architectures emerge as the dominant and most operationally viable paradigm. Classical detectors provide latency guarantees and robustness under adversarial noise, while LLMs augment these pipelines with semantic reasoning and analyst-facing interpretability. At the same time, model scale, agentic complexity, and deployment modality introduce non-trivial trade-offs between capability, cost, and safety. Addressing these trade-offs requires evaluation frameworks that jointly consider accuracy, latency, robustness, and governance-an issue revisited in <xref ref-type="sec" rid="s7">Section 7</xref>.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Operational-risk and failure-mode analysis of representative LLM-based defensive studies.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Study</th>
<th>Defensive Subcategory</th>
<th>Model/System</th>
<th>Methodological Distinction</th>
<th>Demonstrated Performance/Advantage</th>
<th>Observed Limitations and Applicability</th>
</tr>
</thead>
<tbody>
<tr>
<td>Elouardi et al. [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Intrusion Detection</td>
<td>Survey of LLM-IDS</td>
<td>Semantic modelling of cyber telemetry for multi-stage attack detection</td>
<td>Identifies consistent gains over classical IDS for complex, multi-step attacks</td>
<td>Survey-level evidence only; limited operational validation</td>
</tr>
<tr>
<td>Rondanini et al. [<xref ref-type="bibr" rid="ref-160">160</xref>]</td>
<td>Edge Anomaly Detection</td>
<td>Lightweight LLMs</td>
<td>Model compression for edge deployment</td>
<td>Maintains strong accuracy under latency and resource constraints</td>
<td>Generalisation limited across heterogeneous network environments</td>
</tr>
<tr>
<td>Rondanini et al. [<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
<td>Malware Detection</td>
<td>Lightweight transformer IDS</td>
<td>Hybrid static-dynamic malware classification</td>
<td>Robust to obfuscation with low latency on mobile devices</td>
<td>Dataset imbalance impacts stability</td>
</tr>
<tr>
<td>Huynh et al. [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Behavioural Malware Detection</td>
<td>Transformer trace models</td>
<td>Semantic modelling of execution traces</td>
<td>Improved detection of stealthy malware families</td>
<td>Requires large labelled trace corpora</td>
</tr>
<tr>
<td>Che et al. [<xref ref-type="bibr" rid="ref-106">106</xref>]</td>
<td>Domain Adaptation</td>
<td>Adapted LLMs</td>
<td>Cross-family fine-tuning to mitigate dataset drift</td>
<td>Significant improvement in generalisation across malware families</td>
<td>High retraining cost; continuous adaptation required</td>
</tr>
<tr>
<td>Hmimou et al. [<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
<td>Hybrid IDS</td>
<td>LLM-ML fusion</td>
<td>Semantic log embeddings combined with classical ML</td>
<td>Strong robustness across heterogeneous log sources</td>
<td>Reduced interpretability; tuning complexity</td>
</tr>
<tr>
<td>Zou et al. [<xref ref-type="bibr" rid="ref-161">161</xref>]</td>
<td>Traffic Analysis</td>
<td>Transformer NIDS</td>
<td>Packet-level semantic embedding</td>
<td>Outperforms CNN/RNN baselines on long-range dependencies</td>
<td>Latency challenges at high throughput</td>
</tr>
<tr>
<td>Lai et al. [<xref ref-type="bibr" rid="ref-162">162</xref>]</td>
<td>IDS Benchmarking</td>
<td>LLM embeddings &#x002B; ML</td>
<td>Feature enrichment for downstream IDS</td>
<td>Consistent F1-score improvement across datasets</td>
<td>Performance tied to downstream classifier choice</td>
</tr>
<tr>
<td>Usman et al. [<xref ref-type="bibr" rid="ref-163">163</xref>]</td>
<td>Energy-Efficient IDS</td>
<td>Green LLM-based IDS</td>
<td>Power-aware detection strategies</td>
<td>Feasible IDS under strict energy budgets</td>
<td>Accuracy degrades in noisy traffic</td>
</tr>
<tr>
<td>Koide et al. [<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td>Phishing Detection</td>
<td>ChatGPT reasoning</td>
<td>Semantic analysis of phishing content</td>
<td>Effective detection beyond lexical features</td>
<td>English-only focus; limited multimodal scope</td>
</tr>
<tr>
<td>Li et al. [<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td>NIDS</td>
<td>Prompt-learning LLM</td>
<td>Semantic extraction from raw traffic</td>
<td>Competitive classification accuracy with minimal feature engineering</td>
<td>Requires environment-specific prompt adaptation</td>
</tr>
<tr>
<td>Loumachi et al. [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>CTI Extraction</td>
<td>LLM-NER pipelines</td>
<td>Context-aware IOC extraction</td>
<td>Improves extraction precision and analyst efficiency</td>
<td>Entity hallucination in sparse contexts</td>
</tr>
<tr>
<td>Xu et al. [<xref ref-type="bibr" rid="ref-128">128</xref>]</td>
<td>CTI Correlation</td>
<td>RAG-enhanced LLM</td>
<td>Cross-document threat correlation</td>
<td>Improved narrative reconstruction</td>
<td>Highly dependent on retrieval quality</td>
</tr>
<tr>
<td>Srinivas et al. [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>SOC Automation</td>
<td>LLM co-pilot</td>
<td>Automated triage and reporting</td>
<td>Reduced MTTR and analyst workload</td>
<td>Not validated under sustained SOC load</td>
</tr>
<tr>
<td>Karkuzhali and Senthilkumar [<xref ref-type="bibr" rid="ref-158">158</xref>]</td>
<td>Vulnerability Detection</td>
<td>LLM classifier</td>
<td>CWE-aligned semantic classification</td>
<td>High recall across vulnerability classes</td>
<td>False positives remain high</td>
</tr>
<tr>
<td>De Fitero-Dominguez et al. [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>Patch Generation</td>
<td>Fine-tuned repair LLM</td>
<td>Language-specific secure patching</td>
<td>Outperforms prior DL-based repair models</td>
<td>Patch reliability varies by vulnerability type</td>
</tr>
<tr>
<td>Pearce et al. [<xref ref-type="bibr" rid="ref-117">117</xref>]</td>
<td>Code Security Assessment</td>
<td>General-purpose LLMs</td>
<td>Empirical audit of AI-generated code</td>
<td>Reveals frequent silent vulnerabilities</td>
<td>Requires external verification frameworks</td>
</tr>
<tr>
<td>Andreoni-LN et al. [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>Multi-Agent IR</td>
<td>LLM agent collective</td>
<td>Collaborative reasoning across IR stages</td>
<td>Improved incident reconstruction accuracy</td>
<td>Error propagation across agents</td>
</tr>
<tr>
<td>Jaffal et al. [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>SOC Multi-Agent System</td>
<td>LLM agent framework</td>
<td>Task-specialised agent orchestration</td>
<td>Strong applicability to SOC workflows</td>
<td>Not robust under adversarial prompting</td>
</tr>
<tr>
<td>Zhai et al. [<xref ref-type="bibr" rid="ref-130">130</xref>]</td>
<td>Automated Response</td>
<td>Multi-agent responders</td>
<td>Response plan generation</td>
<td>Effective for simple attack scenarios</td>
<td>Fails under complex lateral movement</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Operational-risk and failure-mode synthesis.</bold> Complementing the capability-oriented overview, <xref ref-type="table" rid="table-13">Table 13</xref> provides a study-level synthesis that examines <italic>how LLM-based defences behave</italic> <italic>in practice</italic>, explicitly highlighting methodological distinctions, performance trade-offs, and limitations.</p>

<p>Taken together, <xref ref-type="table" rid="table-12">Tables 12</xref> and <xref ref-type="table" rid="table-13">13</xref> provide a consolidated response to <bold>RQ2</bold>, demonstrating that LLM-based defensive systems derive their primary value as components for semantic reasoning, contextual correlation, and analyst decision support rather than as standalone detection mechanisms. Across intrusion detection, threat intelligence processing, secure coding, and multi-agent response scenarios, the reviewed studies consistently report gains in interpretability, cross-source reasoning, and workflow automation. However, these benefits are intrinsically coupled to hybrid deployment models that integrate LLMs with conventional detectors, maintain robust human-in-the-loop oversight, and incorporate explicit safeguards against hallucination, false-positive escalation, and cascading reasoning failures.</p>

<p>Overall, this synthesis indicates that LLM-based defences should be understood as a complementary evolution of cybersecurity practice-one that reorients IDS and SOC workflows from isolated, signal-level detection toward explainable, context-aware, and decision-oriented defence systems. At the same time, the evidence underscores that sustainable operational adoption depends on evaluation paradigms and governance mechanisms that jointly account for detection effectiveness, latency and computational overhead, adversarial robustness, and deployment context. In the absence of such integrated assessment and control frameworks, the defensive advantages of LLMs risk being offset by new forms of operational fragility and systemic risk.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Security Evaluation Frameworks, Benchmarks, and Red-Teaming</title>
<p>As LLMs transition from experimental tools to embedded components of cybersecurity infra structures&#x2014;spanning SOC co-pilots, threat intelligence platforms, IDS reasoning layers, code analysis utilities, and multi-agent response systems&#x2014;the question of <italic>how</italic> these systems are evaluated becomes as critical as <italic>what</italic> they can do. The reviewed literature reflects a rapidly expanding ecosystem of security benchmarks, red-teaming strategies, and evaluation protocols targeting adversarial robustness, harmful-content refusal, alignment stability, and lifecycle risk. Nevertheless, evaluation practices remain uneven across application domains, modalities, and deployment contexts, with limited standardisation and minimal longitudinal assessment of model evolution.</p>
<p>This section critically organises evaluation research along five interrelated dimensions: (i) dataset usage and experimental realism, (ii) security-oriented benchmarks and datasets, (iii) red-teaming methodologies, (iv) critical benchmarking and ranking of evaluation metrics, and (v) an integrated evaluation synthesis. This structure explicitly differentiates laboratory-driven assessment from operationally meaningful security evaluation.</p>
<sec id="s7_1">
<label>7.1</label>
<title>Dataset Usage and Experimental Realism</title>
<p>Dataset selection and experimental design fundamentally shape the validity, generalisability, and interpretability of reported results in LLM-based cybersecurity research. While benchmark-driven evaluations support controlled comparison and reproducibility, they also introduce structural biases that limit experimental realism-particularly for intrusion detection and SOC-integrated deployments.</p>
<sec id="s7_1_1">
<label>7.1.1</label>
<title>Benchmark-Centric Evaluation Practices</title>
<p>A dominant trend across LLM-assisted intrusion detection and security analytics studies is the reliance on a small set of canonical benchmarks, including CICIDS2017, UNSW-NB15, curated malware corpora, and synthetic phishing datasets. As synthesised in <xref ref-type="sec" rid="s6_5_4">Sections 6.5.4</xref> and <xref ref-type="sec" rid="s6_10">6.10</xref>, these datasets are typically collected under controlled assumptions, with relatively clean labels, balanced class distributions, and isolated attack instances. Although indispensable for methodological comparison, such benchmarks fail to capture the noise, sparsity, temporal overlap, and severe alert imbalance characteristic of operational SOC telemetry [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
</sec>
<sec id="s7_1_2">
<label>7.1.2</label>
<title>Synthetic and Laboratory-Scale Evaluations</title>
<p>Beyond classical benchmarks, several studies evaluate LLM-based systems using synthetic logs, simulated attack traces, or LLM-generated adversarial inputs. These laboratory-scale evaluations facilitate scalable experimentation and targeted stress-testing of reasoning and alignment properties. However, they abstract away critical operational factors, including incomplete logging, policy-driven filtering, analyst feedback loops, and adaptive adversarial behaviour. Consequently, strong performance under synthetic conditions should be interpreted as an upper bound on potential capability rather than as evidence of deployment readiness [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>].</p>
</sec>
<sec id="s7_1_3">
<label>7.1.3</label>
<title>Evidence from Real-World and Longitudinal Evaluations</title>
<p>Only a limited subset of studies evaluates LLM-based systems using real SOC logs, enterprise network traces, or long-duration operational datasets. Even fewer investigate longitudinal effects such as concept drift, evolving attacker strategies, or feedback-induced behavioural shifts. This gap is particularly consequential for LLM-assisted IDS, where reasoning quality and alert correlation depend critically on contextual completeness and temporal continuity. As reflected in <xref ref-type="sec" rid="s7_5">Section 7.5</xref>, the scarcity of SOC-grade and longitudinal evaluations constrains the generalisability of reported gains and weakens claims of operational robustness [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>].</p>
</sec>
<sec id="s7_1_4">
<label>7.1.4</label>
<title>Implications for Experimental Realism</title>
<p>The prevailing reliance on benchmark-centric and laboratory-scale evaluations has direct implications for how reported results should be interpreted. Improvements in detection accuracy, reasoning coherence, or refusal robustness primarily reflect potential capability under idealised conditions rather than guarantees of real-world effectiveness. For IDS-specific and SOC-integrated LLM deployments, realistic evaluation must jointly consider detection performance, latency, alert volume, analyst interaction, and adversarial adaptation-dimensions that remain underrepresented in current evaluation pipelines.</p>
<p><xref ref-type="table" rid="table-14">Table 14</xref> summarises these contrasts and highlights the persistent gap between laboratory experimentation and operational deployment.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Critical comparison of dataset types and experimental realism in LLM-based cybersecurity studies.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dimension</th>
<th>Benchmark/Synthetic Evaluations</th>
<th>Real-World/Deployment-Oriented Evaluations</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Typical Data Sources</bold></td>
<td>CICIDS2017, UNSW-NB15, curated malware corpora, synthetic logs</td>
<td>SOC telemetry, enterprise network traces, operational host logs</td>
</tr>
<tr>
<td><bold>Data Characteristics</bold></td>
<td>Balanced classes, clean labels, isolated attacks</td>
<td>Severe class imbalance, noisy logs, overlapping benign activity</td>
</tr>
<tr>
<td><bold>Temporal Coverage</bold></td>
<td>Static or short-duration snapshots</td>
<td>Continuous, longitudinal data with evolving behaviour</td>
</tr>
<tr>
<td><bold>Adversarial Adaptation</bold></td>
<td>Largely absent or scripted</td>
<td>Implicit and adaptive attacker behaviour</td>
</tr>
<tr>
<td><bold>Evaluation Focus</bold></td>
<td>Detection accuracy, F1-score, reasoning correctness</td>
<td>Operational reliability, alert volume, analyst workload, robustness</td>
</tr>
<tr>
<td><bold>Latency Considerations</bold></td>
<td>Often ignored or qualitatively discussed</td>
<td>Critical for SOC and IDS pipelines</td>
</tr>
<tr>
<td><bold>Reproducibility</bold></td>
<td>High due to public availability</td>
<td>Low due to privacy and confidentiality constraints</td>
</tr>
<tr>
<td><bold>Primary Limitation</bold></td>
<td>Overestimation of performance and generalisation</td>
<td>Limited availability and standardisation</td>
</tr>
<tr>
<td><bold>Implication for LLM-Based IDS</bold></td>
<td>Upper-bound performance under idealised conditions</td>
<td>Realistic assessment of deployment readiness</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s7_2">
<label>7.2</label>
<title>Security-Oriented Benchmarks and Datasets</title>
<p>Security-oriented benchmarks constitute a foundational element of LLM evaluation by providing task-specific corpora designed to surface vulnerabilities under realistic cybersecurity conditions. Recent research has introduced a diverse set of benchmarks targeting phishing generation, harmful-query refusal, exploit reasoning, code-injection robustness, multilingual misinformation, and alignment stability.</p>
<p>Stylometry-aware phishing benchmarks [<xref ref-type="bibr" rid="ref-6">6</xref>] evaluate both AIGC-driven phishing capability and the resilience of downstream classifiers, while comprehensive security evaluation suites [<xref ref-type="bibr" rid="ref-25">25</xref>] span refusal behaviour, exploit-assisted reasoning, code-injection robustness, and semantic safety. High-variance synthetic phishing datasets [<xref ref-type="bibr" rid="ref-100">100</xref>] probe generalisation limits under adversarial diversity, and cross-lingual misinformation benchmarks [<xref ref-type="bibr" rid="ref-164">164</xref>] expose language-dependent robustness gaps. Additional domain-specific datasets target vulnerability reasoning [<xref ref-type="bibr" rid="ref-137">137</xref>,<xref ref-type="bibr" rid="ref-165">165</xref>], scenario-driven manipulation [<xref ref-type="bibr" rid="ref-138">138</xref>], constrained malware detection [<xref ref-type="bibr" rid="ref-105">105</xref>], and cross-model robustness comparison.</p>
<p>Despite their analytical value, most existing benchmarks remain static, predominantly text-centric, and non-adaptive. As a result, multimodal deception, adaptive adversarial strategies, social-context manipulation, and lifecycle-dependent effects are systematically underrepresented. Moreover, frequent vendor-driven model updates invalidate earlier benchmark results, complicating longitudinal comparison and undermining claims of sustained robustness across model generations.</p>
</sec>
<sec id="s7_3">
<label>7.3</label>
<title>Red-Teaming Methodologies</title>
<p>Red-teaming represents the primary mechanism for uncovering misalignment pathways, unsafe behaviours, and systemic vulnerabilities in LLM-based systems. Across the reviewed literature, red-teaming methodologies range from manual expert probing and automated jailbreak generation to conversational manipulation, scenario-based stress testing, and lifecycle-oriented evaluation.</p>
<p>Taxonomies of red-teaming strategies [<xref ref-type="bibr" rid="ref-15">15</xref>] classify attacks based on role-play, multi-turn interaction, and obfuscation techniques, while large-scale automated studies [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>,<xref ref-type="bibr" rid="ref-116">116</xref>] demonstrate the high transferability of jailbreak strategies across model families. Conversational red-teaming approaches [<xref ref-type="bibr" rid="ref-26">26</xref>] reveal failure modes that remain invisible under single-turn testing, and prompt-injection studies [<xref ref-type="bibr" rid="ref-166">166</xref>] expose persistent vulnerabilities in retrieval-augmented and tool-integrated systems. Lifecycle-oriented evaluations [<xref ref-type="bibr" rid="ref-27">27</xref>] further illustrate how vulnerabilities propagate across retrieval, agentic reasoning, and orchestration layers.</p>
<p>Despite this methodological diversity, red-teaming practices remain fragmented. Unified protocols, standardised success criteria, and reproducible baselines are largely absent, and systematic tracking of vulnerability persistence across model updates is rare. These limitations significantly constrain comparative analysis and impede longitudinal assessment of security improvements.</p>
</sec>
<sec id="s7_4">
<label>7.4</label>
<title>Critical Benchmarking and Ranking of Evaluation Metrics</title>
<p>Evaluation metrics play a decisive role in shaping conclusions about LLM security, yet they are frequently applied in isolation. Across the reviewed literature, three broad classes of metrics can be identified.</p>
<p><italic>Capability-centric metrics</italic> (e.g., accuracy, F1-score, exploit success rate) dominate IDS- and malware-oriented evaluations due to their quantitative clarity and ease of comparison [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]. While necessary, these metrics primarily capture task proficiency and often fail to reflect semantic misreasoning, cascading errors, or downstream operational impact.</p>
<p><italic>Behaviour-centric metrics</italic>, including refusal consistency, hallucination rate, calibration error, and explanation fidelity, offer deeper insight into reasoning robustness and alignment stability [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]. Such metrics are particularly relevant for SOC automation and agentic systems, yet remain underutilised in IDS-centric evaluations.</p>
<p><italic>Operational risk-centric metrics</italic>, such as false-positive escalation rate, analyst workload amplification, response latency, and inference overhead, are the most deployment-relevant but also the least standardised [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Their omission systematically inflates perceived readiness for real-world deployment and obscures trade-offs between accuracy, cost, and operational resilience.</p>
<p><xref ref-type="table" rid="table-15">Table 15</xref> ranks evaluation metrics and red-teaming methodologies according to diagnostic power, operational realism, scalability, and suitability for deployment in security-critical environments.</p>
<table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Critical ranking of security metrics and red-teaming methodologies for LLM-based cybersecurity systems.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Evaluation Method</th>
<th>Primary Focus</th>
<th>Diagnostic Power</th>
<th>Operational Realism</th>
<th>Scalability</th>
<th>Deployment Suitability</th>
<th>Overall Rank</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Detection Accuracy/F1</bold></td>
<td>Capability-centric performance</td>
<td>Medium</td>
<td>Low</td>
<td>High</td>
<td>Benchmark-driven IDS</td>
<td><bold>Medium</bold></td>
</tr>
<tr>
<td><bold>Exploit Success Rate</bold></td>
<td>Offensive capability validation</td>
<td>High</td>
<td>Medium</td>
<td>Medium</td>
<td>Malware/exploit analysis</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>Refusal Consistency Metrics</bold></td>
<td>Safety alignment stability</td>
<td>High</td>
<td>Low</td>
<td>High</td>
<td>Policy enforcement evaluation</td>
<td><bold>Medium</bold></td>
</tr>
<tr>
<td><bold>Hallucination &#x0026; Calibration Metrics</bold></td>
<td>Reasoning reliability</td>
<td>High</td>
<td>Medium</td>
<td>Medium</td>
<td>SOC automation, explanation systems</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>False-Positive Escalation Rate</bold></td>
<td>Operational risk amplification</td>
<td>Very High</td>
<td>High</td>
<td>Low</td>
<td>SOC pipelines, IDS triage</td>
<td><bold>Very High</bold></td>
</tr>
<tr>
<td><bold>Latency/Inference Overhead</bold></td>
<td>Runtime feasibility</td>
<td>High</td>
<td>High</td>
<td>Medium</td>
<td>Real-time and hybrid IDS</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>Manual/Human Red-Teaming</bold></td>
<td>Semantic failure discovery</td>
<td>Very High</td>
<td>High</td>
<td>Low</td>
<td>Safety-critical analysis</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>Automated Jailbreak Testing</bold></td>
<td>Alignment robustness</td>
<td>Medium</td>
<td>Low</td>
<td>Very High</td>
<td>Cross-model comparison</td>
<td><bold>Medium</bold></td>
</tr>
<tr>
<td><bold>Prompt Injection Stress Tests</bold></td>
<td>Pipeline integrity</td>
<td>High</td>
<td>Medium</td>
<td>High</td>
<td>RAG/agentic systems</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>Scenario-Based Red-Teaming</bold></td>
<td>System-level realism</td>
<td>Very High</td>
<td>Very High</td>
<td>Low</td>
<td>SOC, IDS, CPS deployments</td>
<td><bold>Very High</bold></td>
</tr>
<tr>
<td><bold>Lifecycle/Update Robustness Testing</bold></td>
<td>Security drift over time</td>
<td>High</td>
<td>High</td>
<td>Low</td>
<td>Production deployments</td>
<td><bold>High</bold></td>
</tr>
<tr>
<td><bold>Multimodal Red-Teaming</bold></td>
<td>Cross-modal safety</td>
<td>Medium</td>
<td>Medium</td>
<td>Medium</td>
<td>Vision-language systems</td>
<td><bold>Medium</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-15fn1" fn-type="other">
<p>Note: <bold>Diagnostic Power:</bold> Ability to expose non-trivial failure modes (hallucination, cascading errors). <bold>Operational Realism:</bold> Alignment with real SOC, IDS, or CPS deployment conditions. <bold>Deployment Suitability:</bold> Relevance for operational cybersecurity systems rather than laboratory benchmarks.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s7_5">
<label>7.5</label>
<title>Evaluation Synthesis</title>
<p>Synthesising evidence across datasets, benchmarks, red-teaming methodologies, and evaluation metrics reveals a persistent disconnect between laboratory-oriented assessment and real-world cybersecurity deployment. Benchmark-centric accuracy measures and automated red-teaming approaches offer scalability and reproducibility, yet they systematically overestimate robustness when models are exposed to operational complexity, adversarial adaptation, and deployment constraints. In contrast, scenario-driven red-teaming and operational risk-centric metrics provide substantially higher diagnostic value for security-critical use cases, but remain difficult to standardise, reproduce, and scale across diverse environments.</p>
<p>As consolidated in <xref ref-type="table" rid="table-16">Tables 16</xref> and <xref ref-type="table" rid="table-17">17</xref>, progress in LLM security evaluation depends on the adoption of multi-dimensional, lifecycle-aware assessment pipelines that jointly measure task capability, behavioural robustness, and operational risk. Without such integrated evaluation frameworks, claims regarding robustness, alignment, and deployment readiness will remain fragmented, difficult to compare across studies, and insufficiently grounded for high-assurance cybersecurity deployment.</p>
<table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Comparative overview of LLM security evaluation frameworks and red-teaming practices.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Evaluation Category</th>
<th>Key References</th>
<th>Primary Evaluation Focus</th>
<th>Key Gaps and Limitations</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Security Benchmarks &#x0026; Datasets</bold></td>
<td>[<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-137">137</xref>,<xref ref-type="bibr" rid="ref-138">138</xref>]</td>
<td>Phishing realism; jailbreak robustness; multilingual deception; refusal behaviour; malware reasoning</td>
<td>Predominantly static and text-centric; weak multimodal coverage; limited adaptive adversary modelling; poor alignment with SOC telemetry</td>
</tr>
<tr>
<td><bold>Red-Teaming &#x0026; Stress Testing</bold></td>
<td>[<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>]</td>
<td>Alignment bypass; multi-turn jailbreaks; prompt injection in RAG and agentic systems; black-box robustness</td>
<td>Lack of standardised protocols; non-deterministic outcomes; weak lifecycle and multi-agent evaluation; inconsistent severity reporting</td>
</tr>
<tr>
<td><bold>Metrics &#x0026; Evaluation Methodology</bold></td>
<td>[<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-164">164</xref>]</td>
<td>Refusal consistency; exploit success; hallucination and calibration metrics; composite safety indices</td>
<td>Highly context-dependent metrics; no robustness thresholds; vendor updates invalidate scores; scarce SOC-grade operational metrics</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-17">
<label>Table 17</label>
<caption>
<title>Comparative study-level analysis of evaluation benchmarks, red-teaming methods, and metrics.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Study</th>
<th>Evaluation Subcategory</th>
<th>Model/Setting</th>
<th>Methodological Distinction</th>
<th>Demonstrated Strengths/<break/>Performance Gains</th>
<th>Observed Limitations and Applicable Scenarios</th>
</tr>
</thead>
<tbody>
<tr>
<td>Opara et al. [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>Security Benchmark</td>
<td>Stylometry-aware phishing corpora</td>
<td>Adversarially diverse stylometric benchmarking</td>
<td>Reveals detector fragility under paraphrasing and stylistic variation</td>
<td>Primarily email-focused; limited multilingual and multimodal realism</td>
</tr>
<tr>
<td>Zhang et al. [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>Cross-Domain Benchmark</td>
<td>Comprehensive LLM security suite</td>
<td>Multi-task evaluation spanning jailbreaks, refusal, and exploit reasoning</td>
<td>One of the broadest comparative security benchmarks</td>
<td>Static tasks; lacks SOC-context and lifecycle realism</td>
</tr>
<tr>
<td>Nguyen et al. [<xref ref-type="bibr" rid="ref-137">137</xref>]</td>
<td>Security Dataset</td>
<td>Exploit and harm assessment corpora</td>
<td>Explicit modelling of harmful query execution chains</td>
<td>Improves reproducibility of exploit-oriented evaluations</td>
<td>Does not capture adaptive attacker evolution</td>
</tr>
<tr>
<td>Sezgin [<xref ref-type="bibr" rid="ref-138">138</xref>]</td>
<td>Scenario Benchmark</td>
<td>Narrative deception datasets</td>
<td>Scenario-driven, multi-step adversarial escalation</td>
<td>Captures progressive manipulation better than static prompts</td>
<td>Synthetic narratives limit ecological validity</td>
</tr>
<tr>
<td>Heiding et al. [<xref ref-type="bibr" rid="ref-100">100</xref>]</td>
<td>Adversarial Spam Benchmark</td>
<td>High-variance phishing datasets</td>
<td>Stress-testing of semantic and structural perturbations</td>
<td>Demonstrates sharp degradation of spam filters</td>
<td>Restricted to email-based attacks</td>
</tr>
<tr>
<td>Kulkarni et al. [<xref ref-type="bibr" rid="ref-164">164</xref>]</td>
<td>Misinformation Benchmark</td>
<td>Multilingual persuasion corpora</td>
<td>Cross-lingual deception robustness testing</td>
<td>Highlights cultural susceptibility differences</td>
<td>Text-only misinformation; no multimodal content</td>
</tr>
<tr>
<td>Toth et al. [<xref ref-type="bibr" rid="ref-165">165</xref>]</td>
<td>Secure Coding Benchmark</td>
<td>Cross-language vulnerability datasets</td>
<td>Language-wise security benchmarking</td>
<td>Exposes uneven security across programming languages</td>
<td>Limited coverage of memory-corruption flaws</td>
</tr>
<tr>
<td>Rondanini et al. [<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
<td>Malware Benchmark</td>
<td>Edge-focused malware corpora</td>
<td>Evaluation under compute-constrained settings</td>
<td>Strong accuracy on mobile/edge platforms</td>
<td>Lacks adversarially generated malware</td>
</tr>
<tr>
<td>Yao et al. [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>Red-Teaming Taxonomy</td>
<td>General-purpose LLMs</td>
<td>Structured categorisation of jailbreak techniques</td>
<td>Provides unified attack taxonomy</td>
<td>Limited quantitative evaluation</td>
</tr>
<tr>
<td>Villa et al. [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>Jailbreak Benchmarking</td>
<td>Open and proprietary LLMs</td>
<td>Large-scale black-box jailbreak evaluation</td>
<td>Shows non-linear size&#x2013;risk relationship</td>
<td>English-only; lacks multimodal testing</td>
</tr>
<tr>
<td>Hilario et al. [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>Conversational Red-Teaming</td>
<td>Human-guided multi-turn attacks</td>
<td>Context layering and narrative manipulation</td>
<td>Uncovers failures unseen in single-turn prompts</td>
<td>Labour-intensive and difficult to scale</td>
</tr>
<tr>
<td>Greshake et al. [<xref ref-type="bibr" rid="ref-166">166</xref>]</td>
<td>Prompt Injection Analysis</td>
<td>Instruction-following LLMs</td>
<td>Indirect injection through templates</td>
<td>Effective even against hardened prompts</td>
<td>Limited evaluation of agentic/RAG systems</td>
</tr>
<tr>
<td>De Maio et al. [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>Lifecycle Stress-Testing</td>
<td>RAG and agent pipelines</td>
<td>Cross-component vulnerability propagation analysis</td>
<td>Exposes system-level fragility</td>
<td>Limited quantitative severity metrics</td>
</tr>
<tr>
<td>Xue et al. [<xref ref-type="bibr" rid="ref-111">111</xref>]</td>
<td>Jailbreak Attacks</td>
<td>Transformer LLMs</td>
<td>Dual-path syntactic-semantic attacks</td>
<td>Significantly higher success than single-vector attacks</td>
<td>Limited testing on small models</td>
</tr>
<tr>
<td>Huang et al. [<xref ref-type="bibr" rid="ref-139">139</xref>]</td>
<td>Capability Stress-Testing</td>
<td>Foundation LLMs</td>
<td>Boundary probing for unsafe behaviours</td>
<td>Identifies emergent alignment failures</td>
<td>No segmentation by alignment method</td>
</tr>
<tr>
<td>Dharmendra et al. [<xref ref-type="bibr" rid="ref-167">167</xref>]</td>
<td>Human-Guided Red-Teaming</td>
<td>Multiple LLMs</td>
<td>Interactive adversarial probing</td>
<td>Finds deeper vulnerabilities than automation</td>
<td>High cost; poor scalability</td>
</tr>
<tr>
<td>Zaydi et al. [<xref ref-type="bibr" rid="ref-168">168</xref>]</td>
<td>Adversarial Stress-Testing</td>
<td>Mixed-model environments</td>
<td>Agent-enabled attack exploration</td>
<td>Reveals weaknesses in tool-augmented systems</td>
<td>Limited RAG-pipeline coverage</td>
</tr>
<tr>
<td>Zhang et al. [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>Safety Metrics</td>
<td>General LLMs</td>
<td>Severity-weighted harm scoring</td>
<td>Nuanced risk stratification</td>
<td>No normative deployment thresholds</td>
</tr>
<tr>
<td>Jaffal et al. [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>Composite Metrics</td>
<td>Multi-task evaluation</td>
<td>Integrated failure-mode indices</td>
<td>Identifies cross-domain vulnerabilities</td>
<td>Dataset heterogeneity affects comparability</td>
</tr>
<tr>
<td>Nguyen et al. [<xref ref-type="bibr" rid="ref-137">137</xref>]</td>
<td>Exploit Metrics</td>
<td>Exploit - evaluation models</td>
<td>Exploit success probability metrics</td>
<td>Operationally meaningful exploit evaluation</td>
<td>Limited multi-model scope</td>
</tr>
<tr>
<td>Chen et al. [<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>Calibration Metrics</td>
<td>General LLMs</td>
<td>Hallucination and consistency scoring</td>
<td>Improves explanation faithfulness</td>
<td>Limited cross-lingual validation</td>
</tr>
<tr>
<td>Kulkarni et al. [<xref ref-type="bibr" rid="ref-164">164</xref>]</td>
<td>Cross-Lingual Metrics</td>
<td>Multilingual LLMs</td>
<td>Cultural deception susceptibility scoring</td>
<td>Exposes strong language-wise variance</td>
<td>Text-only scope</td>
</tr>
<tr>
<td>Omar et al. [<xref ref-type="bibr" rid="ref-169">169</xref>]</td>
<td>Synthetic Behaviour Metrics</td>
<td>Narrative models</td>
<td>Persuasion intensity and deception fidelity</td>
<td>LLMs approach human-level persuasion</td>
<td>Lacks real-world influence validation</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s8">
<label>8</label>
<title>Methodological Orientation in LLM-Cybersecurity Research</title>
<p>Beyond the thematic axes of AIGC-driven threats, defensive capabilities, and security evaluation, the reviewed literature exhibits substantial methodological diversity. This cross-cutting methodological dimension is critical for assessing evidentiary strength, reproducibility, and operational relevance. Across the 167 peer-reviewed studies included in this survey, methodological orientations converge around four dominant classes: (i) empirical studies that measure LLM behaviour under controlled or adversarial conditions, (ii) system and architectural proposals that translate LLM capabilities into operational cybersecurity pipelines, (iii) conceptual and analytical contributions that articulate theoretical, socio-technical, and lifecycle risk frameworks, and (iv) survey and meta-analytic works that consolidate and structure the rapidly expanding research landscape. Each methodological class contributes distinct forms of insight while exhibiting characteristic limitations that influence the maturity and reliability of LLM-driven cybersecurity research.</p>
<sec id="s8_1">
<label>8.1</label>
<title>Empirical Studies</title>
<p>Empirical research constitutes the largest methodological category within the reviewed corpus, reflecting the field&#x2019;s strong emphasis on behavioural evaluation, attack modelling, and detection assessment. These studies span a wide range of threat and defence scenarios, including phishing and impersonation generation [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>], exploit reasoning and code-generation misuse [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>], malware behavioural modelling [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>], and detection enhancement through prompt-learning approaches for network intrusion detection systems [<xref ref-type="bibr" rid="ref-64">64</xref>] or phishing detection [<xref ref-type="bibr" rid="ref-63">63</xref>]. Empirical red-teaming further encompasses scenario-driven deception evaluation [<xref ref-type="bibr" rid="ref-138">138</xref>], dual-path jailbreak attacks [<xref ref-type="bibr" rid="ref-111">111</xref>], multimodal and conversational manipulation [<xref ref-type="bibr" rid="ref-26">26</xref>], and cross-architecture jailbreak robustness testing [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-116">116</xref>].</p>
<p>Despite addressing diverse threat landscapes, empirical studies consistently reveal methodological constraints. Common limitations include reliance on small or synthetic datasets, limited cross-lingual and multimodal evaluation, insufficient modelling of adaptive adversaries, lack of replication across model updates, and weak ecological validity. Many studies employ constrained prompt templates or narrow experimental testbeds that fail to capture realistic operational attacker behaviour. Consequently, empirical research provides high-resolution snapshots of LLM capabilities and vulnerabilities, but often lacks longitudinal depth and generalisability.</p>
</sec>
<sec id="s8_2">
<label>8.2</label>
<title>System and Framework Proposals</title>
<p>System and architectural proposals focus on operationalising LLM capabilities within cybersecurity workflows, including threat intelligence extraction pipelines, SOC triage assistants, autonomous defence agents, and hybrid LLM-ML architectures. Representative contributions include named-entity-recognition-enhanced IOC extraction [<xref ref-type="bibr" rid="ref-20">20</xref>], multi-agent investigative systems for incident response [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>], graph-augmented reasoning for threat-path reconstruction [<xref ref-type="bibr" rid="ref-131">131</xref>], autonomous containment agents [<xref ref-type="bibr" rid="ref-130">130</xref>], and domain-adapted intrusion detection architectures [<xref ref-type="bibr" rid="ref-103">103</xref>,<xref ref-type="bibr" rid="ref-106">106</xref>]. Notably, CTI platforms such as the <bold>RAG-CDI framework</bold> [<xref ref-type="bibr" rid="ref-55">55</xref>] and the real-time detection system <bold>WatchOverGPT</bold> [<xref ref-type="bibr" rid="ref-56">56</xref>] demonstrate the practical integration of LLMs for intelligence fusion. Additional system-level studies propose <bold>customisable LLM frameworks</bold> for vulnerability detection and repair, alert prioritisation mechanisms such as <bold>TailorAlert</bold>, and interpretable frameworks for cybersecurity risk assessment [<xref ref-type="bibr" rid="ref-60">60</xref>].</p>
<p>While these proposals demonstrate substantial potential for applied LLM cybersecurity, they also expose systemic fragilities. Recurrent issues include dependence on brittle prompt patterns, susceptibility to prompt injection in multi-component pipelines, hallucination propagation across chained reasoning steps, and over-reliance on curated retrieval corpora. Only a small fraction of systems are evaluated in live SOC environments or under sustained adversarial pressure; most rely on offline logs, synthetic scenarios, or limited testbeds. As a result, although systems research outlines plausible pathways toward deployment, the maturity and reliability of existing prototypes remain insufficient for high-assurance operational use.</p>
</sec>
<sec id="s8_3">
<label>8.3</label>
<title>Conceptual and Analytical Studies</title>
<p>Conceptual and analytical studies contribute theoretical depth by examining socio-technical implications, risk taxonomies, lifecycle vulnerabilities, and governance structures associated with LLM deployment. Prominent analyses investigate AI-enabled persuasion and its behavioural foundations [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-123">123</xref>], lifecycle propagation of prompt injection in retrieval-augmented and agentic ecosystems [<xref ref-type="bibr" rid="ref-27">27</xref>], and multi-dimensional risk structures governing safety and misuse [<xref ref-type="bibr" rid="ref-3">3</xref>]. Other contributions address <bold>protecting code with LLMs</bold> [<xref ref-type="bibr" rid="ref-134">134</xref>], the use of <bold>generative AI for automated security testing</bold> [<xref ref-type="bibr" rid="ref-62">62</xref>], the fragility of stylometric and watermark-based detection [<xref ref-type="bibr" rid="ref-149">149</xref>], and ethical considerations surrounding automated reasoning and decision support.</p>
<p>These studies excel in conceptual clarity and systemic perspective, yet frequently lack empirical grounding or validation under realistic adversarial conditions. Risk intensities are rarely quantified, formal evaluation protocols are often absent, and integration with behavioural experiments is limited. Nonetheless, conceptual and analytical works provide indispensable frameworks for interpreting systemic vulnerabilities and long-term risks that empirical studies alone cannot reveal.</p>
</sec>
<sec id="s8_4">
<label>8.4</label>
<title>Survey and Meta-Analytic Contributions</title>
<p>Survey and meta-analytic studies synthesise longitudinal developments, identify methodological gaps, and map the evolving landscape of LLM-centric cybersecurity research. Representative examples include comprehensive surveys of AIGC-driven cyber threats [<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>], analyses of LLM safety and robustness [<xref ref-type="bibr" rid="ref-16">16</xref>], systematic reviews of adversarial prompting strategies [<xref ref-type="bibr" rid="ref-15">15</xref>], and surveys focusing on LLM-enabled intrusion detection and CTI processing [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-132">132</xref>].</p>
<p>However, methodological transparency varies considerably across these works. Many surveys rely heavily on preprints, lack PRISMA alignment, or aggregate heterogeneous results without normalising evaluation conditions. Only a limited subset employs explicit inclusion criteria, reproducible search strategies, or structured comparative synthesis. The present survey addresses these shortcomings by adhering to PRISMA guidelines, restricting analysis to peer-reviewed studies, and integrating cross-dimensional evidence spanning threats, defences, and evaluation frameworks.</p>
</sec>
<sec id="s8_5">
<label>8.5</label>
<title>Methodological Synthesis and Implications</title>
<p>The interaction among these methodological orientations reveals a research field advancing rapidly yet unevenly. Empirical studies offer essential behavioural evidence but require more realistic datasets, multimodal evaluation, and longitudinal replication. System proposals demonstrate operational promise but remain constrained by hallucination risks, dependency fragilities, and unresolved security vulnerabilities. Conceptual analyses articulate systemic risks but require empirical validation to inform actionable guidance. Survey studies provide structural clarity yet must adopt more rigorous and standardised methodologies to reduce interpretive variance.</p>
<p>Integrating these orientations within the unified taxonomy proposed in this work enables a coherent assessment of research maturity. It highlights the need for stronger methodological coupling, including empirical validation of conceptual frameworks, system-level evaluation using adversarial benchmarks, and surveys incorporating meta-analytic quantification. Progress toward high-assurance deployment requires reproducible, operationally grounded, and longitudinal methodologies capable of capturing both dynamic adversarial behaviour and the rapid evolution of LLM architectures.</p>
</sec>
<sec id="s8_6">
<label>8.6</label>
<title>Interpretation of Methodological Mapping</title>
<p>The distribution of studies in <xref ref-type="table" rid="table-18">Table 18</xref> reveals pronounced methodological asymmetries across LLM-cybersecurity domains. Research on AIGC-driven threats is predominantly empirical, reflecting the feasibility of controlled phishing, malware, jailbreak, and misinformation experiments, albeit with limited ecological and longitudinal validation. Defensive research exhibits a more balanced mix of empirical and system-oriented approaches, yet many proposed architectures remain at the prototype stage and lack rigorous evaluation in operational environments. Evaluation and red-teaming research is the most diverse but also the most fragmented, underscoring the absence of unified standards, reproducible protocols, and cross-model longitudinal tracking.</p>
<table-wrap id="table-18">
<label>Table 18</label>
<caption>
<title>Methodological mapping of studies across threat, defence, and evaluation dimensions.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Study</th>
<th>A</th>
<th>B</th>
<th>C</th>
<th>D</th>
<th>E</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="6"><bold>A. AIGC-Driven Threats</bold></td>
</tr>
<tr>
<td>Opara et al. [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>Threats</td>
<td>Phishing</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Schmitt and Flechais [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>Threats</td>
<td>Influence Ops</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Aseeri and Bohacek [<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
<td>Threats</td>
<td>Spear-Phishing</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Pham et al. [<xref ref-type="bibr" rid="ref-98">98</xref>]</td>
<td>Threats</td>
<td>Targeted Phishing</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Koide et al. [<xref ref-type="bibr" rid="ref-99">99</xref>]</td>
<td>Threats</td>
<td>Spam Bypass</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Li et al. [<xref ref-type="bibr" rid="ref-170">170</xref>]</td>
<td>Threats</td>
<td>AI-Phishing Survey</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Shrestha et al. [<xref ref-type="bibr" rid="ref-102">102</xref>]</td>
<td>Threats</td>
<td>Adv. Messaging</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Weinz et al. [<xref ref-type="bibr" rid="ref-101">101</xref>]</td>
<td>Threats</td>
<td>Human Factors</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Iturbe et al. [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>Threats</td>
<td>Polymorphic Malware</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Yamin et al. [<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td>Threats</td>
<td>Malware Generation</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Rondanini et al. [<xref ref-type="bibr" rid="ref-160">160</xref>]</td>
<td>Threats</td>
<td>Edge Malware</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Devadiga et al. [<xref ref-type="bibr" rid="ref-107">107</xref>]</td>
<td>Threats</td>
<td>Exploit Derivation</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Maity and Arora [<xref ref-type="bibr" rid="ref-116">116</xref>]</td>
<td>Threats</td>
<td>Jailbreaking</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Villa et al. [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>Threats</td>
<td>Jailbreak Survey</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Pearce et al. [<xref ref-type="bibr" rid="ref-117">117</xref>]</td>
<td>Threats</td>
<td>Prompt Injection</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Greshake et al. [<xref ref-type="bibr" rid="ref-166">166</xref>]</td>
<td>Threats</td>
<td>Indirect Injection</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>De Maio et al. [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>Threats</td>
<td>Lifecycle Risks</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Doumanas et al. [<xref ref-type="bibr" rid="ref-149">149</xref>]</td>
<td>Threats</td>
<td>Watermark Weakness</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Harris et al. [<xref ref-type="bibr" rid="ref-123">123</xref>]</td>
<td>Threats</td>
<td>Fake News Gen.</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Kuntur et al. [<xref ref-type="bibr" rid="ref-122">122</xref>]</td>
<td>Threats</td>
<td>Disinfo</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Valdez et al. [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>Threats</td>
<td>Cross-Cultural Misinfo</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td align="center" colspan="6"><bold>B. LLM-Based Defensive Capabilities</bold></td>
</tr>
<tr>
<td>Elouardi et al. [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Defence</td>
<td>IDS Survey</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Koide et al. [<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td>Defence</td>
<td>Phishing Detection</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Li and Gong [<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td>Defence</td>
<td>NIDS Prompting</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Rondanini et al. [<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
<td>Defence</td>
<td>Edge Malware Det.</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Huynh et al. [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Defence</td>
<td>Behavioural Malware</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Che et al. [<xref ref-type="bibr" rid="ref-106">106</xref>]</td>
<td>Defence</td>
<td>Domain-Adapted IDS</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Hmimou et al. [<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
<td>Defence</td>
<td>Fusion IDS</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Zou et al. [<xref ref-type="bibr" rid="ref-161">161</xref>]</td>
<td>Defence</td>
<td>NIDS Pipelines</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Sarker [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>Defence</td>
<td>CTI Review</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Liu [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>Defence</td>
<td>CTI Review</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Sarker [<xref ref-type="bibr" rid="ref-132">132</xref>]</td>
<td>Defence</td>
<td>CTI Review</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Loumachi et al. [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>Defence</td>
<td>CTI Extraction</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Gandhi [<xref ref-type="bibr" rid="ref-55">55</xref>]</td>
<td>Defence</td>
<td>RAG-CTI (RAG-CDI)</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Shahid et al. [<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>Defence</td>
<td>Real-time CTI</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Xu et al. [<xref ref-type="bibr" rid="ref-128">128</xref>]</td>
<td>Defence</td>
<td>RAG-CTI Summ.</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Srinivas et al. [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Defence</td>
<td>SOC Assistant</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Alharthi and Yasaei [<xref ref-type="bibr" rid="ref-171">171</xref>]</td>
<td>Defence</td>
<td>CTI Automation</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>de Fitero-Dominguez et al. [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>Defence</td>
<td>Patch Gen.</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Babaey and Ravindran [<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>Defence</td>
<td>SQLi Defense</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Che et al. [<xref ref-type="bibr" rid="ref-133">133</xref>]</td>
<td>Defence</td>
<td>Code Analysis</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Moorthy et al. [<xref ref-type="bibr" rid="ref-120">120</xref>]</td>
<td>Defence</td>
<td>Vuln. Discovery</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Chu et al. [<xref ref-type="bibr" rid="ref-134">134</xref>]</td>
<td>Defence</td>
<td>Code Protection</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Corchado et al. [<xref ref-type="bibr" rid="ref-62">62</xref>]</td>
<td>Defence</td>
<td>Automated Sec Test</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Andreoni-LN et al. [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>Defence</td>
<td>Multi-Agent IR</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Krishnamurthy et al. [<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
<td>Defence</td>
<td>Agent Teams for IR</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Seo et al. [<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
<td>Defence</td>
<td>Risk Assessment</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Jaffal et al. [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>Defence</td>
<td>SOC Agents</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Ibrahim and Kashef [<xref ref-type="bibr" rid="ref-131">131</xref>]</td>
<td>Defence</td>
<td>Threat Graph Reasoning</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Zhai et al. [<xref ref-type="bibr" rid="ref-130">130</xref>]</td>
<td>Defence</td>
<td>Response Agents</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Yang et al. [<xref ref-type="bibr" rid="ref-129">129</xref>]</td>
<td>Defence</td>
<td>Policy Verification</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>Defence</td>
<td>General Review</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Lopez et al. [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Defence</td>
<td>General Review</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Uddin et al. [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>Defence</td>
<td>General Review</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td align="center" colspan="6"><bold>C. Evaluation Frameworks, Benchmarks, and Red-Teaming</bold></td>
</tr>
<tr>
<td>Zhang et al. [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>Eval.</td>
<td>Sec. Benchmark</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Kulkarni et al. [<xref ref-type="bibr" rid="ref-164">164</xref>]</td>
<td>Eval.</td>
<td>Misinfo Benchmark</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Nguyen et al. [<xref ref-type="bibr" rid="ref-137">137</xref>]</td>
<td>Eval.</td>
<td>Harm/Vuln. Data</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Sezgin [<xref ref-type="bibr" rid="ref-138">138</xref>]</td>
<td>Eval.</td>
<td>Deception Bench.</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Yao et al. [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>Eval.</td>
<td>Red-Teaming Survey</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Hilario et al. [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>Eval.</td>
<td>Conv. Red-Teaming</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Xue et al. [<xref ref-type="bibr" rid="ref-111">111</xref>]</td>
<td>Eval.</td>
<td>Dual-Path Jailbreak</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Dharmendra et al. [<xref ref-type="bibr" rid="ref-167">167</xref>]</td>
<td>Eval.</td>
<td>Human-Guided RT</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Zaydi and Maleh [<xref ref-type="bibr" rid="ref-168">168</xref>]</td>
<td>Eval.</td>
<td>Adv. Stress-Testing</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Villa et al. [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>Eval.</td>
<td>Jailbreak Testing</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>De Maio et al. [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>Eval.</td>
<td>Lifecycle Risk</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Jaffal et al. [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>Eval.</td>
<td>Risk Taxonomy</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Zhang et al. [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>Eval.</td>
<td>Alignment Metrics</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Chen et al. [<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>Eval.</td>
<td>Calibration Metrics</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Omar et al. [<xref ref-type="bibr" rid="ref-169">169</xref>]</td>
<td>Eval.</td>
<td>Synthetic Behaviour</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Alawida et al. [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>Eval.</td>
<td>Governance Review</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Karras et al. [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>Eval.</td>
<td>Governance/Risk</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-18fn1" fn-type="other">
<p>Note: <bold>A</bold> &#x003D; Research Dimension; <bold>B</bold> &#x003D; Taxonomy Subcategory; <bold>C</bold> &#x003D; Empirical Study (e.g., experiments, data collection, measurement); <bold>D</bold> &#x003D; System or Framework Contribution (e.g., prototype, tool, concrete pipeline); <bold>E</bold> &#x003D; Conceptual, Analytical, or Survey Contribution (e.g., risk model, systematic review, governance analysis).</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>These patterns expose both fragmentation and clear priorities for future work: the development of standardised evaluation pipelines, multimodal and multilingual datasets, field-tested system deployments, and integrated methodologies that bridge empirical, system-level, and conceptual insights. Only through such convergence can the field advance toward reliable, security-critical deployment of LLM-driven cybersecurity capabilities.</p>
</sec>
</sec>
<sec id="s9">
<label>9</label>
<title>Cross-Cutting Challenges and Open Research Gaps</title>
<p>The integrated analysis of AIGC-driven threats, LLM-enabled defensive mechanisms, and security evaluation frameworks reveals a set of deeply interdependent challenges that constrain the secure, reliable, and accountable adoption of LLMs within cybersecurity ecosystems. These challenges do not arise from isolated technical deficiencies; rather, they emerge from the interaction of probabilistic language models with adversarial environments, complex system integrations, and incomplete evaluation and governance practices. Despite rapid advances in individual components, the field remains characterised by methodological fragmentation, uneven empirical grounding, limited operational realism, and a pace of adversarial innovation that consistently outstrips defensive maturity.</p>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> conceptualises these weaknesses as a tightly coupled system of risks rather than independent failure modes. Hallucination and reliability failures interact with adversarial fragility, data governance and privacy risks, evaluation gaps, and regulatory constraints, collectively amplifying systemic vulnerability. Consequently, improvements in a single dimension&#x2014;such as detection accuracy or automation efficiency&#x2014;do not necessarily translate into safer or more trustworthy deployments when other dimensions remain underdeveloped.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Conceptual overview of cross-cutting challenges in LLM-enabled cybersecurity, highlighting interrelated issues spanning hallucination and reliability, adversarial fragility, data governance and privacy, evaluation gaps, and governance constraints.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77367-fig-3.tif"/>
</fig>
<sec id="s9_1">
<label>9.1</label>
<title>Scenario-Grounded Limitations and Open Challenges</title>
<p>When examined under realistic deployment conditions, LLM-based cybersecurity systems exhibit persistent limitations that extend beyond model architecture or training strategy. Empirical evidence indicates that many shortcomings arise from the coupling of LLM reasoning with operational constraints, incomplete context, and adaptive adversarial behaviour [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Grounding these limitations in concrete scenarios&#x2014;SOC automation, intrusion detection, and agentic response systems&#x2014;clarifies why benchmark-level gains frequently fail to translate into dependable operational performance.</p>
<p>In SOC automation and threat intelligence pipelines, LLMs are increasingly employed for alert triage, IOC extraction, report generation, and cross-source correlation [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-132">132</xref>]. While these applications reduce analyst workload and accelerate situational awareness, they introduce operational risks that are largely invisible to accuracy-centric evaluations. A central concern is false-positive escalation: hallucinated or weakly supported indicators may be amplified through automated prioritisation workflows, increasing alert fatigue and diverting analyst attention from genuinely critical incidents [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. As reflected in the comparative analyses in <xref ref-type="table" rid="table-13">Table 13</xref>, retrieval-augmented generation pipelines are particularly sensitive to corpus quality, temporal relevance, and retrieval completeness [<xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>]. In large-scale SOC environments characterised by heterogeneous, noisy, and evolving telemetry, even minor grounding errors can propagate across downstream automation, undermining trust and decision quality [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>].</p>

<p>Comparable limitations arise in LLM-assisted intrusion and anomaly detection. Semantic reasoning over logs, traces, and traffic summaries improves detection of complex or low-and-slow attacks under controlled conditions [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>], yet deployment-scale robustness remains limited. As shown in <xref ref-type="table" rid="table-7">Tables 7</xref> and <xref ref-type="table" rid="table-16">16</xref>, most IDS-oriented studies rely on curated or synthetic datasets that inadequately reflect class imbalance, noise, and concept drift in enterprise environments [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Systems optimised for benchmark performance therefore risk elevated false-positive rates and reduced sensitivity to novel or evolving attacks in production [<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-105">105</xref>]. Moreover, prompt-based and instruction-tuned IDS components remain sensitive to input formatting and contextual framing, raising robustness concerns when exposed to malformed logs, partial traces, or adversarially manipulated inputs [<xref ref-type="bibr" rid="ref-64">64</xref>]. These constraints currently limit the viability of LLM-based IDS for high-throughput or real-time deployment without continuous human oversight.</p>

<p>The limitations are most acute in autonomous and agentic response systems. Multi-agent LLM frameworks aim to enable closed-loop defence through automated investigation, containment planning, and response execution [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]. However, comparative evidence in <xref ref-type="table" rid="table-12">Tables 12</xref> and <xref ref-type="table" rid="table-13">13</xref> highlights substantial safety and verification challenges. Coordination failures between agents, propagation of hallucinated reasoning steps, and ambiguity in natural-language policy interpretation can lead to inappropriate or unsafe response actions [<xref ref-type="bibr" rid="ref-130">130</xref>,<xref ref-type="bibr" rid="ref-131">131</xref>]. Unlike rule-based automation, LLM-driven agents lack formal correctness and safety guarantees, complicating verification under adversarial or unforeseen conditions [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. These risks are particularly severe in cyber-physical and critical infrastructure environments&#x2014;such as smart grids and industrial control systems&#x2014;where erroneous actions may cause service disruption or physical harm [<xref ref-type="bibr" rid="ref-47">47</xref>]. Consequently, current agentic defence architectures are best understood as decision-support systems rather than fully autonomous responders.</p>

<p>Beyond technical constraints, the literature reveals socio-technical barriers that directly constrain trustworthy deployment and are central to <bold>RQ4</bold>. Integration costs remain substantial, encompassing compute-intensive inference, continuous model adaptation, RAG corpus curation, and tooling integration within existing SOC infrastructures [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Analyst trust calibration further complicates adoption, as probabilistic outputs, non-deterministic explanations, and hallucination risk can lead to either over-reliance or systematic distrust [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>]. Organisational resistance is particularly pronounced in regulated and safety-critical environments, where accountability, auditability, and liability concerns limit acceptance of AI-assisted workflows lacking enforceable oversight [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Together, these frictions explain why many LLM-based defences remain confined to decision-support roles despite strong benchmark performance.</p>
</sec>
<sec id="s9_2">
<label>9.2</label>
<title>Hallucination, Reliability, and Decision-Support Risks</title>
<p>Across nearly all LLM-enabled cybersecurity workflows, hallucination and output instability represent foundational risks. Fabricated indicators, incorrect remediation guidance, and inconsistent vulnerability explanations can mislead analysts when LLMs are embedded in SOC pipelines, CTI extraction, or multi-agent coordination. Empirical studies show that grounding failures in retrieval-augmented systems propagate confident yet incorrect conclusions [<xref ref-type="bibr" rid="ref-128">128</xref>], while surveys of generative AI for CTI report persistent reliability risks during data processing and report generation [<xref ref-type="bibr" rid="ref-132">132</xref>]. Similar issues arise in vulnerability repair [<xref ref-type="bibr" rid="ref-22">22</xref>] and automated SOC triage [<xref ref-type="bibr" rid="ref-23">23</xref>], with error rates increasing under incomplete telemetry, evolving malware semantics, and collaborative agent workflows [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-103">103</xref>,<xref ref-type="bibr" rid="ref-129">129</xref>,<xref ref-type="bibr" rid="ref-130">130</xref>]. Despite emerging work on calibration and uncertainty estimation, security-specific reliability frameworks capable of supporting high-assurance decision-making remain largely absent.</p>
</sec>
<sec id="s9_3">
<label>9.3</label>
<title>Adversarial Fragility and Alignment Limits</title>
<p>Adversarial prompting, jailbreaks, and prompt injection remain among the most consistently demonstrated vulnerabilities. Large-scale evaluations show that role-play manipulation, persona steering, and multi-turn coercion can bypass alignment safeguards even in heavily guarded models [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>]. Obfuscation-based and syntactic jailbreaks exhibit high transferability across model sizes and training regimes [<xref ref-type="bibr" rid="ref-116">116</xref>]. Indirect prompt injection&#x2014;via untrusted documents, code, or retrieved web content&#x2014;remains particularly effective in RAG and tool-augmented systems [<xref ref-type="bibr" rid="ref-117">117</xref>,<xref ref-type="bibr" rid="ref-166">166</xref>]. Lifecycle-oriented analyses further demonstrate how injection vectors propagate across retrieval chains, agent communication layers, and automated planners [<xref ref-type="bibr" rid="ref-27">27</xref>]. Collectively, these findings indicate that model-level alignment is insufficient in isolation and must be complemented by system-level safeguards, including input validation, provenance tracking, sandboxing, and continuous red-teaming.</p>
</sec>
<sec id="s9_4">
<label>9.4</label>
<title>Data Governance, Privacy, and Model Adaptation Risks</title>
<p>Data governance and privacy risks remain substantially under-addressed in LLM-based cybersecurity systems. SOC logs, malware traces, and proprietary source code frequently contain sensitive operational information, creating exposure risks through memorisation, cross-user leakage, and unintended disclosure [<xref ref-type="bibr" rid="ref-33">33</xref>]. Studies on organisational fine-tuning and domain adaptation show that incorporating sensitive datasets into LLM training pipelines introduces additional compliance and leakage risks [<xref ref-type="bibr" rid="ref-120">120</xref>]. Lifecycle analyses further demonstrate how privacy vulnerabilities propagate across chained RAG and multi-agent systems [<xref ref-type="bibr" rid="ref-27">27</xref>]. Despite these concerns, the surveyed literature reports limited adoption of differential privacy, secure inference, federated deployment, or redaction-aware training, leaving organisations exposed to unresolved regulatory and operational risks.</p>
</sec>
<sec id="s9_5">
<label>9.5</label>
<title>Evaluation Gaps and the Absence of Standards</title>
<p>Security evaluation practices remain fragmented and insufficiently standardised. Existing benchmarks capture isolated threat vectors&#x2014;such as phishing realism [<xref ref-type="bibr" rid="ref-6">6</xref>], harmful-query handling [<xref ref-type="bibr" rid="ref-137">137</xref>], or scenario-based deception [<xref ref-type="bibr" rid="ref-138">138</xref>]&#x2014;but rarely incorporate multimodal, multilingual, or adaptive adversarial dynamics. Red-teaming methodologies vary widely, from manual probing to automated jailbreak testing and multi-step manipulation, yet lack shared coverage criteria or reporting standards [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>]. Frequent vendor-side model updates further invalidate prior evaluation results, underscoring the need for lifecycle-integrated and continuously updated assessment frameworks.</p>
</sec>
<sec id="s9_6">
<label>9.6</label>
<title>Regulatory, Ethical, and Governance Challenges</title>
<p>Regulatory and governance frameworks lag behind the technical complexity of LLM deployment in cybersecurity contexts. Existing policies emphasise transparency and accountability but provide limited guidance on adversarial robustness, autonomous agent behaviour, or cross-organisational data governance. Analyses of digital deception and systemic risk highlight gaps in current governance models for managing interconnected LLM pipelines and agentic systems [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. Ethical principles such as fairness, explainability, and auditability remain difficult to operationalise in security-critical environments, leaving unresolved questions around liability, oversight, and risk ownership.</p>
</sec>
<sec id="s9_7">
<label>9.7</label>
<title>Synthesis of Cross-Cutting Challenges</title>
<p>Across the surveyed literature, a consistent pattern emerges: AIGC-enabled offensive capabilities evolve more rapidly than defensive maturity, evaluation frameworks lack operational breadth and standardisation, and governance mechanisms trail both technical innovation and adversarial adaptation. Reliability failures, fragmented red-teaming practices, and limited lifecycle-oriented risk modelling collectively amplify systemic fragility. Addressing these gaps requires an integrated, end-to-end perspective that combines robust threat modelling, provenance validation, privacy-preserving design, continuous adversarial evaluation, and standardised metrics aligned with real-world deployment constraints.</p>
</sec>
</sec>
<sec id="s10">
<label>10</label>
<title>Future Research Directions</title>
<p>The synthesis of 167 peer-reviewed studies spanning AIGC-driven threats, LLM-enabled defensive capabilities, and security evaluation frameworks reveals a research landscape that is advancing rapidly yet remains structurally imbalanced. While offensive applications of LLMs&#x2014;including phishing automation [<xref ref-type="bibr" rid="ref-6">6</xref>], polymorphic malware and exploit generation [<xref ref-type="bibr" rid="ref-8">8</xref>], jailbreak and alignment evasion [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-116">116</xref>], and large-scale misinformation and influence operations [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>]&#x2014;have progressed in sophistication, scalability, and autonomy, corresponding advances in defensive systems, evaluation methodologies, and governance mechanisms have not occurred at a comparable pace. As demonstrated throughout <xref ref-type="sec" rid="s5">Sections 5</xref>&#x2013;<xref ref-type="sec" rid="s7">7</xref>, this imbalance manifests in brittle defensive deployments, fragmented and benchmark-centric evaluation practices, and governance discussions that remain weakly coupled to enforceable technical controls.</p>
<p>A central finding of this review is that many limitations observed in current LLM-based cybersecurity systems do not arise from isolated algorithmic deficiencies, but from systemic gaps spanning the model&#x2013;system&#x2013;evaluation&#x2013;governance lifecycle. Evaluation practices remain difficult to compare due to inconsistent metrics, narrow benchmark assumptions, and limited consideration of adaptive adversarial behaviour. Concurrently, emerging deployment paradigms&#x2014;such as retrieval-augmented generation (RAG), tool-augmented reasoning, and multi-agent defence architectures&#x2014;introduce novel failure modes that are insufficiently addressed by existing assurance and governance frameworks [<xref ref-type="bibr" rid="ref-27">27</xref>]. These structural asymmetries amplify systemic risk and motivate the need for coordinated, evidence-driven research priorities that integrate technical, organisational, and regulatory perspectives.</p>
<p>To address these challenges in a principled and actionable manner, this review advances a structured future research agenda that explicitly links empirically observed weaknesses to targeted research interventions. <xref ref-type="table" rid="table-19">Table 19</xref> consolidates these priorities into a unified roadmap organised around five interdependent research pillars, capturing the technical, operational, evaluative, and governance dimensions of LLM-enabled cybersecurity. Although the pillars are framed in technical terms, their realisation inherently depends on interdisciplinary enablers&#x2014;most notably human&#x2013;computer interaction (HCI), cognitive ergonomics, policy studies, and regulatory science&#x2014;particularly for analyst-facing systems and compliance-critical deployments.</p>
<table-wrap id="table-19">
<label>Table 19</label>
<caption>
<title>Roadmap of key research directions and solutions for LLM-enabled cybersecurity.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Research Pillar</th>
<th>Key Challenges</th>
<th>Research Directions</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>1. Domain-Specialised &#x0026; Security-Aligned LLMs</bold></td>
<td>Hallucinated exploit reasoning and inconsistent remediation in generic models [<xref ref-type="bibr" rid="ref-16">16</xref>]; limited grounding in security ontologies (e.g., CVE, ATT&#x0026;CK, SIGMA); brittleness of RAG pipelines in CTI extraction [<xref ref-type="bibr" rid="ref-55">55</xref>,<xref ref-type="bibr" rid="ref-128">128</xref>]; need for customisable vulnerability detection and repair.</td>
<td>Neuro-symbolic architectures integrating attack graphs and security taxonomies; security-aware pre-training and retrieval-anchored verification; policy-constrained decoding enforcing cyber-specific rules; memory-scoped and isolation-aware reasoning; auditable reasoning traces for forensic analysis.</td>
</tr>
<tr>
<td><bold>2. Robust Multi-Agent &#x0026; Tool-Augmented Systems</bold></td>
<td>Error and hallucination propagation across agent teams [<xref ref-type="bibr" rid="ref-24">24</xref>]; tool misuse and privilege escalation [<xref ref-type="bibr" rid="ref-117">117</xref>]; lack of verifiable inter-agent communication protocols [<xref ref-type="bibr" rid="ref-65">65</xref>]; fragility in alert prioritisation and risk assessment [<xref ref-type="bibr" rid="ref-60">60</xref>].</td>
<td>Formal agent-role and privilege specifications; symbolic verification and model checking of workflows; sandboxed tool execution with runtime invariants; confidence-aware human-in-the-loop escalation for high-impact decisions.</td>
</tr>
<tr>
<td><bold>3. Next-Generation Evaluation &#x0026; Red-Teaming</bold></td>
<td>Static, text-centric benchmarks; lack of adversary-adaptive and longitudinal robustness assessment; fragmented reporting of robustness, refusal fidelity, and hallucination metrics.</td>
<td>Multilingual and multimodal adversary-adaptive benchmarks; automated red-teaming agents with evolving attack strategies; longitudinal audits across model updates; shared repositories of failure cases and safety regressions.</td>
</tr>
<tr>
<td><bold>4. Privacy, Data Governance &#x0026; Lifecycle Management</bold></td>
<td>Memorisation and leakage of sensitive SOC data [<xref ref-type="bibr" rid="ref-33">33</xref>]; vulnerability propagation across RAG layers and agent pipelines [<xref ref-type="bibr" rid="ref-27">27</xref>].</td>
<td>Differential privacy for SOC telemetry; secure enclaves and encrypted inference; redaction-aware training; tamper-evident provenance and audit logs; lifecycle-aware data governance policies.</td>
</tr>
<tr>
<td><bold>5. Governance, Regulation &#x0026; Accountability</bold></td>
<td>Misalignment between technical capabilities and legal expectations [<xref ref-type="bibr" rid="ref-13">13</xref>]; absence of certification standards; limited mechanisms for auditing autonomous decisions.</td>
<td>Certification and conformity standards for LLM-based defence tools; risk-scoring aligned with regulatory harm categories; traceable audit interfaces; multi-stakeholder governance frameworks.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Crucially, the proposed agenda emphasises that meaningful progress must extend across the full LLM lifecycle. Early-stage advances in domain-specialised and security-aligned model design must be complemented by adversary-aware evaluation, system-level verification, runtime monitoring, and compliance-ready governance mechanisms. As evidenced by the scenario-grounded limitations discussed in <xref ref-type="sec" rid="s6_9">Section 6.9</xref>, failure to address any single layer&#x2014;such as evaluation realism, human trust calibration, or policy enforcement&#x2014;can undermine otherwise strong technical performance. By grounding each research direction in recurring limitations identified across the evidence base, the roadmap promotes methodologically rigorous advances while encouraging deployment practices that integrate technical safeguards with procedural, organisational, and human-centred controls.</p>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> visualises this agenda as a vertically structured roadmap organised around five interconnected thematic pillars: (i) domain-specialised and security-aligned LLMs, (ii) robust multi-agent and tool-augmented defence architectures, (iii) next-generation evaluation and red-teaming ecosystems, (iv) privacy-preserving data governance and lifecycle management, and (v) governance- and regulation-aligned assurance frameworks. Notably, several pillars explicitly require interdisciplinary collaboration: HCI and cognitive ergonomics are essential for explanation design, trust calibration, and analyst&#x2013;LLM interaction within SOC environments, while policy studies and regulatory science underpin the translation of technical safeguards into auditable and enforceable compliance mechanisms.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Research roadmap for LLM-enabled cybersecurity, outlining future directions across domain-specialised security-aligned models, robust multi-agent and tool-augmented systems, adversary-aware evaluation and red-teaming, privacy- and lifecycle-aware data governance, and governance- and regulation-aligned assurance frameworks.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77367-fig-4.tif"/>
</fig>
<p><xref ref-type="table" rid="table-19">Table 19</xref> operationalises this vision by mapping each research pillar to the corresponding empirical challenges and prioritised research directions. A key insight emerging from this synthesis is that resilience cannot be achieved through progress along a single dimension. Instead, robust and trustworthy LLM-enabled cybersecurity will depend on coordinated advances across modelling, system integration, evaluation realism, human&#x2013;system interaction design, privacy protection, and governance alignment&#x2014;an insight that directly informs both near-term research investment and longer-term standardisation and policy efforts.</p>

<p>Beyond this high-level roadmap, the literature consistently points to the need for prioritisation across time horizons. Near-term research should focus on operational realism, evaluation standardisation, and safe hybrid deployment models that incorporate explicit human oversight. Mid-term efforts must address systemic robustness in multi-agent and tool-augmented systems, alongside privacy-preserving lifecycle management. Long-term investments are required to develop domain-specialised LLMs with formal reasoning guarantees and to establish certification and assurance regimes suitable for safety-critical and national infrastructure contexts.</p>
<p>To make this prioritisation explicit, <xref ref-type="table" rid="table-20">Table 20</xref> aligns the proposed research directions with the four research questions (RQ1-RQ4) and categorises them by expected time horizon and impact. This staged alignment ensures that future work not only advances the state of the art but also systematically closes the methodological, operational, and governance gaps identified throughout the review.</p>
<table-wrap id="table-20">
<label>Table 20</label>
<caption>
<title>Prioritised future research roadmap explicitly aligned with research questions (RQ1&#x2013;RQ4).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Time Horizon</th>
<th>Research Priority</th>
<th>RQ Link</th>
<th>Rationale and Expected Impact</th>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td>Operationally grounded evaluation and red-teaming ecosystems</td>
<td>RQ3</td>
<td>Improves robustness assessment and reproducibility through adversary-adaptive, multilingual, and SOC-aligned benchmarks, directly addressing fragmentation in current evaluation practices.</td>
</tr>
<tr>
<td><bold>Near-Term (1&#x2013;3 years)</bold></td>
<td>Hybrid IDS architectures with explicit human-in-the-loop controls</td>
<td>RQ2</td>
<td>Reflects evidence that LLMs are most effective as post-detection reasoning layers, mitigating hallucination and false-positive escalation while enabling safe deployment.</td>
</tr>
<tr>
<td></td>
<td>Governance-aligned deployment safeguards</td>
<td>RQ4</td>
<td>Operationalises ethical and regulatory principles via audit logs, escalation workflows, and compliance-ready controls.</td>
</tr>
<tr>
<td rowspan="2"><bold>Mid-Term (3&#x2013;5 years)</bold></td>
<td>Verified multi-agent and tool-augmented defence systems</td>
<td>RQ2, RQ3</td>
<td>Targets systemic failure modes such as error propagation and tool misuse, advancing verification, runtime monitoring, and controlled autonomy.</td>
</tr>
<tr>

<td>Privacy-preserving lifecycle and data governance mechanisms</td>
<td>RQ4</td>
<td>Addresses data leakage and compliance risks, enabling trustworthy deployment in regulated and long-lived environments.</td>
</tr>
<tr>
<td><bold>Long-Term (5&#x002B; years)</bold></td>
<td>Domain-specialised, security-aligned LLMs with formal reasoning guarantees</td>
<td>RQ1, RQ2</td>
<td>Moves beyond prompt-level mitigation toward foundational solutions for hallucination, exploit reasoning errors, and misuse risk.</td>
</tr>
<tr>
<td></td>
<td>Standardised certification and assurance frameworks</td>
<td>RQ4</td>
<td>Enables adoption in safety-critical contexts by linking technical robustness to governance and accountability structures.</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-20fn1" fn-type="other">
<p>Note: RQ1 addresses AIGC-enabled threats, RQ2 defensive capabilities, RQ3 evaluation and benchmarking, and RQ4 governance and systemic challenges.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s11">
<label>11</label>
<title>Conclusion</title>
<p>This PRISMA-guided systematic review synthesised 167 peer-reviewed studies published between 2022 and 2025 to examine the transformative impact of Large Language Models (LLMs) on cybersecurity across four tightly coupled dimensions: threat generation, defensive capabilities, security evaluation, and governance. Anchored in a unified threat&#x2013;defence&#x2013;evaluation taxonomy and supported by transparent screening, coding, and synthesis procedures, the review provides an evidence-driven response to the four research questions (RQ1&#x2013;RQ4) that motivated this study.</p>
<p>In addressing <bold>RQ1</bold>&#x2014;how LLMs enable or amplify cybersecurity threats&#x2014;the review demonstrates that AIGC-driven attacks represent a qualitative shift rather than an incremental extension of prior techniques. Across phishing and social engineering, malware and exploit development, jailbreaks and prompt injection, and large-scale misinformation and influence operations, LLMs introduce unprecedented scalability, contextual awareness, and adaptability. These properties substantially lower the technical barrier for adversaries while eroding the effectiveness of traditional signature-based detection, stylometric attribution, and static defence mechanisms. The convergence of multilingual generation, semantic reasoning, and rapid iteration fundamentally reshapes the attacker&#x2013;defender balance.</p>
<p>In response to <bold>RQ2</bold>, which examined the defensive role of LLMs, the evidence indicates that LLMs deliver their greatest value when deployed as <italic>cognitive and reasoning layers</italic> within hybrid security architectures. Meaningful advances are observed in threat intelligence extraction, SOC automation, alert correlation, anomaly interpretation, vulnerability analysis, and multi-agent investigation workflows. At the same time, the review identifies persistent and non-trivial limitations, including hallucinated indicators, false-positive escalation, sensitivity to noisy or adversarial inputs, and significant computational overhead. Crucially, LLMs rarely function effectively as primary detectors; instead, they augment conventional IDS and security tools through semantic abstraction, contextual explanation, and decision support. These findings reinforce the necessity of defence-in-depth designs and robust human-in-the-loop oversight for operational deployment.</p>
<p>With respect to <bold>RQ3</bold>, which focused on security evaluation frameworks and red-teaming methodologies, the review exposes a fragmented and methodologically uneven evaluation landscape. Existing benchmarks are predominantly text-centric, static, and weakly aligned with real-world SOC, IDS, and OT/ICS operating conditions. Red-teaming practices vary widely in scope, reproducibility, and reporting rigor, while frequent model updates undermine longitudinal robustness assessment. Comparative synthesis demonstrates that accuracy-centric metrics alone are insufficient; operational indicators such as inference latency, false-positive amplification, reasoning stability, and system-level failure propagation provide more meaningful insight into deployment readiness. The absence of standardised, adversary-aware, and lifecycle-oriented evaluation pipelines remains a critical barrier to trustworthy assessment.</p>
<p>Addressing <bold>RQ4</bold>, which examined methodological and governance challenges, the review identifies systemic gaps that constrain the safe and accountable deployment of LLM-based cybersecurity systems. These include overreliance on synthetic or laboratory datasets, limited analysis of real-world deployment dynamics, insufficient reporting of operational cost and energy impact, and weak integration between technical design decisions and regulatory obligations. Although ethical considerations are frequently discussed, they are often decoupled from enforceable controls. By explicitly mapping observed failure modes to auditable safeguards and regulatory frameworks&#x2014;such as the EU NIS-2 Directive and guidance from CISA&#x2014;the review demonstrates that governance must be operationalised through accountability, transparency, human oversight, auditability, and change-management mechanisms rather than treated as an abstract overlay.</p>
<p>Taken together, the answers to RQ1&#x2013;RQ4 reveal that LLMs constitute a dual-use inflection point in cybersecurity. They simultaneously accelerate offensive sophistication and enable higher-level defensive reasoning, shifting security practice from isolated signal-level detection toward context-aware, explainable, and decision-oriented defence systems. The central contribution of this review lies in systematically integrating these dimensions through a unified taxonomy and critical synthesis, clarifying the trade-offs among capability, robustness, cost, and governance. By connecting empirical evidence with evaluation realism and regulatory alignment, this work provides a structured foundation for future research and practice aimed at developing secure, accountable, and resilient LLM-enabled cybersecurity systems.</p>
</sec>
</body>
<back>
<ack>
<p>The authors acknowledge support from their respective institutions.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The author, Hamed Alqahtani extends his appreciation to the Deanship of Scientific Research at King Khalid University for funding this work through large group under grant number (GRP.2/663/46).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Hamed Alqahtani and Gulshan Kumar; methodology, Hamed Alqahtani and Gulshan Kumar; validation, Hamed Alqahtani and Gulshan Kumar; formal analysis, Hamed Alqahtani; investigation, Hamed Alqahtani; data curation, Hamed Alqahtani; writing&#x2014;original draft preparation, Hamed Alqahtani; writing&#x2014;review and editing, Hamed Alqahtani and Gulshan Kumar; visualization, Hamed Alqahtani; supervision, Gulshan Kumar; project administration, Gulshan Kumar. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>This article does not involve data availability, and this section is not applicable.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Brown</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mann</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ryder</surname> <given-names>N</given-names></string-name>, <string-name><surname>Subbiah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kaplan</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Language models are few-shot learners</article-title>. <source>Adv Neural Inform Process Syst</source>. <year>2020</year>;<volume>33</volume>:<fpage>1877</fpage>&#x2013;<lpage>901</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2021.mrl-1.1</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alqahtani</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Large language models for effective detection of algorithmically generated domains: a comprehensive review</article-title>. <source>Comput Model Eng Sci</source>. <year>2025</year>;<volume>144</volume>(<issue>2</issue>):<fpage>1439</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2025.067738</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jaffal</surname> <given-names>NO</given-names></string-name>, <string-name><surname>Alkhanafseh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mohaisen</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Large language models in cybersecurity: a survey of applications, vulnerabilities, and defense techniques</article-title>. <source>AI</source>. <year>2025</year>;<volume>6</volume>(<issue>9</issue>):<fpage>216</fpage>. doi:<pub-id pub-id-type="doi">10.3390/ai6090216</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Karras</surname> <given-names>A</given-names></string-name>, <string-name><surname>Theodorakopoulos</surname> <given-names>L</given-names></string-name>, <string-name><surname>Karras</surname> <given-names>C</given-names></string-name>, <string-name><surname>Theodoropoulou</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kalliampakou</surname> <given-names>I</given-names></string-name>, <string-name><surname>Kalogeratos</surname> <given-names>G</given-names></string-name></person-group>. <article-title>LLMs for cybersecurity in the big data era: a comprehensive review of applications, challenges, and future directions</article-title>. <source>Information</source>. <year>2025</year>;<volume>16</volume>(<issue>11</issue>):<fpage>957</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info16110957</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Agarwal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Trivedi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bhardwaj</surname> <given-names>A</given-names></string-name></person-group>. <chapter-title>Use cases of ChatGPT and other AI tools with security concerns</chapter-title>. In: <source>Applications, challenges, and the future of ChatGPT</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2024</year>. p. <fpage>216</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3693-6824-4.ch012</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Opara</surname> <given-names>C</given-names></string-name>, <string-name><surname>Modesti</surname> <given-names>P</given-names></string-name>, <string-name><surname>Golightly</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Evaluating spam filters and Stylometric Detection of AI-generated phishing emails</article-title>. <source>Expert Syst Appl</source>. <year>2025</year>;<volume>276</volume>:<fpage>127044</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2025.127044</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Aseeri</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Bohacek</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Using ensembles of LLMs to detect phishing emails</article-title>. In: <conf-name>International Conference on Advanced Information Networking and Applications</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>21</fpage>&#x2013;<lpage>34</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Iturbe</surname> <given-names>E</given-names></string-name>, <string-name><surname>Llorente-Vazquez</surname> <given-names>O</given-names></string-name>, <string-name><surname>Rego</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rios</surname> <given-names>E</given-names></string-name>, <string-name><surname>Toledo</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Unleashing offensive artificial intelligence: automated attack technique code generation</article-title>. <source>Comput Secur</source>. <year>2024</year>;<volume>147</volume>:<fpage>104077</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2024.104077</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yamin</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Hashmi</surname> <given-names>E</given-names></string-name>, <string-name><surname>Katt</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Combining uncensored and censored LLMS for ransomware generation</article-title>. In: <conf-name>International Conference on Web Information Systems Engineering</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>189</fpage>&#x2013;<lpage>202</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shandilya</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Prharsha</surname> <given-names>G</given-names></string-name>, <string-name><surname>Datta</surname> <given-names>A</given-names></string-name>, <string-name><surname>Choudhary</surname> <given-names>G</given-names></string-name>, <string-name><surname>Park</surname> <given-names>H</given-names></string-name>, <string-name><surname>You</surname> <given-names>I</given-names></string-name></person-group>. <article-title>GPT based malware: unveiling vulnerabilities and creating a way forward in digital space</article-title>. In: <conf-name>2023 International Conference on Data Security and Privacy Protection (DSPP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>164</fpage>&#x2013;<lpage>73</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hyun</surname> <given-names>J</given-names></string-name></person-group>. <chapter-title>Utilizing prompt engineering to operationalize cybersecurity</chapter-title>. In: <source>Generative AI security: theories and practices</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>271</fpage>&#x2013;<lpage>303</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-54252-7_9</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Alauthman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Almomani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Aoudi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Al-Qerem</surname> <given-names>A</given-names></string-name>, <string-name><surname>Aldweesh</surname> <given-names>A</given-names></string-name></person-group>. <chapter-title>Automated vulnerability discovery generative AI in offensive security</chapter-title>. In: <source>Examining cybersecurity risks produced by generative AI</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>309</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3373-0832-6.ch013</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Schmitt</surname> <given-names>M</given-names></string-name>, <string-name><surname>Flechais</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Digital deception: generative artificial intelligence in social engineering and phishing</article-title>. <source>Artif Intell Rev</source>. <year>2024</year>;<volume>57</volume>(<issue>12</issue>):<fpage>324</fpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Valdez</surname> <given-names>HPD</given-names></string-name>, <string-name><surname>Abri</surname> <given-names>F</given-names></string-name>, <string-name><surname>Webb</surname> <given-names>J</given-names></string-name>, <string-name><surname>Austin</surname> <given-names>TH</given-names></string-name></person-group>. <article-title>Exploring the use and misuse of large language models</article-title>. <source>Information</source>. <year>2025</year>;<volume>16</volume>(<issue>9</issue>):<fpage>758</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info16090758</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A survey on large language model (LLM) security and privacy: the good, the bad, and the ugly</article-title>. <source>High-Confid Comput</source>. <year>2024</year>;<volume>4</volume>(<issue>2</issue>):<fpage>100211</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.hcc.2024.100211</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fei</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xi</surname> <given-names>R</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>When LLMS meet cybersecurity: a systematic literature review</article-title>. <source>Cybersecurity</source>. <year>2025</year>;<volume>8</volume>(<issue>1</issue>):<fpage>55</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s42400-025-00361-w</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Page</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>McKenzie</surname> <given-names>JE</given-names></string-name>, <string-name><surname>Bossuyt</surname> <given-names>PM</given-names></string-name>, <string-name><surname>Boutron</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hoffmann</surname> <given-names>TC</given-names></string-name>, <string-name><surname>Mulrow</surname> <given-names>CD</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The PRISMA, 2020 statement: an updated guideline for reporting systematic reviews</article-title>. <source>BMJ</source>. <year>2021</year>;<volume>372</volume>. doi:<pub-id pub-id-type="doi">10.31222/osf.io/v7gm2_v1</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Elouardi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Motii</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jouhari</surname> <given-names>M</given-names></string-name>, <string-name><surname>Amadou</surname> <given-names>ANH</given-names></string-name>, <string-name><surname>Hedabou</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A survey on Hybrid-CNN and LLMs for intrusion detection systems: recent IoT datasets</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>1</issue>):<fpage>180009</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3506604</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huynh</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jayasundera</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jeon</surname> <given-names>W</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bi</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Detecting code vulnerabilities using LLMs</article-title>. In: <conf-name>2025 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>401</fpage>&#x2013;<lpage>14</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Loumachi</surname> <given-names>FY</given-names></string-name>, <string-name><surname>Ghanem</surname> <given-names>MC</given-names></string-name>, <string-name><surname>Ferrag</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Advancing cyber incident timeline analysis through retrieval-augmented generation and large language models</article-title>. <source>Computers</source>. <year>2025</year>;<volume>14</volume>(<issue>67</issue>):<fpage>1</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.3390/computers14020067</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Sarker</surname> <given-names>IH</given-names></string-name></person-group>. <chapter-title>Introduction to AI-driven cybersecurity and threat intelligence</chapter-title>. In: <source>AI-driven cybersecurity and threat intelligence: cyber automation, intelligent decision-making and explainability</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>3</fpage>&#x2013;<lpage>19</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>de Fitero-Dominguez</surname> <given-names>D</given-names></string-name>, <string-name><surname>Garcia-Lopez</surname> <given-names>E</given-names></string-name>, <string-name><surname>Garcia-Cabot</surname> <given-names>A</given-names></string-name>, <string-name><surname>Martinez-Herraiz</surname> <given-names>JJ</given-names></string-name></person-group>. <article-title>Enhanced automated code vulnerability repair using large language models</article-title>. <source>Eng Appl Artif Intell</source>. <year>2024</year>;<volume>138</volume>(<issue>1</issue>):<fpage>109291</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2024.109291</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Srinivas</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kirk</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zendejas</surname> <given-names>J</given-names></string-name>, <string-name><surname>Espino</surname> <given-names>M</given-names></string-name>, <string-name><surname>Boskovich</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bari</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AI-augmented SOC: a survey of LLMs and agents for security automation</article-title>. <source>J Cybersecur Priv</source>. <year>2025</year>;<volume>5</volume>(<issue>4</issue>):<fpage>95</fpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Andreoni</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lunardi</surname> <given-names>WT</given-names></string-name>, <string-name><surname>Lawton</surname> <given-names>G</given-names></string-name>, <string-name><surname>Thakkar</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Enhancing autonomous system security and resilience with generative AI: a comprehensive survey</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>1</issue>):<fpage>109470</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3439363</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>P</given-names></string-name>, <string-name><surname>London</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tenney</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Benchmarking and evaluating large language models in phishing detection for small and midsize enterprises: a comprehensive analysis</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>2</issue>):<fpage>28335</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3540075</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hilario</surname> <given-names>E</given-names></string-name>, <string-name><surname>Azam</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sundaram</surname> <given-names>J</given-names></string-name>, <string-name><surname>Imran Mohammed</surname> <given-names>K</given-names></string-name>, <string-name><surname>Shanmugam</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Generative AI for pentesting: the good, the bad, the ugly</article-title>. <source>Int J Inform Secur</source>. <year>2024</year>;<volume>23</volume>(<issue>3</issue>):<fpage>2075</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10207-024-00835-x</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>De Maio</surname> <given-names>C</given-names></string-name>, <string-name><surname>Di Gisi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fenza</surname> <given-names>G</given-names></string-name>, <string-name><surname>Gallo</surname> <given-names>M</given-names></string-name>, <string-name><surname>Loia</surname> <given-names>V</given-names></string-name></person-group>. <article-title>A lifecycle-oriented survey of emerging threats and vulnerabilities in large language models</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>176482</fpage>&#x2013;<lpage>500</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3619764</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Y</given-names></string-name>, <string-name><surname>TJj</surname> <given-names>Li</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Clear: towards contextual LLM-empowered privacy policy analysis and risk generation for large language model applications</article-title>. In: <conf-name>Proceedings of the 30th International Conference on Intelligent User Interfaces</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2025</year>. p. <fpage>277</fpage>&#x2013;<lpage>97</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Campanile</surname> <given-names>L</given-names></string-name>, <string-name><surname>De Fazio</surname> <given-names>R</given-names></string-name>, <string-name><surname>Di Giovanni</surname> <given-names>M</given-names></string-name>, <string-name><surname>Marulli</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Beyond the hype: toward a concrete adoption of the fair and responsible use of AI</article-title>. In: <conf-name>Ital-IA 2024: 4th National Conference on Artificial Intelligence; 2024 May 29&#x2013;30</conf-name>; <publisher-loc>Naples, Italy</publisher-loc>; <year>2024</year>. p. <fpage>60</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Humphreys</surname> <given-names>D</given-names></string-name>, <string-name><surname>Koay</surname> <given-names>A</given-names></string-name>, <string-name><surname>Desmond</surname> <given-names>D</given-names></string-name>, <string-name><surname>Mealy</surname> <given-names>E</given-names></string-name></person-group>. <article-title>AI hype as a cyber security risk: the moral responsibility of implementing generative AI in business</article-title>. <source>AI Ethics</source>. <year>2024</year>;<volume>4</volume>(<issue>3</issue>):<fpage>791</fpage>&#x2013;<lpage>804</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>The ethical security of large language models: a systematic review</article-title>. <source>Front Eng Manag</source>. <year>2025</year>;<volume>12</volume>(<issue>1</issue>):<fpage>128</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s42524-025-4082-6</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Villa</surname> <given-names>C</given-names></string-name>, <string-name><surname>Mirza</surname> <given-names>S</given-names></string-name>, <string-name><surname>P&#x00F6;pper</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Exposing the guardrails: reverse-engineering and jailbreaking safety filters in DALL&#x22C5;E text-to-image pipelines</article-title>. In: <conf-name>34th USENIX Security Symposium (USENIX Security 25)</conf-name>. <publisher-loc>Berkeley, CA, USA</publisher-loc>: <publisher-name>USENIX Association</publisher-name>; <year>2025</year>. p. <fpage>897</fpage>&#x2013;<lpage>916</lpage>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alawida</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mejri</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mehmood</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chikhaoui</surname> <given-names>B</given-names></string-name>, <string-name><surname>Isaac Abiodun</surname> <given-names>O</given-names></string-name></person-group>. <article-title>A comprehensive study of ChatGPT: advancements, limitations, and ethical considerations in natural language processing and cybersecurity</article-title>. <source>Information</source>. <year>2023</year>;<volume>14</volume>(<issue>8</issue>):<fpage>462</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info14080462</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Esmradi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yip</surname> <given-names>DW</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>CF</given-names></string-name></person-group>. <article-title>A comprehensive survey of attack techniques, implementation, and mitigation strategies in large language models</article-title>. In: <conf-name>International Conference on Ubiquitous Security</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>76</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lopez Delgado</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Lopez Ramos</surname> <given-names>JA</given-names></string-name></person-group>. <article-title>A comprehensive survey on generative AI solutions in IoT security</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>24</issue>):<fpage>4965</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13244965</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nawara</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kashef</surname> <given-names>R</given-names></string-name></person-group>. <article-title>A comprehensive survey on LLM-powered recommender systems: from discriminative, generative to multi-modal paradigms</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>145772</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3599832</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Du</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Backdoor attacks and countermeasures in natural language processing models: a comprehensive security review</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2025</year>;<volume>36</volume>(<issue>8</issue>):<fpage>13628</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2025.3540303</pub-id>; <pub-id pub-id-type="pmid">40031656</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Uddin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Arfeen</surname> <given-names>SU</given-names></string-name>, <string-name><surname>Alanazi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mazhar</surname> <given-names>T</given-names></string-name>, <string-name><surname>Arafatur Rahman</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A critical analysis of generative AI: challenges, opportunities, and future research directions</article-title>. <source>Arch Comput Methods Eng</source>. <year>2025</year>. doi:<pub-id pub-id-type="doi">10.1007/s11831-025-10355-z</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ferrag</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Ndhlovu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tihanyi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Cordeiro</surname> <given-names>LC</given-names></string-name>, <string-name><surname>Debbah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lestable</surname> <given-names>T</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Revolutionizing cyber threat detection with large language models: a privacy-preserving bert-based lightweight model for IoT/IIoT devices</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>23733</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3363469</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hasanov</surname> <given-names>I</given-names></string-name>, <string-name><surname>Virtanen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hakkala</surname> <given-names>A</given-names></string-name>, <string-name><surname>Isoaho</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Application of large language models in cybersecurity: a systematic literature review</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>1</issue>):<fpage>176751</fpage>&#x2013;<lpage>78</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3505983</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alqahtani</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A comprehensive review of generative AI techniques and their impact on cybersecurity</article-title>. <source>Soft Comput</source>. <year>2025</year>;<volume>29</volume>(<issue>13</issue>):<fpage>4945</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00500-025-10702-z</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gunda</surname> <given-names>M</given-names></string-name>, <string-name><surname>Manda</surname> <given-names>V</given-names></string-name>, <string-name><surname>Naradasu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Mekala</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bhattacharya</surname> <given-names>S</given-names></string-name></person-group>. <article-title>ChatGPT in cyber onslaught and fortification: past, present, and future</article-title>. In: <conf-name>2024 IEEE 9th International Conference for Convergence in Technology (I2CT)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Lebed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Namiot</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zubareva</surname> <given-names>E</given-names></string-name>, <string-name><surname>Khenkin</surname> <given-names>P</given-names></string-name>, <string-name><surname>Vorobeva</surname> <given-names>A</given-names></string-name>, <string-name><surname>Svichkar</surname> <given-names>D</given-names></string-name></person-group>. <chapter-title>Large language models in cyberattacks</chapter-title>. In: <source>Doklady mathematics</source>. Vol. <volume>110</volume>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>S510</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Siemerink</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jansen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Labunets</surname> <given-names>K</given-names></string-name></person-group>. <article-title>The dual-edged sword of large language models in phishing</article-title>. In: <conf-name>Nordic Conference on Secure IT Systems</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>258</fpage>&#x2013;<lpage>79</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>De Queiroz</surname> <given-names>HJDS</given-names></string-name></person-group>. <chapter-title>Phishing and social engineering attack prevention with LLMs</chapter-title>. In: <source>Revolutionizing cybersecurity with deep learning and large language models</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>133</fpage>&#x2013;<lpage>64</lpage>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Han</surname> <given-names>W</given-names></string-name></person-group>. <article-title>A survey of cybersecurity knowledge base and its automatic labeling</article-title>. In: <conf-name>International Conference on Network Simulation and Evaluation</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>53</fpage>&#x2013;<lpage>70</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Imanifard</surname> <given-names>A</given-names></string-name>, <string-name><surname>Majidi</surname> <given-names>B</given-names></string-name>, <string-name><surname>Shamisa</surname> <given-names>A</given-names></string-name></person-group>. <article-title>SmartGridAgent: an educational framework for reliable digital twin-based smart grid workforce training with locally hosted LLMs</article-title>. <source>Smart Grids Sustain Energy</source>. <year>2025</year>;<volume>10</volume>(<issue>2</issue>):<fpage>41</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s40866-025-00274-0</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A review of advancements and applications of pre-trained language models in cybersecurity</article-title>. In: <conf-name>2024 12th International Symposium on Digital Forensics and Security (ISDFS)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Motlagh</surname> <given-names>FN</given-names></string-name>, <string-name><surname>Hajizadeh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Majd</surname> <given-names>M</given-names></string-name>, <string-name><surname>Najafi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Meinel</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Large language models in cybersecurity: state-of-the-art</article-title>. <comment>arXiv:2402.00891. 2024</comment>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Sarker</surname> <given-names>IH</given-names></string-name></person-group>. <chapter-title>CyberAI: a comprehensive summary of AI variants, explainable and responsible AI for cybersecurity</chapter-title>. In: <source>AI-driven cybersecurity and threat intelligence: cyber automation, intelligent decision-making and explainability</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>173</fpage>&#x2013;<lpage>200</lpage>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Palma</surname> <given-names>G</given-names></string-name>, <string-name><surname>Cecchi</surname> <given-names>G</given-names></string-name>, <string-name><surname>Caronna</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rizzo</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Leveraging large language models for scalable and explainable cybersecurity log analysis</article-title>. <source>J Cybersecur Priv</source>. <year>2025</year>;<volume>5</volume>(<issue>3</issue>):<fpage>55</fpage>. doi:<pub-id pub-id-type="doi">10.3390/jcp5030055</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>N</given-names></string-name>, <string-name><surname>Sinhababu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chakravarty</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Impact of conversational and generative AI systems on libraries: a use case large language model (LLM)</article-title>. <source>Sci Technol Librar</source>. <year>2024</year>;<volume>43</volume>(<issue>4</issue>):<fpage>319</fpage>&#x2013;<lpage>33</lpage>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mohsin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Janicke</surname> <given-names>H</given-names></string-name>, <string-name><surname>Nepal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Holmes</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Digital twins and the future of their use enabling shift left and shift right cybersecurity operations</article-title>. In: <conf-name>2023 5th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems and Applications (TPS-ISA)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>277</fpage>&#x2013;<lpage>86</lpage>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Luntovskyy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Winkler</surname> <given-names>U</given-names></string-name></person-group>. <article-title>Advanced software technology paradigms and AI deployment</article-title>. In: <conf-name>Intelligent Systems Conference</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>590</fpage>&#x2013;<lpage>605</lpage>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gandhi</surname> <given-names>ST</given-names></string-name></person-group>. <article-title>RAG-driven cybersecurity intelligence: leveraging semantic search for improved threat detection</article-title>. <source>Int J Res Appl Innov</source>. <year>2023</year>;<volume>6</volume>(<issue>3</issue>):<fpage>8889</fpage>&#x2013;<lpage>97</lpage>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shahid</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Hasan</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Kankanamge</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Hossain</surname> <given-names>MZ</given-names></string-name>, <string-name><surname>Imteaj</surname> <given-names>A</given-names></string-name></person-group>. <article-title>WatchOverGPT: a framework for real-time crime detection and response using wearable camera and large language model</article-title>. In: <conf-name>2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>2189</fpage>&#x2013;<lpage>94</lpage>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Moongela</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mayayise</surname> <given-names>T</given-names></string-name></person-group>. <article-title>The impact of large language models on cybersecurity</article-title>. In: <conf-name>Annual Conference of South African Institute of Computer Scientists and Information Technologists</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2026</year>. p. <fpage>129</fpage>&#x2013;<lpage>46</lpage>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Muzammal</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Mahadevappa</surname> <given-names>P</given-names></string-name>, <string-name><surname>Tayyab</surname> <given-names>M</given-names></string-name></person-group>. <chapter-title>Exploring security challenges in generative AI for web engineering</chapter-title>. In: <source>Generative AI for web engineering models</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global</publisher-name>; <year>2025</year>. p. <fpage>331</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3693-3703-5.ch016</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Babaey</surname> <given-names>V</given-names></string-name>, <string-name><surname>Ravindran</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Gensqli: a generative artificial intelligence framework for automatically securing web application firewalls against structured query language injection attacks</article-title>. <source>Future Internet</source>. <year>2024</year>;<volume>17</volume>(<issue>1</issue>):<fpage>8</fpage>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Seo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Rong</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Flexible and secure code deployment in federated learning using large language models: prompt engineering to enhance malicious code detection</article-title>. In: <conf-name>2023 IEEE International Conference on Cloud Computing Technology and Science (CloudCom)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>341</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Martell</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Baweja</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Dreslin</surname> <given-names>BD</given-names></string-name></person-group>. <article-title>Mitigative strategies for recovering from large language model trust violations</article-title>. <source>J Cognit Eng Decis Mak</source>. <year>2025</year>;<volume>19</volume>(<issue>1</issue>):<fpage>76</fpage>&#x2013;<lpage>95</lpage>. doi:<pub-id pub-id-type="doi">10.1177/15553434241303577</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Corchado</surname> <given-names>JM</given-names></string-name>, <string-name><surname>L&#x00F3;pez</surname> <given-names>S</given-names></string-name>, <string-name><surname>Garcia</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chamoso</surname> <given-names>P</given-names></string-name>, <string-name><surname>Juan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nunez</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Generative artificial intelligence: fundamentals</article-title>. <source>ADCAIJ</source>. <year>2023</year>;<volume>12</volume>(<issue>1</issue>):<fpage>e31704</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.14201/adcaij.31704</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koide</surname> <given-names>T</given-names></string-name>, <string-name><surname>Nakano</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chiba</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Chatphishdetector: detecting phishing sites using large language models</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>154381</fpage>&#x2013;<lpage>400</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3483905</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Prompting large language models for malicious webpage detection</article-title>. In: <conf-name>2023 IEEE 4th International Conference on Pattern Recognition and Machine Learning (PRML)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>393</fpage>&#x2013;<lpage>400</lpage>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Krishnamurthy</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Enhancing cyber security enhancement through generative AI</article-title>. <source>Int J Univ Sci Eng</source>. <year>2023</year>;<volume>9</volume>(<issue>1</issue>):<fpage>35</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chowdhury</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Rifat</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ahsan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Latif</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gomes</surname> <given-names>R</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>ChatGPT: a threat against the CIA triad of cyber security</article-title>. In: <conf-name>2023 IEEE International Conference on Electro Information Technology (eIT)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Iyengar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kundu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Large language models and computer security</article-title>. In: <conf-name>2023 5th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems and Applications (TPS-ISA)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>307</fpage>&#x2013;<lpage>13</lpage>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Security and privacy challenges of AIGC in metaverse: a comprehensive survey</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>57</volume>(<issue>10</issue>):<fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3729419</pub-id>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>McIntosh</surname> <given-names>TR</given-names></string-name>, <string-name><surname>Susnjak</surname> <given-names>T</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Watters</surname> <given-names>P</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>From google gemini to openai q&#x002A; (q-star): a survey on reshaping the generative artificial intelligence (AI) research landscape</article-title>. <source>Technologies</source>. <year>2025</year>;<volume>13</volume>(<issue>2</issue>):<fpage>51</fpage>. doi:<pub-id pub-id-type="doi">10.3390/technologies13020051</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pahuja</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kukreja</surname> <given-names>S</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Comprehensive review of generative artificial intelligence: mechanisms, models, and applications</article-title>. <source>Procedia Comput Sci</source>. <year>2025</year>;<volume>258</volume>(<issue>8</issue>):<fpage>3731</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.procs.2025.04.628</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bengesi</surname> <given-names>S</given-names></string-name>, <string-name><surname>El-Sayed</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sarker</surname> <given-names>MK</given-names></string-name>, <string-name><surname>Houkpati</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Irungu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Oladunni</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Advancements in generative AI: a comprehensive review of GANs, GPT, autoencoders, diffusion model, and transformers</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>69812</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3397775</pub-id>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Mohammed</surname> <given-names>D</given-names></string-name>, <string-name><surname>MacLennan</surname> <given-names>H</given-names></string-name></person-group>. <chapter-title>Secure authentication and identity management with AI</chapter-title>. In: <source>Revolutionizing cybersecurity with deep learning and large language models</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>271</fpage>&#x2013;<lpage>306</lpage>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>F</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Moving target defense meets artificial intelligence-driven network: a comprehensive survey</article-title>. <source>IEEE Int Things J</source>. <year>2025</year>;<volume>12</volume>(<issue>10</issue>):<fpage>13384</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JIOT.2025.3533016</pub-id>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sakib</surname> <given-names>MN</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Pathak</surname> <given-names>R</given-names></string-name>, <string-name><surname>Arifin</surname> <given-names>MM</given-names></string-name></person-group>. <article-title>Risks, causes, and mitigations of widespread deployments of large language models (LLMS): a survey</article-title>. In: <conf-name>2024 2nd International Conference on Artificial Intelligence, Blockchain, and Internet of Things (AIBThings)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Gondal</surname> <given-names>I</given-names></string-name>, <string-name><surname>Yi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Susnjak</surname> <given-names>T</given-names></string-name>, <string-name><surname>Watters</surname> <given-names>P</given-names></string-name>, <string-name><surname>McIntosh</surname> <given-names>TR</given-names></string-name></person-group>. <article-title>The erosion of cybersecurity zero-trust principles through generative AI: a survey on the challenges and future directions</article-title>. <source>J Cybersecur Priv</source>. <year>2025</year>;<volume>5</volume>(<issue>4</issue>):<fpage>87</fpage>. doi:<pub-id pub-id-type="doi">10.3390/jcp5040087</pub-id>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lv</surname> <given-names>P</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Rag-WM: an efficient black-box watermarking approach for retrieval-augmented generation of large language models</article-title>. In: <conf-name>Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2025</year>. p. <fpage>1709</fpage>&#x2013;<lpage>23</lpage>.</mixed-citation></ref>
<ref id="ref-77"><label>[77]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mudgal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wouhaybi</surname> <given-names>R</given-names></string-name></person-group>. <article-title>An assessment of ChatGPT on log data</article-title>. In: <conf-name>International Conference on AI-generated Content</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>148</fpage>&#x2013;<lpage>69</lpage>.</mixed-citation></ref>
<ref id="ref-78"><label>[78]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Hazra</surname> <given-names>S</given-names></string-name></person-group>. <chapter-title>Review on social and ethical concerns of generative AI and IoT</chapter-title>. In: <source>Generative AI: current trends and applications</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>257</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-97-8460-8_13</pub-id>.</mixed-citation></ref>
<ref id="ref-79"><label>[79]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zeb</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rehman</surname> <given-names>FU</given-names></string-name>, <string-name><surname>Bin Othayman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rabnawaz</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Artificial intelligence and ChatGPT are fostering knowledge sharing, ethics, academia and libraries</article-title>. <source>Int J Inform Learn Technol</source>. <year>2025</year>;<volume>42</volume>(<issue>1</issue>):<fpage>67</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.1108/ijilt-03-2024-0046</pub-id>.</mixed-citation></ref>
<ref id="ref-80"><label>[80]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>YC</given-names></string-name></person-group>. <article-title>Balancing innovation and regulation in the age of generative artificial intelligence</article-title>. <source>J Inform Policy</source>. <year>2024</year>;<volume>14</volume>:<fpage>385</fpage>&#x2013;<lpage>416</lpage>. doi:<pub-id pub-id-type="doi">10.5325/jinfopoli.14.2024.0012</pub-id>.</mixed-citation></ref>
<ref id="ref-81"><label>[81]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Carvalko</surname> <given-names>JR</given-names></string-name></person-group>. <article-title>Generative AI, ingenuity, and law</article-title>. <source>IEEE Trans Technol Soc</source>. <year>2024</year>;<volume>5</volume>(<issue>2</issue>):<fpage>169</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tts.2024.3413591</pub-id>.</mixed-citation></ref>
<ref id="ref-82"><label>[82]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Iyengar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Nabavirazavi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hariprasad</surname> <given-names>Y</given-names></string-name>, <string-name><surname>P.</surname> <given-names>HB</given-names></string-name>, <string-name><surname>Mohan</surname> <given-names>CK</given-names></string-name></person-group>. <chapter-title>The convergence of AI/ML and cybersecurity: advancing digital forensic techniques</chapter-title>. In: <source>Artificial intelligence in practice: theory and application for cyber security and forensics</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>139</fpage>&#x2013;<lpage>59</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-89327-8_4</pub-id>.</mixed-citation></ref>
<ref id="ref-83"><label>[83]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sleiman</surname> <given-names>JP</given-names></string-name></person-group>. <article-title>Generative artificial intelligence and large language models for digital banking: first outlook and perspectives</article-title>. <source>J Digital Bank</source>. <year>2023</year>;<volume>8</volume>(<issue>2</issue>):<fpage>102</fpage>&#x2013;<lpage>17</lpage>. doi:<pub-id pub-id-type="doi">10.69554/cnmi7720</pub-id>.</mixed-citation></ref>
<ref id="ref-84"><label>[84]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lambert</surname> <given-names>J</given-names></string-name>, <string-name><surname>Stevens</surname> <given-names>M</given-names></string-name></person-group>. <article-title>ChatGPT and generative AI technology: a mixed bag of concerns and new opportunities</article-title>. <source>Comput Schools</source>. <year>2024</year>;<volume>41</volume>(<issue>4</issue>):<fpage>559</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.1080/07380569.2023.2256710</pub-id>.</mixed-citation></ref>
<ref id="ref-85"><label>[85]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tihanyi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Bisztray</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ferrag</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cordeiro</surname> <given-names>LC</given-names></string-name></person-group>. <article-title>How secure is AI-generated code: a large-scale comparison of large language models</article-title>. <source>Empir Softw Eng</source>. <year>2025</year>;<volume>30</volume>(<issue>2</issue>):<fpage>47</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10664-024-10590-1</pub-id>.</mixed-citation></ref>
<ref id="ref-86"><label>[86]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>The dual role of large language models in network security: survey and research trends</article-title>. In: <conf-name>Proceedings of the 2025 ACM Workshop on Wireless Security and Machine Learning</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2025</year>. p. <fpage>20</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-87"><label>[87]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Atlam</surname> <given-names>HF</given-names></string-name></person-group>. <article-title>LLMs in cyber security: bridging practice and education</article-title>. <source>Big Data Cogn Comput</source>. <year>2025</year>;<volume>9</volume>(<issue>7</issue>):<fpage>184</fpage>. doi:<pub-id pub-id-type="doi">10.3390/bdcc9070184</pub-id>.</mixed-citation></ref>
<ref id="ref-88"><label>[88]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Celik</surname> <given-names>A</given-names></string-name>, <string-name><surname>Eltawil</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>At the dawn of generative AI era: a tutorial-cum-survey on new frontiers in 6G wireless intelligence</article-title>. <source>IEEE Open J Communicat Soc</source>. <year>2024</year>;<volume>5</volume>:<fpage>2433</fpage>&#x2013;<lpage>89</lpage>. doi:<pub-id pub-id-type="doi">10.1109/OJCOMS.2024.3362271</pub-id>.</mixed-citation></ref>
<ref id="ref-89"><label>[89]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>PY</given-names></string-name>, <string-name><surname>Picek</surname> <given-names>S</given-names></string-name></person-group>. <article-title>An overview of trustworthy AI: advances in IP protection, privacy-preserving federated learning, security verification, and GAI safety alignment</article-title>. <source>IEEE J Emerg Select Topics Circ Syst</source>. <year>2024</year>;<volume>14</volume>(<issue>4</issue>):<fpage>582</fpage>&#x2013;<lpage>607</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jetcas.2024.3477348</pub-id>.</mixed-citation></ref>
<ref id="ref-90"><label>[90]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alwahedi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Aldhaheri</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ferrag</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Battah</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tihanyi</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Machine learning techniques for IoT security: current research and future vision with generative AI and large language models</article-title>. <source>Internet Things Cyber-Phys Syst</source>. <year>2024</year>;<volume>4</volume>:<fpage>167</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iotcps.2023.12.003</pub-id>.</mixed-citation></ref>
<ref id="ref-91"><label>[91]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Arokodare</surname> <given-names>O</given-names></string-name>, <string-name><surname>Wimmer</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Large language models for phishing and spam detection: a BERT approach</article-title>. In: <conf-name>1st Annual Fall National Conference on Creativity, Innovation, and Technology Conference (NCCIT); 2023 Nov 15&#x2013;16</conf-name>; <publisher-loc>Online</publisher-loc>.</mixed-citation></ref>
<ref id="ref-92"><label>[92]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>&#x015E;ent&#x00FC;rk</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Bahtiyar</surname> <given-names>&#x015E;</given-names></string-name></person-group>. <article-title>A survey on large language models in phishing detection</article-title>. In: <conf-name>2025 12th IFIP International Conference on New Technologies, Mobility and Security (NTMS)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>196</fpage>&#x2013;<lpage>204</lpage>.</mixed-citation></ref>
<ref id="ref-93"><label>[93]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>XW</given-names></string-name>, <string-name><surname>See</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kok</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Anticipate, simulate, reason (ASR): a comprehensive generative AI framework for combating messaging scams</article-title>. <comment>arXiv:2507.17543. 2025</comment>.</mixed-citation></ref>
<ref id="ref-94"><label>[94]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Shamoo</surname> <given-names>Y</given-names></string-name></person-group>. <chapter-title>Cybercrime investigation and fraud detection with AI</chapter-title>. In: <source>Digital forensics in the age of AI</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>83</fpage>&#x2013;<lpage>114</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3373-0857-9.ch004</pub-id>.</mixed-citation></ref>
<ref id="ref-95"><label>[95]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Towards federated large language models: motivations, methods, and future directions</article-title>. <source>IEEE Communicat Surv Tutor</source>. <year>2025</year>;<volume>27</volume>(<issue>4</issue>):<fpage>2733</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1109/COMST.2024.3503680</pub-id>.</mixed-citation></ref>
<ref id="ref-96"><label>[96]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Novelo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Silva</surname> <given-names>RR</given-names></string-name>, <string-name><surname>Bernardino</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A literature review of personalized large language models for email generation and automation</article-title>. <source>Future Internet</source>. <year>2025</year>;<volume>17</volume>(<issue>12</issue>):<fpage>536</fpage>. doi:<pub-id pub-id-type="doi">10.3390/fi17120536</pub-id>.</mixed-citation></ref>
<ref id="ref-97"><label>[97]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Reti</surname> <given-names>D</given-names></string-name>, <string-name><surname>Becker</surname> <given-names>N</given-names></string-name>, <string-name><surname>Angeli</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chattopadhyay</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schneider</surname> <given-names>D</given-names></string-name>, <string-name><surname>Vollmer</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Act as a honeytoken generator! an investigation into honeytoken generation with large language models</article-title>. In: <conf-name>Proceedings of the 11th ACM Workshop on Adaptive and Autonomous Cyber Defense</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>12</lpage>.</mixed-citation></ref>
<ref id="ref-98"><label>[98]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pham</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bui</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name></person-group>. <article-title>LLM-based browser extension for phishing detection</article-title>. In: <conf-name>2025 Silicon Valley Cybersecurity Conference (SVCC)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>1</fpage>&#x2013;<lpage>3</lpage>.</mixed-citation></ref>
<ref id="ref-99"><label>[99]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Koide</surname> <given-names>T</given-names></string-name>, <string-name><surname>Fukushi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Nakano</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chiba</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Chatspamdetector: leveraging large language models for effective phishing email detection</article-title>. In: <conf-name>International Conference on Security and Privacy in Communication Systems</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>297</fpage>&#x2013;<lpage>319</lpage>.</mixed-citation></ref>
<ref id="ref-100"><label>[100]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Heiding</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lermen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kao</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schneier</surname> <given-names>B</given-names></string-name>, <string-name><surname>Vishwanath</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Evaluating large language models&#x2019; capability to launch fully automated spear phishing campaigns: validated on human subjects</article-title>. <comment>arXiv:2412.00586. 2024</comment>.</mixed-citation></ref>
<ref id="ref-101"><label>[101]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Weinz</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zannone</surname> <given-names>N</given-names></string-name>, <string-name><surname>Allodi</surname> <given-names>L</given-names></string-name>, <string-name><surname>Apruzzese</surname> <given-names>G</given-names></string-name></person-group>. <article-title>The impact of emerging phishing threats: assessing quishing and LLM-generated phishing emails against organizations</article-title>. In: <conf-name>Proceedings of the 20th ACM Asia Conference on Computer and Communications Security</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2025</year>. p. <fpage>1550</fpage>&#x2013;<lpage>66</lpage>.</mixed-citation></ref>
<ref id="ref-102"><label>[102]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shrestha</surname> <given-names>L</given-names></string-name>, <string-name><surname>Balogun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>AI-driven phishing: techniques, threats, and defence strategies</article-title>. In: <conf-name>International Conference on Global Security, Safety, and Sustainability</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>121</fpage>&#x2013;<lpage>43</lpage>.</mixed-citation></ref>
<ref id="ref-103"><label>[103]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hmimou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tabaa</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khiat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hidila</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A multi-agent system for cybersecurity threat detection and correlation using large language models</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>5</issue>):<fpage>150199</fpage>&#x2013;<lpage>215</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3602681</pub-id>.</mixed-citation></ref>
<ref id="ref-104"><label>[104]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Deshpande</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>S</given-names></string-name></person-group>. <article-title>GenAI in the cyber kill chain: a comprehensive review of risks, threat operative strategies and adaptive defense approaches</article-title>. In: <conf-name>2023 IEEE International Conference on ICT in Business Industry &#x0026; Government (ICTBIG)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-105"><label>[105]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rondanini</surname> <given-names>C</given-names></string-name>, <string-name><surname>Carminati</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ferrari</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kundu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gaudiano</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Malware detection at the edge with lightweight LLMs: a performance evaluation</article-title>. <source>ACM Trans Internet Technol</source>. <year>2026</year>;<volume>26</volume>(<issue>1</issue>):<fpage>15</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3769681</pub-id>.</mixed-citation></ref>
<ref id="ref-106"><label>[106]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Che</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A domain-adaptive large language model with refinement framework for IoT cybersecurity</article-title>. In: <conf-name>2024 IEEE International Conferences on Internet of Things (iThings) and IEEE Green Computing &#x0026; Communications (GreenCom) and IEEE Cyber, Physical &#x0026; Social Computing (CPSCom) and IEEE Smart Data (SmartData) and IEEE Congress on Cybermatics</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>224</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-107"><label>[107]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Devadiga</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>G</given-names></string-name>, <string-name><surname>Potdar</surname> <given-names>B</given-names></string-name>, <string-name><surname>Koo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Han</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shringi</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Gleam: GAN and LLM for evasive adversarial malware</article-title>. In: <conf-name>2023 14th International Conference on Information and Communication Technology Convergence (ICTC)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>53</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-108"><label>[108]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Sikos</surname> <given-names>LF</given-names></string-name></person-group>. <chapter-title>Defensive generative AI</chapter-title>. In: <source>Generative AI in cybersecurity</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>1</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-032-05250-6_1</pub-id>.</mixed-citation></ref>
<ref id="ref-109"><label>[109]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Anglano</surname> <given-names>C</given-names></string-name></person-group>. <article-title>A review of mobile surveillanceware: capabilities, countermeasures, and research challenges</article-title>. <source>Electronics</source>. <year>2025</year>;<volume>14</volume>(<issue>14</issue>):<fpage>2763</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics14142763</pub-id>.</mixed-citation></ref>
<ref id="ref-110"><label>[110]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Al Maqousi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Basu</surname> <given-names>K</given-names></string-name></person-group>. <chapter-title>Cybersecurity in smart cities generative AI as threat and solution</chapter-title>. In: <source>Examining cybersecurity risks produced by generative AI</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>357</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3373-0832-6.ch015</pub-id>.</mixed-citation></ref>
<ref id="ref-111"><label>[111]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Dual intention escape: penetrating and toxic jailbreak attack against large language models</article-title>. In: <conf-name>Proceedings of the ACM on Web Conference 2025</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2025</year>. p. <fpage>863</fpage>&#x2013;<lpage>71</lpage>.</mixed-citation></ref>
<ref id="ref-112"><label>[112]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Farghaly</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Aslan</surname> <given-names>HK</given-names></string-name>, <string-name><surname>Abdel Halim</surname> <given-names>IT</given-names></string-name></person-group>. <article-title>A hybrid human-AI model for enhanced automated vulnerability scoring in modern vehicle sensor systems</article-title>. <source>Future Internet</source>. <year>2025</year>;<volume>17</volume>(<issue>8</issue>):<fpage>339</fpage>. doi:<pub-id pub-id-type="doi">10.3390/fi17080339</pub-id>.</mixed-citation></ref>
<ref id="ref-113"><label>[113]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Katukam</surname> <given-names>R</given-names></string-name></person-group>. <article-title>AI-driven log summarization for security operations centers: a web-based approach using Gemini API</article-title>. <source>Int J Emerg Res Eng Technol</source>. <year>2025</year>;<volume>6</volume>(<issue>3</issue>):<fpage>136</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.63282/3050-922X.IJERET-V6I3P117</pub-id>.</mixed-citation></ref>
<ref id="ref-114"><label>[114]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Harrath</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Adohinzin</surname> <given-names>O</given-names></string-name>, <string-name><surname>Kaabi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Saathoff</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Bridging domains: advances in explainable, automated, and privacy-preserving AI for computer science and cybersecurity</article-title>. <source>Computers</source>. <year>2025</year>;<volume>14</volume>(<issue>9</issue>):<fpage>374</fpage>. doi:<pub-id pub-id-type="doi">10.3390/computers14090374</pub-id>.</mixed-citation></ref>
<ref id="ref-115"><label>[115]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Grov</surname> <given-names>G</given-names></string-name>, <string-name><surname>Halvorsen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Eckhoff</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Hansen</surname> <given-names>BJ</given-names></string-name>, <string-name><surname>Eian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mavroeidis</surname> <given-names>V</given-names></string-name></person-group>. <article-title>On the use of neurosymbolic AI for defending against cyber attacks</article-title>. In: <conf-name>International Conference on Neural-Symbolic Learning and Reasoning</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>119</fpage>&#x2013;<lpage>40</lpage>.</mixed-citation></ref>
<ref id="ref-116"><label>[116]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Maity</surname> <given-names>S</given-names></string-name>, <string-name><surname>Arora</surname> <given-names>J</given-names></string-name></person-group>. <article-title>The colossal defense: security challenges of large language models</article-title>. In: <conf-name>2024 3rd Edition of IEEE Delhi Section Flagship Conference (DELCON)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/delcon64804.2024.10866433</pub-id>.</mixed-citation></ref>
<ref id="ref-117"><label>[117]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pearce</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>B</given-names></string-name>, <string-name><surname>Karri</surname> <given-names>R</given-names></string-name>, <string-name><surname>Dolan-Gavitt</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Examining zero-shot vulnerability repair with large language models</article-title>. In: <conf-name>2023 IEEE Symposium on Security and Privacy (SP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>2339</fpage>&#x2013;<lpage>56</lpage>.</mixed-citation></ref>
<ref id="ref-118"><label>[118]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>AlOmar</surname> <given-names>B</given-names></string-name>, <string-name><surname>Trabelsi</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Integrating generative AI in cybersecurity curricula</article-title>. In: <conf-name>2025 IEEE Global Engineering Education Conference (EDUCON)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-119"><label>[119]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>T</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Joint optimization of prompt security and system performance in edge-cloud LLM systems</article-title>. In: <conf-name>IEEE INFOCOM 2025-IEEE Conference on Computer Communications</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>1</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-120"><label>[120]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Moorthy</surname> <given-names>V</given-names></string-name>, <collab>Rupen</collab>, <string-name><surname>Chopra</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Uncovering hidden risks: a vulnerability analysis of advanced conversational AI models</article-title>. In: <conf-name>International Conference on Smart Trends for Information Technology and Computer Communications</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>377</fpage>&#x2013;<lpage>88</lpage>.</mixed-citation></ref>
<ref id="ref-121"><label>[121]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yashvardhan</surname> <given-names>U</given-names></string-name>, <string-name><surname>Chamola</surname> <given-names>V</given-names></string-name>, <string-name><surname>Sikdar</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Generative AI for cyber security: analyzing the potential of ChatGPT, DALL-E, and other models for enhancing the security space</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>53497</fpage>&#x2013;<lpage>516</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3385107</pub-id>.</mixed-citation></ref>
<ref id="ref-122"><label>[122]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kuntur</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wr&#x00F3;blewska</surname> <given-names>A</given-names></string-name>, <string-name><surname>Paprzycki</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ganzha</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Under the influence: a survey of large language models in fake news detection</article-title>. <source>IEEE Trans Artif Intell</source>. <year>2025</year>;<volume>6</volume>(<issue>2</issue>):<fpage>458</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TAI.2024.3471735</pub-id>.</mixed-citation></ref>
<ref id="ref-123"><label>[123]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Harris</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hadi</surname> <given-names>HJ</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>N</given-names></string-name>, <string-name><surname>Alshara</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Fake news detection revisited: an extensive review of theoretical frameworks, dataset assessments, model constraints, and forward-looking research agendas</article-title>. <source>Technologies</source>. <year>2024</year>;<volume>12</volume>(<issue>11</issue>):<fpage>222</fpage>. doi:<pub-id pub-id-type="doi">10.3390/technologies12110222</pub-id>.</mixed-citation></ref>
<ref id="ref-124"><label>[124]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>N</given-names></string-name>, <string-name><surname>Tani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jha</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Can LLM-generated misinformation be detected: a study on Cyber Threat Intelligence</article-title>. <source>Future Gener Comput Syst</source>. <year>2025</year>;<volume>173</volume>:<fpage>107877</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2025.107877</pub-id>.</mixed-citation></ref>
<ref id="ref-125"><label>[125]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Han</surname> <given-names>QL</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>The security of using large language models: a survey with emphasis on ChatGPT</article-title>. <source>IEEE/CAA J Automat Sinica</source>. <year>2025</year>;<volume>12</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JAS.2024.124983</pub-id>.</mixed-citation></ref>
<ref id="ref-126"><label>[126]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Golec</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hachaj</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Ten natural language processing tasks with generative artificial intelligence</article-title>. <source>Appl Sci</source>. <year>2025</year>;<volume>15</volume>(<issue>16</issue>):<fpage>9057</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app15169057</pub-id>.</mixed-citation></ref>
<ref id="ref-127"><label>[127]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Okey</surname> <given-names>OD</given-names></string-name>, <string-name><surname>Udo</surname> <given-names>EU</given-names></string-name>, <string-name><surname>Rosa</surname> <given-names>RL</given-names></string-name>, <string-name><surname>Rodr&#x00ED;guez</surname> <given-names>DZ</given-names></string-name>, <string-name><surname>Kleinschmidt</surname> <given-names>JH</given-names></string-name></person-group>. <article-title>Investigating ChatGPT and cybersecurity: a perspective on topic modeling and sentiment analysis</article-title>. <source>Comput Secur</source>. <year>2023</year>;<volume>135</volume>:<fpage>103476</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2023.103476</pub-id>.</mixed-citation></ref>
<ref id="ref-128"><label>[128]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>N</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>K</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Large language models for cyber security: a systematic literature review</article-title>. <source>ACM Trans Softw Eng Methodol</source>. <year>2024</year>. doi:<pub-id pub-id-type="doi">10.1145/3769676</pub-id>.</mixed-citation></ref>
<ref id="ref-129"><label>[129]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Song</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>AI-driven safety and security for UAVs: from machine learning to large language models</article-title>. <source>Drones</source>. <year>2025</year>;<volume>9</volume>(<issue>6</issue>):<fpage>392</fpage>. doi:<pub-id pub-id-type="doi">10.3390/drones9060392</pub-id>.</mixed-citation></ref>
<ref id="ref-130"><label>[130]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Miao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Automatic generation of cybersecurity teaching cases using large language models</article-title>. <source>Comput Appl Eng Educ</source>. <year>2025</year>;<volume>33</volume>(<issue>5</issue>):<fpage>e70081</fpage>. doi:<pub-id pub-id-type="doi">10.1002/cae.70081</pub-id>.</mixed-citation></ref>
<ref id="ref-131"><label>[131]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ibrahim</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kashef</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Exploring the emerging role of large language models in smart grid cybersecurity: a survey of attacks, detection mechanisms, and mitigation strategies</article-title>. <source>Front Energy Res</source>. <year>2025</year>;<volume>13</volume>:<fpage>1531655</fpage>. doi:<pub-id pub-id-type="doi">10.3389/fenrg.2025.1531655</pub-id>.</mixed-citation></ref>
<ref id="ref-132"><label>[132]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Sarker</surname> <given-names>IH</given-names></string-name></person-group>. <chapter-title>Generative AI and large language modeling in cybersecurity</chapter-title>. In: <source>AI-driven cybersecurity and threat intelligence: cyber automation, intelligent decision-making and explainability</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>79</fpage>&#x2013;<lpage>99</lpage>.</mixed-citation></ref>
<ref id="ref-133"><label>[133]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Che</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Text-enhanced method for LLMs domain adaptation in cybersecurity</article-title>. <source>J Comput Inf Syst</source>. <year>2025</year>. doi:<pub-id pub-id-type="doi">10.1080/08874417.2025.2554853</pub-id>.</mixed-citation></ref>
<ref id="ref-134"><label>[134]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>C</given-names></string-name></person-group>. <article-title>How to protect copyright data in optimization of large language models?</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence</conf-name>. Vol. 38; <publisher-loc>Palo Alto, CA, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>; <year>2024</year>. p. <fpage>17871</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-135"><label>[135]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Kaushik</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kaushik</surname> <given-names>A</given-names></string-name></person-group>. <chapter-title>Decoding potential of ChatGPT: a comprehensive exploration of AI generated contents and challenges</chapter-title>. In: <source>Textual intelligence: large language models and their real-world applications</source>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons, Inc.</publisher-name>; <year>2025</year>. p. <fpage>177</fpage>&#x2013;<lpage>99</lpage>.</mixed-citation></ref>
<ref id="ref-136"><label>[136]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Usman</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Oladipupo</surname> <given-names>H</given-names></string-name>, <string-name><surname>During</surname> <given-names>AD</given-names></string-name>, <string-name><surname>Robert</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chataut</surname> <given-names>R</given-names></string-name></person-group>. <article-title>AI, ML, and LLM integration in 5G/6G networks: a comprehensive survey of architectures, challenges, and future directions</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>168914</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-137"><label>[137]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>TT</given-names></string-name>, <string-name><surname>Vu</surname> <given-names>HT</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>HN</given-names></string-name></person-group>. <article-title>Security, privacy, and ethical challenges of artificial intelligence in large language model scope: a comprehensive survey</article-title>. In: <conf-name>2024 1st International Conference on Cryptography and Information Security (VCRIS)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-138"><label>[138]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sezgin</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Scenario-driven evaluation of autonomous agents: integrating large language model for UAV mission reliability</article-title>. <source>Drones</source>. <year>2025</year>;<volume>9</volume>(<issue>3</issue>):<fpage>213</fpage>. doi:<pub-id pub-id-type="doi">10.3390/drones9030213</pub-id>.</mixed-citation></ref>
<ref id="ref-139"><label>[139]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Dave</surname> <given-names>D</given-names></string-name>, <string-name><surname>Cody</surname> <given-names>T</given-names></string-name>, <string-name><surname>Beling</surname> <given-names>PA</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>M</given-names></string-name></person-group>. <article-title>From capabilities to performance: evaluating key functional properties of LLM architectures in penetration testing</article-title>. In: <conf-name>Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; 2025 Nov 4&#x2013;9</conf-name>; <publisher-loc>Suzhou, China</publisher-loc>. p. <fpage>15890</fpage>&#x2013;<lpage>916</lpage>.</mixed-citation></ref>
<ref id="ref-140"><label>[140]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Narula</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ghasemigol</surname> <given-names>M</given-names></string-name>, <string-name><surname>Carnerero-Cano</surname> <given-names>J</given-names></string-name>, <string-name><surname>Minnich</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lupu</surname> <given-names>E</given-names></string-name>, <string-name><surname>Takabi</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Exploring research and tools in AI security: a systematic mapping study</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>2</issue>):<fpage>84057</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3567195</pub-id>.</mixed-citation></ref>
<ref id="ref-141"><label>[141]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Moenks</surname> <given-names>N</given-names></string-name>, <string-name><surname>Penava</surname> <given-names>P</given-names></string-name>, <string-name><surname>Buettner</surname> <given-names>R</given-names></string-name></person-group>. <article-title>A systematic literature review of large language model applications in industry</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>4</issue>):<fpage>160010</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3608650</pub-id>.</mixed-citation></ref>
<ref id="ref-142"><label>[142]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mohawesh</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ottom</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Salameh</surname> <given-names>HB</given-names></string-name></person-group>. <article-title>A data-driven risk assessment of cybersecurity challenges posed by generative AI</article-title>. <source>Decis Anal J</source>. <year>2025</year>;<volume>15</volume>(<issue>6</issue>):<fpage>100580</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dajour.2025.100580</pub-id>.</mixed-citation></ref>
<ref id="ref-143"><label>[143]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ihekweazu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>B</given-names></string-name>, <string-name><surname>Adelowo</surname> <given-names>EA</given-names></string-name></person-group>. <article-title>Ethics-driven education: integrating AI responsibly for academic excellence</article-title>. <source>Inform Syst Educat J</source>. <year>2024</year>;<volume>22</volume>(<issue>3</issue>):<fpage>36</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.62273/JWXX9525</pub-id>.</mixed-citation></ref>
<ref id="ref-144"><label>[144]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>MU</given-names></string-name>, <string-name><surname>Ullah</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>SU</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Ethical principles of integrating ChatGPT into IoT-based software wearables: a fuzzy-TOPSIS ranking and analysis approach</article-title>. <source>Int J Intell Syst</source>. <year>2025</year>;<volume>2025</volume>(<issue>1</issue>):<fpage>6660868</fpage>. doi:<pub-id pub-id-type="doi">10.1155/int/6660868</pub-id>.</mixed-citation></ref>
<ref id="ref-145"><label>[145]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Czaja</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gdowski</surname> <given-names>B</given-names></string-name>, <string-name><surname>Niemiec</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mees</surname> <given-names>W</given-names></string-name>, <string-name><surname>Stoianov</surname> <given-names>N</given-names></string-name>, <string-name><surname>Votis</surname> <given-names>K</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Cybersecurity challenges and opportunities of machine learning-based artificial intelligence</article-title>. <source>Neural Comput Appl</source>. <year>2025</year>;<volume>37</volume>:<fpage>27931</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-025-11604-9</pub-id>.</mixed-citation></ref>
<ref id="ref-146"><label>[146]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xiong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Du</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>When search engine services meet large language models: visions and challenges</article-title>. <source>IEEE Trans Serv Comput</source>. <year>2024</year>;<volume>17</volume>(<issue>6</issue>):<fpage>4558</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tsc.2024.3451185</pub-id>.</mixed-citation></ref>
<ref id="ref-147"><label>[147]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>N</given-names></string-name>, <string-name><surname>Miao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Large language models for cybersecurity education: a survey of current practices and future directions</article-title>. In: <conf-name>Pacific-Asia Conference on Knowledge Discovery and Data Mining</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>3</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-148"><label>[148]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mateo Sanguino</surname> <given-names>TDJ</given-names></string-name></person-group>. <article-title>Enhancing security in industrial application development: case study on self-generating artificial intelligence</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>9</issue>):<fpage>3780</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app14093780</pub-id>.</mixed-citation></ref>
<ref id="ref-149"><label>[149]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Doumanas</surname> <given-names>D</given-names></string-name>, <string-name><surname>Karakikes</surname> <given-names>A</given-names></string-name>, <string-name><surname>Soularidis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mainas</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kotis</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Emerging threat vectors: how malicious actors exploit LLMs to undermine border security</article-title>. <source>AI</source>. <year>2025</year>;<volume>6</volume>(<issue>9</issue>):<fpage>232</fpage>.</mixed-citation></ref>
<ref id="ref-150"><label>[150]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Truong</surname> <given-names>TC</given-names></string-name>, <string-name><surname>Diep</surname> <given-names>QB</given-names></string-name>, <string-name><surname>Zelinka</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Artificial intelligence in the cyber domain: offense and defense</article-title>. <source>Symmetry</source>. <year>2020</year>;<volume>12</volume>(<issue>3</issue>):<fpage>410</fpage>. doi:<pub-id pub-id-type="doi">10.3390/sym12030410</pub-id>.</mixed-citation></ref>
<ref id="ref-151"><label>[151]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Divakaran</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Peddinti</surname> <given-names>ST</given-names></string-name></person-group>. <article-title>Large language models for cybersecurity: new opportunities</article-title>. <source>IEEE Secur Priv</source>. <year>2024</year>;<volume>23</volume>:<fpage>38</fpage>&#x2013;<lpage>45</lpage>.</mixed-citation></ref>
<ref id="ref-152"><label>[152]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kasri</surname> <given-names>W</given-names></string-name>, <string-name><surname>Himeur</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Alkhazaleh</surname> <given-names>HA</given-names></string-name>, <string-name><surname>Tarapiah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Atalla</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mansoor</surname> <given-names>W</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>From vulnerability to defense: the role of large language models in enhancing cybersecurity</article-title>. <source>Computation</source>. <year>2025</year>;<volume>13</volume>(<issue>2</issue>):<fpage>30</fpage>. doi:<pub-id pub-id-type="doi">10.3390/computation13020030</pub-id>.</mixed-citation></ref>
<ref id="ref-153"><label>[153]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Friha</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ferrag</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Kantarci</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cakmak</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ozgun</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ghoualmi-Zine</surname> <given-names>N</given-names></string-name></person-group>. <article-title>LLM-based edge intelligence: a comprehensive survey on architectures, applications, security and trustworthiness</article-title>. <source>IEEE Open J Communicat Soc</source>. <year>2024</year>;<volume>5</volume>:<fpage>5799</fpage>&#x2013;<lpage>856</lpage>. doi:<pub-id pub-id-type="doi">10.1109/OJCOMS.2024.3456549</pub-id>.</mixed-citation></ref>
<ref id="ref-154"><label>[154]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Javaid</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fahim</surname> <given-names>H</given-names></string-name>, <string-name><surname>He</surname> <given-names>B</given-names></string-name>, <string-name><surname>Saeed</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Large language models for UAVS: current state and pathways to the future</article-title>. <source>IEEE Open J Vehic Technol</source>. <year>2024</year>;<volume>5</volume>:<fpage>1166</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1109/OJVT.2024.3446799</pub-id>.</mixed-citation></ref>
<ref id="ref-155"><label>[155]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Goertzel</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wright</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ponnapalli</surname> <given-names>J</given-names></string-name></person-group>. <source>Generative AI security: theories and practices (future of business and finance)</source>. <publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>.</mixed-citation></ref>
<ref id="ref-156"><label>[156]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>Ismail</collab>, <string-name><surname>Kurnia</surname> <given-names>R</given-names></string-name>, <string-name><surname>Widyatama</surname> <given-names>F</given-names></string-name>, <string-name><surname>Wibawa</surname> <given-names>IM</given-names></string-name>, <string-name><surname>Brata</surname> <given-names>ZA</given-names></string-name>, <collab>Ukasyah</collab>, <etal>et al</etal></person-group>. <article-title>Enhancing security operations center: wazuh security event response with retrieval-augmented-generation-driven copilot</article-title>. <source>Sensors</source>. <year>2025</year>;<volume>25</volume>(<issue>3</issue>):<fpage>870</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s25030870</pub-id>; <pub-id pub-id-type="pmid">39943508</pub-id></mixed-citation></ref>
<ref id="ref-157"><label>[157]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Uddin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Irshad</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Kandhro</surname> <given-names>IA</given-names></string-name>, <string-name><surname>Alanazi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>F</given-names></string-name>, <string-name><surname>Maaz</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Generative AI revolution in cybersecurity: a comprehensive review of threat intelligence and operations</article-title>. <source>Artif Intell Rev</source>. <year>2025</year>;<volume>58</volume>(<issue>8</issue>):<fpage>236</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-025-11219-5</pub-id>.</mixed-citation></ref>
<ref id="ref-158"><label>[158]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Karkuzhali</surname> <given-names>S</given-names></string-name>, <string-name><surname>Senthilkumar</surname> <given-names>S</given-names></string-name></person-group>. <chapter-title>LLM-powered security solutions in healthcare, government, and industrial cybersecurity</chapter-title>. In: <source>Revolutionizing cybersecurity with deep learning and large language models</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>97</fpage>&#x2013;<lpage>132</lpage>.</mixed-citation></ref>
<ref id="ref-159"><label>[159]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Haque</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Siddique</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Hasan</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Das</surname> <given-names>LR</given-names></string-name>, <string-name><surname>Kamal</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SOK: exploring hallucinations and security risks in AI-assisted software development with insights for LLM deployment</article-title>. In: <conf-name>2025 Sixth International Conference on Intelligent Data Science Technologies and Applications (IDSTA)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>57</fpage>&#x2013;<lpage>64</lpage>.</mixed-citation></ref>
<ref id="ref-160"><label>[160]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rondanini</surname> <given-names>C</given-names></string-name>, <string-name><surname>Carminati</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ferrari</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kundu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jajoo</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Large language models to enhance malware detection in edge computing</article-title>. In: <conf-name>2024 IEEE 6th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS-ISA)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-161"><label>[161]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zou</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Singhal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Deep learning for detecting logic-flaw-exploiting network attacks: an end-to-end approach</article-title>. <source>J Comput Secur</source>. <year>2022</year>;<volume>30</volume>(<issue>4</issue>):<fpage>541</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.3233/jcs-210101</pub-id>.</mixed-citation></ref>
<ref id="ref-162"><label>[162]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lai</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Intrusion detection technology based on large language models</article-title>. In: <conf-name>2023 International Conference on Evolutionary Algorithms and Soft Computing Techniques (EASCT)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-163"><label>[163]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Usman</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ihejirika</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Offor</surname> <given-names>SN</given-names></string-name>, <string-name><surname>Robert</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chataut</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Green cybersecurity: leveraging AI, ML, and LLMs to optimize energy, threat detection, and sustainability Frameworks</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>1</issue>):<fpage>159345</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3602451</pub-id>.</mixed-citation></ref>
<ref id="ref-164"><label>[164]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kulkarni</surname> <given-names>A</given-names></string-name>, <string-name><surname>Balachandran</surname> <given-names>V</given-names></string-name>, <string-name><surname>Divakaran</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Das</surname> <given-names>T</given-names></string-name></person-group>. <article-title>From ML to LLM: evaluating the robustness of phishing web page detection models against adversarial attacks</article-title>. <source>Digital Threats Res Pract</source>. <year>2025</year>;<volume>6</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3737295</pub-id>.</mixed-citation></ref>
<ref id="ref-165"><label>[165]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Toth</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bisztray</surname> <given-names>T</given-names></string-name>, <string-name><surname>Erdodi</surname> <given-names>L</given-names></string-name></person-group>. <article-title>LLMS in web development: evaluating LLM-generated PHP code unveiling vulnerabilities and limitations</article-title>. In: <conf-name>International Conference on Computer Safety, Reliability, and Security</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>425</fpage>&#x2013;<lpage>37</lpage>.</mixed-citation></ref>
<ref id="ref-166"><label>[166]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Greshake</surname> <given-names>K</given-names></string-name>, <string-name><surname>Abdelnabi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mishra</surname> <given-names>S</given-names></string-name>, <string-name><surname>Endres</surname> <given-names>C</given-names></string-name>, <string-name><surname>Holz</surname> <given-names>T</given-names></string-name>, <string-name><surname>Fritz</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Not what you&#x2019;ve signed up for: compromising real-world LLM-integrated applications with indirect prompt injection</article-title>. In: <conf-name>Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2023</year>. p. <fpage>79</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-167"><label>[167]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Dharmendra</surname> <given-names>H</given-names></string-name>, <string-name><surname>Raghunandan</surname> <given-names>G</given-names></string-name>, <string-name><surname>Sindhu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Samanvitha</surname> <given-names>C</given-names></string-name>, <string-name><surname>Nethravathi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Elango</surname> <given-names>D</given-names></string-name></person-group>. <chapter-title>Human evaluation in large language model testing: assessing the quality of AI model output</chapter-title>. In: <source>Advancements in intelligent process automation</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global</publisher-name>; <year>2025</year>. p. <fpage>553</fpage>&#x2013;<lpage>74</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3693-5380-6.ch022</pub-id>.</mixed-citation></ref>
<ref id="ref-168"><label>[168]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zaydi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Maleh</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Empowering red teams with generative AI: transforming penetration testing through adaptive intelligence</article-title>. <source>EDPACS</source>. <year>2025</year>;<volume>70</volume>(<issue>2</issue>):<fpage>41</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.1080/07366981.2024.2439628</pub-id>.</mixed-citation></ref>
<ref id="ref-169"><label>[169]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Omar</surname> <given-names>KO</given-names></string-name>, <string-name><surname>Zraqou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>JM</given-names></string-name></person-group>. <chapter-title>From synthetic text to real threats: unraveling the security risks of generative AI</chapter-title>. In: <source>Examining cybersecurity risks produced by generative AI</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>1</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3373-0832-6.ch001</pub-id>.</mixed-citation></ref>
<ref id="ref-170"><label>[170]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Manickam</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chong</surname> <given-names>YW</given-names></string-name>, <string-name><surname>Leng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Nanda</surname> <given-names>P</given-names></string-name></person-group>. <article-title>A state-of-the-art review on phishing website detection techniques</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>4</issue>):<fpage>187976</fpage>&#x2013;<lpage>8012</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3514972</pub-id>.</mixed-citation></ref>
<ref id="ref-171"><label>[171]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Alharthi</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yasaei</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Llm-powered automated cloud forensics: from log analysis to investigation</article-title>. In: <conf-name>2025 IEEE 18th International Conference on Cloud Computing (CLOUD)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>12</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>