<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">74566</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.074566</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Integration of Large Language Models (LLMs) and Static Analysis for Improving the Efficacy of Security Vulnerability Detection in Source Code</article-title>
<alt-title alt-title-type="left-running-head">Integration of Large Language Models (LLMs) and Static Analysis for Improving the Efficacy of Security Vulnerability Detection in Source Code</alt-title>
<alt-title alt-title-type="right-running-head">Integration of Large Language Models (LLMs) and Static Analysis for Improving the Efficacy of Security Vulnerability Detection in Source Code</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Santas Ciavatta</surname><given-names>Jos&#x00E9; Armando</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Bermejo Higuera</surname><given-names>Juan Ram&#x00F3;n</given-names></name><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>juanramon.bermejo@unir.net</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Bermejo Higuera</surname><given-names>Javier</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Montalvo</surname><given-names>Juan Antonio Sicilia</given-names></name></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Riera</surname><given-names>Tom&#x00E1;s Sureda</given-names></name></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>P&#x00E9;rez Melero</surname><given-names>Jes&#x00FA;s</given-names></name></contrib>
<aff id="aff-1">
<institution>School of Engineering and Technology, International University of La Rioja, Avda.de La Paz, 137</institution>, <addr-line>Logro&#x00F1;o, 26006, La Rioja</addr-line>, <country>Spain</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Juan Ram&#x00F3;n Bermejo Higuera. Email: <email>juanramon.bermejo@unir.net</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>12</day><month>1</month><year>2026</year>
</pub-date>
<volume>86</volume>
<issue>3</issue>
<elocation-id>11</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>11</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_74566.pdf"></self-uri>
<abstract>
<p>As artificial Intelligence (AI) continues to expand exponentially, particularly with the emergence of generative pre-trained transformers (GPT) based on a transformer&#x2019;s architecture, which has revolutionized data processing and enabled significant improvements in various applications. This document seeks to investigate the security vulnerabilities detection in the source code using a range of large language models (LLM). Our primary objective is to evaluate the effectiveness of Static Application Security Testing (SAST) by applying various techniques such as prompt persona, structure outputs and zero-shot. To the selection of the LLMs (CodeLlama 7B, DeepSeek coder 7B, Gemini 1.5 Flash, Gemini 2.0 Flash, Mistral 7b Instruct, Phi 3 8b Mini 128K instruct, Qwen 2.5 coder, StartCoder 2 7B) with comparison and combination with Find Security Bugs. The evaluation method will involve using a selected dataset containing vulnerabilities, and the results to provide insights for different scenarios according to the software criticality (Business critical, non-critical, minimum effort, best effort) In detail, the main objectives of this study are to investigate if large language models outperform or exceed the capabilities of traditional static analysis tools, if the combining LLMs with Static Application Security Testing (SAST) tools lead to an improvement and the possibility that local machine learning models on a normal computer produce reliable results. Summarizing the most important conclusions of the research, it can be said that while it is true that the results have improved depending on the size of the LLM for business-critical software, the best results have been obtained by SAST analysis. This differs in &#x201C;Non-Critical,&#x201D; &#x201C;Best Effort,&#x201D; and &#x201C;Minimum Effort&#x201D; scenarios, where the combination of LLM (Gemini) &#x002B; SAST has obtained better results.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>AI &#x002B; SAST</kwd>
<kwd>secure code</kwd>
<kwd>LLM</kwd>
<kwd>benchmarking LLM</kwd>
<kwd>vulnerability detection</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Software security vulnerabilities are rising, with open-source software reporting an 89% annual increase [<xref ref-type="bibr" rid="ref-1">1</xref>]; the most common issues are Cross-site Scripting (XSS), SQL Injection, and Memory Corruption [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>]. This trend, combined with organizations&#x2019; efforts to raise awareness about implementing adequate system defenses [<xref ref-type="bibr" rid="ref-4">4</xref>], contributes to the growing economic impact of breaches, with the average cost reaching USD 4.88 million&#x2014;a 10% increase over the previous year&#x2014;and global IT investment expected to grow 9.8% in 2025 [<xref ref-type="bibr" rid="ref-5">5</xref>], reflecting the expanding scope of vulnerabilities affecting all sectors worldwide [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Artificial intelligence technologies, particularly Generative Pretrained Transformers (GPT), have demonstrated human-like abilities in text generation, comprehension, and recognition [<xref ref-type="bibr" rid="ref-8">8</xref>], and are being exploited in cybercrime through phishing, DeepFakes, and malicious GPT attacks [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>], while also enabling AI-assisted vulnerability detection [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>The increasing complexity of technologies and methodologies together with the AI revolution, remarks the need for a secure and speed the process of Software Development Life Cycle (S-SDLC) to address the high demand, reducing the overall cost of vulnerability management [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>] and the number of manual activities without reducing the quantity of controls. Lack of controls can lead to the leakage of sensitive information, privilege escalation, and a detrimental impact on system performance and efficiency, affecting both corporate entities and individual users [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Implementing a robust SAST solution requires not only cybersecurity expertise but also a deep understanding of the entire technology stack, increasing effort and costs as systems evolve [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. The absence of a proper SSDLC contributes to the growing frequency and impact of cyberattacks, highlighting the need to integrate secure development practices from the earliest stages to mitigate risks and enhance resilience against emerging threats [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
<p>Around 68% of the world&#x2019;s population&#x2014;about 5.5 billion people&#x2014;are now online [<xref ref-type="bibr" rid="ref-23">23</xref>]. As AI adoption can lower overall cybersecurity costs [<xref ref-type="bibr" rid="ref-24">24</xref>], this expanding digital landscape underscores the need for secure and scalable technologies. However, research on integrating artificial intelligence into Static Application Security Testing (SAST) remains limited. Existing studies often focus on specific vulnerability categories or small datasets and explore only a narrow range of model architectures. This study aims to evaluate the benefits and limitations of integrating or alone the AI into static code security analysis, examining how it can improve defect detection, vulnerability identification, and overall efficiency compared to traditional tools, considering factors such as accuracy, analysis time, and impact on development workflows.</p>
<p>Specifically:
<list list-type="simple">
<list-item><label>1.</label><p>Evaluate different language models (LLM) solutions that are specifically trained for code analysis and can be executed locally.</p></list-item>
<list-item><label>2.</label><p>Apply a range of prompt engineering techniques to determine whether these techniques improve the results compared with widely used open-source static analysis tools (SAST).
<list list-type="simple">
<list-item><label>a.</label><p>Tools SAST.</p></list-item>
<list-item><label>b.</label><p>LLM (minimum adjustments).</p></list-item>
<list-item><label>c.</label><p>LLM with prompt engineering.</p></list-item>
</list></p></list-item>
</list></p>
<p>The experiment will aim to answer the following key research questions or hypotheses:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>Q1. Do the results produced by LLMs outperform those of traditional static analysis tools?</p></list-item>
<list-item><label>&#x2022;</label><p>Q2. Is there an improvement when combining LLMs with SAST tools?</p></list-item>
<list-item><label>&#x2022;</label><p>Q3. Can reliable results be achieved with locally deployed models on a standard desktop computer?</p></list-item>
</list></p>
<p>Data privacy remains a central concern in adopting AI-based technologies. Most experiments were conducted on standard laptops or small servers, except for Gemini. This approach can benefit governments and corporations, as running analyses on local machines reduces dependency on large infrastructures and enables early detection of potential vulnerabilities during implementation.</p>
<p>The main contributions of this work are:
<list list-type="simple">
<list-item><label>1.</label><p>Compare each local LLM, a cloud based LLM service (Gemini), and the top performing SAST tool (FindSecBugs) to address Question 1.</p></list-item>
<list-item><label>2.</label><p>Evaluate the impact of combining local LLMs with cloud services on detection quality, thereby tackling Question 2.</p></list-item>
<list-item><label>3.</label><p>Benchmark a broad set of locally deployable LLMs (CodeLlama 7B, DeepSeek Cod-er 7B, Gemini 1.5 Flash, Gemini 2.0 Flash, Mistral 7B Instruct, Phi 3 8B Mini 128K Instruct, Qwen 2.5 Coder, StarCoder 2 7B) to answer Question 3.</p></list-item>
<list-item><label>4.</label><p>Restrict the study to a single, widely used programming language (Java), eliminating confounding variables introduced by heterogeneous technology stacks.</p></list-item>
<list-item><label>5.</label><p>Use a standardized test set: the OWASP benchmark project, which contains both vulnerable and non-vulnerable Java programs. To improve the OWASP benchmark functionality, it has been developed an adaptation that enables LLMs to participate in the analysis workflow, thereby extending the tool&#x2019;s functionality and introducing new automated evaluation modes over Benchmark&#x2019;s test suite.</p></list-item>
<list-item><label>6.</label><p>Assess the results across multiple operational contexts&#x2014;business critical, non-critical, best effort, and minimum effort environments [<xref ref-type="bibr" rid="ref-25">25</xref>].</p></list-item>
</list></p>
<p>The rest of the work is structured as follows. <xref ref-type="sec" rid="s2">Section 2</xref> describes the cybersecurity context, specifically the Secure Software Development Life Cycle (S-SDLC), and the AI landscape, including the large language models (LLMs) used in the experiment and their salient characteristics. Also, it provides the relative work of this experiment. <xref ref-type="sec" rid="s3">Section 3</xref> details the phases and the activities used to configure, run and collect the experiment results, as well as the preprocessing steps required before the analysis. <xref ref-type="sec" rid="s4">Section 4</xref> shows the results analysis of the experiment interpreting the data and answering the research questions. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> offers the conclusions, highlights the main contributions and the further path of proposed research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Background and Relative Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Static Analysis and AI</title>
<p>Static Application Security Testing (SAST) analyzes source code without execution to detect vulnerabilities early in the SDLC [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>]. Its core, taint analysis, tracks untrusted data through program paths to identify flaws such as SQL injection and XSS [<xref ref-type="bibr" rid="ref-29">29</xref>&#x2013;<xref ref-type="bibr" rid="ref-31">31</xref>]. While SAST is cost-effective and fast [<xref ref-type="bibr" rid="ref-32">32</xref>], its real-world detection capability is limited, often producing many false positives and highlighting the need to improve recall [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>SAST tools, while limited in detecting some vulnerabilities, are crucial for analyzing all code paths [<xref ref-type="bibr" rid="ref-28">28</xref>]. Their main drawback is high false-positive rates [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>], which can cause developer fatigue and reduce confidence in the tool [<xref ref-type="bibr" rid="ref-19">19</xref>]. Research addresses this by enhancing precision and usability through machine learning, enabling better vulnerability inference, false-positive prediction, and scalable analyzers for industrial applications [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>].</p>
<p>Large Language Models (LLMs) are reshaping Static Application Security Testing (SAST) by addressing traditional limitations such as low recall and lack of repository-level context [<xref ref-type="bibr" rid="ref-35">35</xref>]. Models like GPT-4 achieve higher F1 scores in vulnerability detection due to their ability to reason across broader code contexts and understand complex semantics [<xref ref-type="bibr" rid="ref-36">36</xref>]. Hybrid approaches, such as LSAST, combine conventional data-flow analysis (e.g., CodeQL) with LLM-driven taint inference and contextual reasoning, previously achievable only manually [<xref ref-type="bibr" rid="ref-37">37</xref>]. However, LLMs often produce high false-positive rates, sometimes exceeding 60% [<xref ref-type="bibr" rid="ref-38">38</xref>]. Nevertheless, recent studies challenge this assumption, showing that high false-positive rates stem not from intrinsic LLM limitations but from inadequate or noisy code context. When equipped with precise and complete contextual information&#x2014;such as through frameworks like LLM4FPM&#x2014;LLMs can achieve F1 scores above 99% on benchmark datasets (e.g., Juliet) and reduce false positives by over 85% in real-world projects [<xref ref-type="bibr" rid="ref-39">39</xref>]. Research now leverages LLMs for filtering and alert adjudication, with prototypes like FPShield significantly reducing false positives and developer effort [<xref ref-type="bibr" rid="ref-36">36</xref>]. The emerging trend integrates LLMs with traditional SAST, using analyzers for data flow and LLMs for semantic validation, prioritization, and correction suggestions [<xref ref-type="bibr" rid="ref-37">37</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Datasets</title>
<p>The Juliet test suite [<xref ref-type="bibr" rid="ref-40">40</xref>] is a widely used NSA-developed benchmark for evaluating static code analysis tools in C/C&#x002B;&#x002B; and Java. It consists of thousands of synthetic test cases covering a broad range of CWE-classified vulnerabilities, with both vulnerable and fixed versions to facilitate tool validation [<xref ref-type="bibr" rid="ref-34">34</xref>]. Similarly, the OWASP Benchmark dataset [<xref ref-type="bibr" rid="ref-41">41</xref>], which we employ in this work, contains 2740 synthetic Java test cases focusing on web application vulnerabilities such as SQL Injection, XSS, Command Injection, LDAP Injection, Path Traversal, Weak Encryption/Hashing, Insecure Cookies, Trust Boundary, Weak Randomness, and XPath Injection, with a balanced mix of true and false positives. Both datasets use synthetic code to provide controlled and reproducible benchmarks for evaluating security tools.</p>
<p>Below, <xref ref-type="table" rid="table-1">Table 1</xref> shows the categories of vulnerabilities and percentage based on the total.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Dataset type of vulnerabilities</title>
</caption>
<table>
<colgroup>
<col align="center" width="38mm"/>
<col align="center" width="30mm"/>
<col align="center" width="45mm"/> </colgroup>
<thead>
<tr>
<th>Category vulnerability</th>
<th>Vulnerabilities</th>
<th>Type of vulnerability (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CMDI</td>
<td>251</td>
<td>9.16%</td>
</tr>
<tr>
<td>CRYPTO</td>
<td>246</td>
<td>8.98%</td>
</tr>
<tr>
<td>HASH</td>
<td>236</td>
<td>8.61%</td>
</tr>
<tr>
<td>LDAPI</td>
<td>59</td>
<td>2.15%</td>
</tr>
<tr>
<td>PATHTRAVER</td>
<td>268</td>
<td>9.78%</td>
</tr>
<tr>
<td>SECURECOOKIE</td>
<td>67</td>
<td>2.45%</td>
</tr>
<tr>
<td>SQLI</td>
<td>504</td>
<td>18.39%</td>
</tr>
<tr>
<td>TRUSTBOUND</td>
<td>126</td>
<td>4.60%</td>
</tr>
<tr>
<td>WEAKRAND</td>
<td>493</td>
<td>17.99%</td>
</tr>
<tr>
<td>XPATHI</td>
<td>35</td>
<td>1.28%</td>
</tr>
<tr>
<td>XSS</td>
<td>455</td>
<td>16.61%</td>
</tr>
<tr>
<td>Total</td>
<td>2740</td>
<td>100.00%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-2">Table 2</xref> shows the percentage of true vulnerabilities and identifies the false positives for each category.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Percentages of TP and FP in the dataset</title>
</caption>
<table>
<colgroup>
<col align="center" width="38mm"/>
<col align="center" width="30mm"/>
<col align="center" width="32mm"/> </colgroup>
<thead>
<tr>
<th>Category vulnerability</th>
<th>False positive (%)</th>
<th>True positive (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CMDI</td>
<td>49.80%</td>
<td>50.20%</td>
</tr>
<tr>
<td>CRYPTO</td>
<td>47.15%</td>
<td>52.85%</td>
</tr>
<tr>
<td>HASH</td>
<td>45.34%</td>
<td>54.66%</td>
</tr>
<tr>
<td>LDAPI</td>
<td>54.24%</td>
<td>45.76%</td>
</tr>
<tr>
<td>PATHTRAVER</td>
<td>50.37%</td>
<td>49.63%</td>
</tr>
<tr>
<td>SECURECOOKIE</td>
<td>46.27%</td>
<td>53.73%</td>
</tr>
<tr>
<td>SQLI</td>
<td>46.03%</td>
<td>53.97%</td>
</tr>
<tr>
<td>TRUSTBOUND</td>
<td>34.13%</td>
<td>65.87%</td>
</tr>
<tr>
<td>WEAKRAND</td>
<td>55.78%</td>
<td>44.22%</td>
</tr>
<tr>
<td>XPATHI</td>
<td>57.14%</td>
<td>42.86%</td>
</tr>
<tr>
<td>XSS</td>
<td>45.93%</td>
<td>54.07%</td>
</tr>
<tr>
<td>Total</td>
<td>48.36%</td>
<td>51.64%</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Large Language Models (LLM)</title>
<p>The continual expansion of different LLM variants and sizes yields a wide array of models, with new ones being trained every day using diverse techniques and/or architectural typologies. To compare the corpora of these models and select a reasonably sized subset that can be run in the near term on a variety of laptops, the following models were chosen, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>LLMs</title>
</caption>
<table>
<colgroup>
<col align="center" width="60mm"/>
<col align="center" width="85mm"/> </colgroup>
<thead>
<tr>
<th>LLM</th>
<th>Main features</th>
</tr>
</thead>
<tbody>
<tr>
<td>CodeLlama 7B instruct FP16</td>
<td>A code-specialized language model built on the Llama architecture.</td>
</tr>
<tr>
<td>DeepSeek coder 6.7B instruct FP16</td>
<td>Trained from scratch on 87% code and 13% natural language in English and Chinese.</td>
</tr>
<tr>
<td>Mistral 7B instruct FP16</td>
<td>Features a 32 K-token context window and is 15 months older than DeepSeek-V3.</td>
</tr>
<tr>
<td>Phi3 3.8B mini 128K instruct FP16</td>
<td>A compact variant offering a 128 K-token context window.</td>
</tr>
<tr>
<td>Qwen2-5 coder 7B instruct FP16</td>
<td>Part of the Qwen 2.5 family, which includes models tailored for mathematics and coding.</td>
</tr>
<tr>
<td>StarCoder2 7B FP16</td>
<td>Trained on the Stack V2 dataset, which is seven times larger than the original StarCoder dataset.</td>
</tr>
<tr>
<td>Gemini 2.0 flash 001</td>
<td>LLM from Google, which core features include multimodal understanding, real-time streaming, native tool integration and code-processing capacity.</td>
</tr>
<tr>
<td>Gemini 1.5 flash 002</td>
<td>LLM from Google, which core features are real-time data transformation, live translation supporting multilingual interaction, and text summarization.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For static code analysis, models that can interpret code are essential. Accordingly, the &#x201C;Instruct&#x201D; variant was selected, as it contains largely code-centric training data and the most accurate Floating-Point 16 (FP16) checkpoint. This provides a superior balance between precision and performance compared to the quantization 3 (Q3), quantization 4 (Q4), quantization 6 (Q6), and quantization 8 (Q8) quantized models, which reduce accuracy to shrink the model size. In <xref ref-type="table" rid="table-4">Table 4</xref>, the detailed available specification of each LLM used in the experiment are presented.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>LLMs characteristics</title>
</caption>
<table>
<colgroup>
<col align="center" width="25mm"/>
<col align="center" width="25mm"/>
<col align="center" width="35mm"/>
<col align="center" width="55mm"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Parameters &#x0026; format</th>
<th>Training data</th>
<th>Key strengths</th>
</tr>
</thead>
<tbody>
<tr>
<td>CodeLlama 7B instruct FP16 [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>7B, FP16</td>
<td>Llama 2 architecture; fine-tuned on 500 B tokens (85% code, 8% related text)</td>
<td>First Meta AI code-specialized model; SOTA on HumanEval, MBPP, MultiPL-E; long context (100K tokens); bug detection and explanation</td>
</tr>
<tr>
<td>DeepSeek coder 6-7B instruct FP16 [<xref ref-type="bibr" rid="ref-43">43</xref>,<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>6-7B, FP16</td>
<td>2 trillion tokens (87% code, 13% English/Chinese)</td>
<td>Leads HumanEval, MBPP, MultiPL-E, DS-1000, APPS; project-level completion; models inter-file dependencies</td>
</tr>
<tr>
<td>Mistral 7B instruct FP16 [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>7B, FP16</td>
<td>General-purpose corpus: web data, books, open-source code</td>
<td>Best public 7B model; outperforms Llama 2-13B and Llama 1-34B on general benchmarks; approaches CodeLlama-7B on code tasks</td>
</tr>
<tr>
<td>Phi-3 3.8B mini 128K instruct FP16 [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>3.8B, FP16</td>
<td>Phi-3 datasets: high-quality synthetic &#x002B; curated web content; supervised fine-tuning &#x0026; preference optimization</td>
<td>Strong reasoning and code understanding; long context (128K tokens); reliable prompt following</td>
</tr>
<tr>
<td>Qwen2-5 coder 7B instruct FP16 [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>7B, FP16</td>
<td>5.5 trillion tokens across 92 languages (70% code, 20% text, 10% math)</td>
<td>Latest Qwen model; improved generation, reasoning, code correction</td>
</tr>
<tr>
<td>Starcoder 2 7B FP16 [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>7B, FP16</td>
<td>The Stack V2: 3.5 trillion code tokens, 17 languages</td>
<td>Multilingual; surpasses CodeLlama 7B on multiple code benchmarks</td>
</tr>
<tr>
<td>Gemini 2.0 flash 001 [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>N/A</td>
<td>Google datasets &#x002B; internal knowledge bases (up to Aug 2024)</td>
<td>Multimodal understanding, real-time streaming, native tool integration, code-processing</td>
</tr>
<tr>
<td>Gemini 1.5 flash 002 [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>N/A</td>
<td>Google datasets &#x002B; internal knowledge bases (up to Sep 2024)</td>
<td>Visual, video, and audio comprehension; real-time transformation; multilingual translation; summarization</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Prompt Techniques</title>
<p>Prompt engineering techniques are structured methods for designing inputs (prompts) to guide a language model&#x2019;s behavior to produce desired outputs. These techniques aim to improve accuracy, relevance, and format consistency by explicitly instructing the model on how to respond. Common strategies include role prompting, instruction framing, output formatting, and contextual conditioning. Same prompt template was provided to all LLMs using the techniques described in the <xref ref-type="table" rid="table-5">Table 5</xref> [<xref ref-type="bibr" rid="ref-51">51</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>].</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Prompt techniques</title>
</caption>
<table>
<colgroup>
<col align="center" width="30mm"/>
<col align="center" width="55mm"/>
<col align="center" width="60mm"/> </colgroup>
<thead>
<tr>
<th>Technique</th>
<th>Definition</th>
<th>Provided prompt</th>
</tr>
</thead>
<tbody>
<tr>
<td>Role prompting</td>
<td>Assigns a specific identity or role to the model to influence its behavior and expertise.</td>
<td>&#x201C;You are a security assistant.&#x201D;</td>
</tr>
<tr>
<td>Instruction framing</td>
<td>Explicitly tells the model what to do or not do, often including constraints.</td>
<td>&#x201C;ONLY answer the following X questions. No reasoning, no explanation.&#x201D;</td>
</tr>
<tr>
<td>Contextual conditioning</td>
<td>Provides contextual information that the model should use to generate the answer.</td>
<td>&#x201C;Considering this category of vulnerabilities: {category}&#x201D;<break/>&#x201C;Result of static analysis of FindSecBugs:&#x005C;n{static_analysis}&#x201D;</td>
</tr>
<tr>
<td>Output formatting/structured output</td>
<td>Specifies exactly how the model should structure its output, often using templates or delimiters.</td>
<td>&#x201C;Strictly follow this format:&#x005C;n&#x005C;nQ1: Is there any vulnerability? &#x0003C;TRUE or FALSE&#x0003E;&#x005C;n...&#x201D;</td>
</tr>
<tr>
<td>Internal vs. external instruction separation</td>
<td>Distinguishes instructions for internal reasoning from what should appear in the final output.</td>
<td>&#x201C;Use angle brackets &#x0003C;...&#x003E; only in your internal response, but the final JSON output must NOT include them.&#x201D;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To ensure controlled and consistent outputs during the experiment, the following generation parameters were used. The temperature was set to 0.0, making the model&#x2019;s responses deterministic and minimizing randomness. Top-p sampling was set to 0.9 and top-k sampling to 20, balancing the diversity of outputs while restricting the model to plausible continuations. A repeat penalty of 1.0 was applied, indicating no explicit discouragement of token repetition. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> shows the prompt used in the experiment, while blue shows the difference in the second prompt.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Prompts used in the experiment</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-1.tif"/>
</fig>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Metrics</title>
<p>Below are the metrics presented in the <xref ref-type="table" rid="table-6">Table 6</xref> that will be used to evaluate static vulnerability analysis. All of them have wide acceptance in academia and are commonly employed for static code verification.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Metrics</title>
</caption>
<table>
<colgroup>
<col align="center" width="30mm"/>
<col align="center" width="50mm"/>
<col align="center" width="60mm"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Description</th>
<th>Calculation</th>
</tr>
</thead>
<tbody>
<tr>
<td>True positive (TP)</td>
<td>The vulnerability is real, and the tool correctly detects it.</td>
<td><inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>False negative (FN)</td>
<td>A real vulnerability is not detected (worst case).</td>
<td><inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>False positive (FP)</td>
<td>The tool flags a vulnerability that does not exist.</td>
<td><inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>True negative (TN)</td>
<td>No vulnerability is present, and the tool correctly gives no alert.</td>
<td><inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Precision</td>
<td>How many of the alerts are correct (low FP).</td>
<td><inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>Recall</td>
<td>How many of the real vulnerabilities are found (low FN).</td>
<td><inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>Accuracy</td>
<td>Overall proportion of correct predictions.</td>
<td><inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>F1 score</td>
<td>Harmonic mean of precision and recall.</td>
<td><inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow></mml:math></inline-formula><inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Ratio false positive</td>
<td>Fraction of real vulnerabilities that are detected.</td>
<td><inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>Ratio true positive</td>
<td>Fraction of non-vulnerable code that is correctly ignored.</td>
<td><inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>Ratio true negative</td>
<td>Fraction of non-vulnerable code that is incorrectly flagged.</td>
<td><inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>Markedness</td>
<td>Indicates the reliability of the predictions.<break/>1 &#x2192; Correct<break/>0 &#x2192; Random<break/>0 &#x003E; &#x2192; Predict better than expected</td>
<td><inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>V</mml:mi><mml:mo>+</mml:mo><mml:mi>N</mml:mi><mml:mi>P</mml:mi><mml:mi>V</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula><break/><inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>P</mml:mi><mml:mi>P</mml:mi><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula><break/><inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>N</mml:mi><mml:mi>P</mml:mi><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>Informedness</td>
<td>Indicates how well it distinguishes between classes (vulnerabilities and non-vulnerabilities).<break/>1 &#x2192; Correct<break/>0 &#x2192; Random<break/>0 &#x003E; &#x2192; Predict worse than expected</td>
<td><inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>McNemar</td>
<td>Indicate no statistically significant difference:<break/><inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x003C;</mml:mo><mml:mn>3.84</mml:mn></mml:math></inline-formula><break/>Indicate a highly significant difference in classification patterns<break/><inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x226B;</mml:mo><mml:mn>3.85</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>On the other hand, the results will be categorized according to different objectives:
<list list-type="bullet">
<list-item>
<p>Business-critical applications: Tools that detect the largest number of vulnerabilities or the fewest vulnerabilities per detection. These scenarios have the necessary resources to fix all vulnerabilities and to verify or correct false positives.</p></list-item>
<list-item>
<p>Non-critical applications: Tools that detect a high quantity of vulnerabilities while minimizing false positives. This category includes companies where a leak or vulnerability could cause significant losses.</p></list-item>
<list-item>
<p>Best effort: Tools that detect a large number of vulnerabilities while reporting few false positives.</p></list-item>
<list-item>
<p>Minimum effort: Tools that detect the fewest false positives. Focused on small and medium-sized businesses. [<xref ref-type="bibr" rid="ref-25">25</xref>].</p></list-item>
</list></p>
<p>The metrics for the different scenarios will be evaluated as shown in the <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Metric per scenario</title>
</caption>
<table>
<colgroup>
<col align="center" width="27mm"/>
<col align="center" width="39mm"/>
<col align="center" width="45mm"/> </colgroup>
<thead>
<tr>
<th>Scenario</th>
<th>Recommended metric</th>
<th>Recommended tiebreaker</th>
</tr>
</thead>
<tbody>
<tr>
<td>Business critical</td>
<td>Recall</td>
<td>Precision</td>
</tr>
<tr>
<td>Non critical</td>
<td>Informedness</td>
<td>Recall</td>
</tr>
<tr>
<td>Best effort</td>
<td>F-measure</td>
<td>Recall</td>
</tr>
<tr>
<td>Minimum effort</td>
<td>Markedness</td>
<td>Precision</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_6">
<label>2.6</label>
<title>Relative Work</title>
<p>In the context of SAST tools, the most recent and relevant work related to advances in the use of LLMs for static detection of security vulnerabilities in software without requiring application execution has been researched and analyzed.</p>
<p>Khare et al. [<xref ref-type="bibr" rid="ref-52">52</xref>] evaluated LLMs across languages and datasets, finding moderate accuracy&#x2014;strong on simple vulnerabilities but weak on complex reasoning tasks. Prompting techniques like step-by-step reasoning improve results. While LLMs sometimes outperform static analysis tools, they are not yet reliable for full-scale vulnerability detection, but it shows great potential as complementary tools. Similarly, Das Purba et al. [<xref ref-type="bibr" rid="ref-53">53</xref>] compared GPT-3.5, GPT-4, Davinci, and CodeGen for SQL injection and buffer overflow detection, observing high recall, but low precision due to false positives. Guo et al. [<xref ref-type="bibr" rid="ref-54">54</xref>], Yin et al. [<xref ref-type="bibr" rid="ref-55">55</xref>], and Shimmi et al. [<xref ref-type="bibr" rid="ref-56">56</xref>] further showed that performance varies widely depending on prompting strategy, dataset quality, and model fine-tuning, while Almeida [<xref ref-type="bibr" rid="ref-57">57</xref>] demonstrated that few-shot and chain-of-thought prompts increase consistency. Nevertheless, two recent, comprehensive surveys by Sheng et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] and Zhou et al. [<xref ref-type="bibr" rid="ref-58">58</xref>] caution that LLMs are currently limited by insufficient contextual awareness for complex, inter-file dependencies, leading research to focus on function-level detection rather than practical repository-level analysis. In contrast, our work addresses these limitations by evaluating a broader set of vulnerability categories, integrating context from a Static Application Security Testing (SAST) tool, and utilizing multiple LLMs with a unified taxonomy and consistent dataset, enabling robust cross-model comparisons under similar conditions.</p>
<p>Some studies focus on the security awareness of LLMs. For example, Sajadi et al. [<xref ref-type="bibr" rid="ref-59">59</xref>] found that GPT-4, Claude 3, and LLaMA 3 rarely issue security warnings unless explicitly prompted. Other studies go further, examining LLM-generated code; Ayd&#x0131;n and Bahtiyar [<xref ref-type="bibr" rid="ref-60">60</xref>] revealed that JavaScript produced by LLMs often contains vulnerabilities such as XSS or unsafe use of eval(). In contrast, our study focuses on evaluating LLMs&#x2019; ability to detect vulnerabilities in existing source code, rather than generating or self-auditing new code.</p>
<p>Hybrid approaches have shown in combining LLMs with traditional analyzers. Jaoua et al. [<xref ref-type="bibr" rid="ref-61">61</xref>] demonstrated that integrating static analyzer results during training or inference enhances review accuracy, and Munson et al. [<xref ref-type="bibr" rid="ref-62">62</xref>] used Semgrep with LLMs to reduce false positives. This necessity for hybrid models is further supported by Zhou et al. [<xref ref-type="bibr" rid="ref-58">58</xref>], whose survey highlights that blending LLMs with program analysis or other modules is a primary adaptation technique to overcome the contextual limitations of language models. Our work extends this hybrid direction by integrating the FindSecBugs analyzer with multiple LLM architectures and prompt strategies, applied to the OWASP Benchmark dataset.</p>
<p>Beyond detection, LLMs have been used for localization and malicious code analysis. Wu et al. [<xref ref-type="bibr" rid="ref-63">63</xref>] proposed VFFinder to identify vulnerable functions in source code by analyzing natural language CVE descriptions, which in experiments on open-source projects improves vulnerability localization accuracy compared to traditional methods. Hossain et al. [<xref ref-type="bibr" rid="ref-64">64</xref>] trained a Mixtral-based LLM to detect malicious Java snippets, outperforming traditional SAST tools but requiring several iterative refinements. Additionally, Blefari et al. [<xref ref-type="bibr" rid="ref-65">65</xref>] proposed SecFlow, an agentic LLM-based framework for modular post-event cyberattack analysis and explainability from raw logs, utilizing a RAG component for contextualized reasoning. Similarly, Belcastro et al. [<xref ref-type="bibr" rid="ref-66">66</xref>] introduced KLAGE, a methodology that integrates Knowledge Graphs, XAI (LIME) and LLMs to enhance network threat detection, classification, and generation of explainable reports from network traffic logs. He et al. [<xref ref-type="bibr" rid="ref-67">67</xref>] showed that combining LLM-derived features with machine learning improves defect detection, particularly with few-shot prompting. While these works enhance specific tasks such as localization or malware detection, our study targets comprehensive vulnerability identification. The study of Li et al. [<xref ref-type="bibr" rid="ref-36">36</xref>] leverages LLM reasoning to automatically inspect the results of very broad SAST tools. They develop a GPT-based prototype, called FPShield, to automatically identify and eliminate potential false positives from SAST results.</p>
<p>A separate line of research of general LLM used on highly specialized fine-tuned transformers is addressed by some studies. Smaili et al. [<xref ref-type="bibr" rid="ref-68">68</xref>] proposed a transformer-based framework (CodeGATNet) that utilizes a specialized model (CodeBERT) in combination with an Attention-Driven Convolutional Neural Network component, this together with long-range data-flow dependencies through a Convolutional Attention Network (CAN) offers an alternative than relying on complex prompting of LLMs, However, this specialization is a limitation: adapting CodeGATNet to new languages or vulnerability classes typically demands a full, costly model retraining, whereas general LLMs can often integrate new context via simple prompting and inference-time augmentation.</p>
<p>Finally, [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed evaluating detection tools using four effort-based categories&#x2014;Business Critical, Non-Critical, Best Effort, and Minimum Effort&#x2014;which we adopt to contextualize our results.</p>
<p>Overall, our work presents an evaluation of vulnerability detection using lightweight LLMs that can be run on standard laptops, as well as LLMs accessed via API-based services. The study integrates one of the most widely used static application security testing (SAST) tools, FindSecBugs, and employs a dataset derived from OWASP&#x2019;s benchmark test suite. The evaluation focuses on measuring detection precision and categorizing the results into four practical risk categories: Business Critical, Non-Critical, Best Effort, and Minimum Effort from an overall perspective, taking into account all security vulnerabilities, and also from the perspective of each specific vulnerability. While we restrict our analysis to Java programming language and a synthetic dataset, this choice allowed us to validate and compare the performance of multiple LLM architectures under the exact same conditions and scenarios, significantly reducing external interference from variables like differing code bases or language complexities. Furthermore, the OWASP Benchmark provides good ground truth, which is essential for the precise measurement of detection metrics and the robust categorization of results into our four risks.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methods</title>
<p>In the following chapter the experiment is described in the <xref ref-type="sec" rid="s3_1">Section 3.1</xref>, along with its results at the <xref ref-type="sec" rid="s3_2">Section 3.2</xref>.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Research Questions Definition</title>
<p>The experiment aims to investigate the effectiveness of using LLMs for software vulnerability, Specifically, we address the following questions or hypotheses:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>Q1. Do the results produced by LLMs outperform those of traditional static analysis tools?</p>
<p>This question focuses on comparing the accuracy of LLMs with SAST tools in detecting software vulnerabilities, rather than evaluating computational performance or runtime efficiency. The goal is to assess whether LLMs could potentially serve as an accurate alternative or complement to traditional tools.</p></list-item>
<list-item><label>&#x2022;</label><p>Q2. Is there an improvement when combining LLMs with SAST tools?</p>
<p>Hybrid approach to integrate LLMs with SAST tool in order to evaluate the combination also in a bigger LLMs.</p></list-item>
<list-item><label>&#x2022;</label><p>Q3. Can reliable results be achieved with locally deployed models on a standard desktop computer?</p></list-item>
</list></p>
<p>These questions focus on comparing local LLMs which could improve the privacy in large companies as well as the anticipation of vulnerabilities.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Description of the Methodology</title>
<p>The preparation of the experiment involves several key phases, as outlined in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> below:</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Phases of the methodology</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-2.tif"/>
</fig>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Selection of LLM, SAST and Dataset</title>
<p>At the time the experiment was conducted, the most relevant model architectures that could feasibly run on a laptop (and were specifically trained for code generation and understanding) were: CodeLlama 7B, DeepSeek Coder 6.7B, Mistral 7B, Phi-3 Mini, Qwen2-5 Coder 7B, and StarCoder2.</p>
<p>To select an appropriate static application security testing (SAST) tool for Java, several open-source solutions were evaluated using the OWASP Benchmark. FindSecBugs, in its latest version, demonstrated the highest detection accuracy among the tools tested. Its widespread adoption and continued community support further reinforce its suitability for integration into secure software development workflows [<xref ref-type="bibr" rid="ref-41">41</xref>].</p>
<p>To ensure focused evaluation and minimize noise from unrelated factors, a synthetic dataset specifically designed for Java was used. This dataset provides a wide range of vulnerability categories relevant to static analysis [<xref ref-type="bibr" rid="ref-41">41</xref>,<xref ref-type="bibr" rid="ref-62">62</xref>].</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Dataset Preparation</title>
<p>To reduce bias and prevent direct influence on the models, the dataset will be cleaned by removing all explicit references to specific vulnerabilities, Common Weakness Enumeration (CWE) identifiers, and expected outcomes. This de-contamination ensures that the models operate solely on the contextual information available, allowing us to evaluate their reasoning and generalization abilities in environments that lack explicit prior knowledge.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Elements of the Experiment</title>
<p>A script will be developed in Python, taking advantage of its extensive library ecosystem and its seamless integration with large-language-model frameworks such as Ollama. This choice helps reduce errors that can arise from unstable or incompatible libraries.</p>
<p>For the evaluation phase, we will employ Benchmark Java-OWASP [<xref ref-type="bibr" rid="ref-41">41</xref>], an official OWASP project that provides a curated dataset and a standardized framework for running test classes on SAST tools. The benchmark will be extended to include large-language-model (LLM) evaluations, as the original repository does not support comparisons with LLM-based systems. For this reason, we will develop additional components that execute the benchmark scenarios and capture the predictions generated by the LLMs, enabling a comparison between traditional SAST tools and the proposed LLM approach.</p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>Experiment Execution</title>
<p>A total of 2740 scenarios will be processed. For every scenario the original input and the LLM&#x2019;s response will be captured and stored in a structured format, enabling a detailed comparison after the fact. The entire workflow will be automated to guarantee reproducibility, consistency of results, and full traceability of the experiment. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> illustrates the process of the first prompt execution, while the <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the process of the second prompt execution including the SAST analysis.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Execution of first prompt</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-3.tif"/>
</fig><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Execution of the second prompt</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-4.tif"/>
</fig>
<p>Two full test-suite runs will be performed:
<list list-type="order">
<list-item>
<p>LLM with detection vulnerability questions illustrated in the <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p></list-item>
<list-item>
<p>Combination of the results in SAST Tool (FindSecBugs) and LLM illustrated in the <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, while the <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows an example of SAST result into the prompt.</p>
</list-item>
</list></p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Example of content included as SAST into the prompt</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_2_5">
<label>3.2.5</label>
<title>Results Collection</title>
<p>The results produced by the models will be processed to normalize their format and facilitate comparison with other techniques such as OWASP Benchmark. This process will include tasks such as text cleaning, semantic labeling, and extraction of key performance indicators. The resulting data will be stored in a structured database, ready for quantitative and qualitative analysis using selecting metrics.</p>
</sec>
<sec id="s3_2_6">
<label>3.2.6</label>
<title>Analysis and Discussion</title>
<p>The evaluation will use standard performance metrics (precision, recall, F1-score, and true-positive ratio) that are widely adopted for vulnerability detection [<xref ref-type="bibr" rid="ref-25">25</xref>]. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, all results will be aggregated and grouped to facilitate a direct comparison with conventional static-analysis tools. In addition, we will conduct a preliminary quality assessment to confirm that the LLM outputs lie within the acceptable range.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Process to compare results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-6.tif"/>
</fig>
<p>The analyzed results and conclusions will be presented in a clear and structured manner, using comparative charts, summary tables, and qualitative interpretations to facilitate understanding of the findings. The goal is to provide a comprehensive view of the model performance, highlighting strengths, identifying limitations, and pointing out areas for improvement in future work.</p>
</sec>
<sec id="s3_2_7">
<label>3.2.7</label>
<title>Method Limitations and Considerations</title>
<p>While we limit our analysis to the Java programming language and a synthetic dataset, this controlled setting allows for a fair comparison of multiple LLM architectures under identical conditions, minimizing external factors such as varying codebases or language-specific complexities. The dataset covers a representative spectrum of security issues, including command injection, path traversal, weak hash algorithms, cryptographic algorithm misuse, cross-site scripting (XSS), LDAP-related flaws, and trust-boundary violations, thereby providing a comprehensive evaluation of model capabilities. Although our experiments focus on open-source LLMs, the framework could be extended to include newer open-source or commercial LLMs of different sizes, as well as traditional SAST tools, offering broader benchmarking opportunities. Nevertheless, practical considerations&#x2014;such as dataset size, vulnerability diversity, and computational cost&#x2014;necessitate imposing limits when selecting models and scenarios to ensure consistency, reproducibility, and meaningful results.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Implementation of the Methodology</title>
<p>The steps of the implementation are the following:</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Dataset Preparation</title>
<p>An exhaustive review of all the classes in the dataset was carried out to remove references that might disclose the type of vulnerability involved. In particular, patterns such as those defined in the @WebServelet annotation were identified, for example:</p>
<p><italic>@WebServlet (value &#x003D; &#x201C;/pathtraver-00/BenchmarkTest00001&#x201D;)</italic></p>
<p>To automate this cleansing step, we developed a dedicated script that performs bulk removal of such references, thereby ensuring the semantic neutrality of the classes with respect to the vulnerability type they represent. The script is available in the repository provided in this paper &#x201C;CleanReferences.py&#x201D;.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Elements of the Experiment</title>
<p>The element that has been required to implement or configure are the following:
<list list-type="bullet">
<list-item>
<p>Environment preparation</p></list-item>
<list-item>
<p>Script Implementation</p></list-item>
<list-item>
<p>Extension of the application OWASP Benchmark</p></list-item>
</list></p>
<p>Environment preparation. For the development and execution of the tests, two environments were prepared. The first is a local workstation equipped with Apple Silicon (M1) architecture, primarily used for analysis, development, data processing, and running certain LLMs. The second environment consists of a remote server specifically configured for running the experiment and collecting associated metrics. Further technical details about the experiment hardware are provided in <xref ref-type="table" rid="table-8">Table 8</xref>.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Hardware specifications</title>
</caption>
<table>
<colgroup>
<col align="center" width="31mm"/>
<col align="center" width="31mm"/>
<col align="left" width="45mm"/> </colgroup>
<thead>
<tr>
<th>Hardware</th>
<th>Type</th>
<th align="center">Technical specifications</th>
</tr>
</thead>
<tbody>
<tr>
<td>Apple M1 (32 GB RAM) de 2021</td>
<td>Local</td>
<td>Apple M1 Pro chip<break/>&#x2022;&#x2003; 10-core CPU with 8 performance cores and 2 efficiency cores<break/>&#x2022;&#x2003; 16-core GPU<break/>&#x2022;&#x2003; 16-core Neural Engine</td>
</tr>
<tr>
<td>Server&#x2013;Runpod</td>
<td>In cloud</td>
<td>1 x RTX 4080 SUPER<break/>21 vCPU<break/>41 GB RAM</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Script implementation. The script is designed to scan multiple Java source files, leveraging large-scale language models (LLMs) to identify potential security vulnerabilities. Analysis is performed via prompt engineering: each file is sent to an LLM (Gemini or Ollama), and both the prompt and the model&#x2019;s response are stored in structured text files.</p>
<p>Ollama supports executing language models in both local and remote environments and provides a broad catalog of pre-trained models accessible through its libraries. In this work, a variety of strategies were applied to ensure the robustness, reproducibility, and efficiency of the automated analysis:
<list list-type="bullet">
<list-item>
<p>Standardization of the output format: The models were forced to return responses in JSON format, thereby facilitating subsequent automated analysis and reducing ambiguity in result interpretation.</p></list-item>
<list-item>
<p>Execution in multiple iterations: For some models, several runs were performed on the same source files to assess the stability and consistency of their responses, as well as to correctly configure the model&#x2019;s deterministic outputs.</p></list-item>
<list-item>
<p>Integration with Gemini and usage-limit management: The Gemini model was added as an additional backend, with programmed delays between requests to respect API usage restrictions.</p></list-item>
<list-item>
<p>Optimization toward deterministic outputs: Generation parameters were tuned (e.g., temperature &#x003D; 0.0) to prioritize more deterministic responses over random or probabilistic ones, enabling a more precise and reproducible evaluation</p></list-item>
</list></p>
<p>Extension of the application OWASP Benchmark. The OWASP Benchmark [<xref ref-type="bibr" rid="ref-41">41</xref>] is a widely adopted standard reference for comparative evaluation of security-analysis tools, including static, dynamic, and hybrid scanners. However, it does not natively support integration with large-scale language models (LLMs).</p>
<p>In addition, this Benchmark includes a ranking system that allows tools to be compared across distinct vulnerability categories. To leverage these capabilities, we have developed an adaptation that enables LLMs to participate in the analysis workflow, thereby extending the tool&#x2019;s functionality and introducing new automated evaluation modes over Benchmark&#x2019;s test suite.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Experiment Execution</title>
<p><xref ref-type="table" rid="table-9">Table 9</xref> shows the execution time of the single LLM and the SAST with the LLM. The One-Shot strategy has demonstrated faster performance compared to using SAST results, even without accounting for the delay introduced by the free-tier limitations of LLM as a Service.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Execution time approx</title>
</caption>
<table>
<colgroup>
<col align="center" width="60mm"/>
<col align="center" width="30mm"/>
<col align="center" width="55mm"/> </colgroup>
<thead>
<tr>
<th>Model LLM</th>
<th>Type</th>
<th>Execution time approx.</th>
</tr>
</thead>
<tbody>
<tr>
<td>codellama-7b-instruct</td>
<td>One shot</td>
<td>50 min</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>One shot</td>
<td>34 min</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>One shot</td>
<td>274 min&#x002A;</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>One shot</td>
<td>277 min&#x002A;</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>One shot</td>
<td>55 min</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>One shot</td>
<td>173 min</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>One shot</td>
<td>58 min</td>
</tr>
<tr>
<td>starcoder2-7b-fp16</td>
<td>One shot</td>
<td>46 min</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>One shot &#x002B; SAST</td>
<td>196 min</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>One shot &#x002B; SAST</td>
<td>300 min</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-9fn1" fn-type="other">
<p>Note: &#x002A;Include delay for the free version.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3_3_4">
<label>3.3.4</label>
<title>Results Collection</title>
<p>All LLM models were executed, generating 2740 outputs for subsequent analysis. Based on the outputs a script is implemented to convert the results into a CSV file, enabling the OWASP Benchmark application (along with its built-in extensions) to perform the comparative analysis.</p>
<p>The process begins with a consistency check of the outputs, verifying that each LLM has produced a response that falls within the permitted value ranges.</p>
<p>Following this verification, preliminary analyses are presented before the final results. The hallucinations generated by each LLM are flagged, indicating either invented categories or CWEs that lie outside the acceptable range.</p>
<p><xref ref-type="table" rid="table-10">Table 10</xref> displays the number of &#x201C;hallucinations&#x201D;, when the output was out for the range proposed in the prompt, and their percentage relative to the total. To ensure that the results are more deterministic than probabilistic, since LLMs have a probability factor, there are several instances of model execution to test that the values are very similar and to verify with certainty that the result is not a matter of luck.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Quantity of hallucination per LLM in the first prompt</title>
</caption>
<table>
<colgroup>
<col align="center" width="50mm"/>
<col align="center" width="10mm"/>
<col align="center" width="13mm"/>
<col align="center" width="14mm"/>
<col align="center" width="16mm"/>
<col align="center" width="12mm"/> </colgroup>
<thead>
<tr>
<th>Modelo LLM</th>
<th>Cat.</th>
<th>CWE</th>
<th>Cat. (%)</th>
<th>CWE (%)</th>
<th>Env.</th>
</tr>
</thead>
<tbody>
<tr>
<td>codellama-7b-instruct</td>
<td>769</td>
<td>2298</td>
<td>28</td>
<td>83.8</td>
<td>runPod</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>885</td>
<td>2329</td>
<td>32.3</td>
<td>85</td>
<td>runPod</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>229</td>
<td>1103</td>
<td>8.4</td>
<td>40.2</td>
<td>M1</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>204</td>
<td>1092</td>
<td>7.4</td>
<td>39.8</td>
<td>M1</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>149</td>
<td>1045</td>
<td>5.4</td>
<td>38.1</td>
<td>M1</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>111</td>
<td>2670</td>
<td>4</td>
<td>97.4</td>
<td>Google</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>81</td>
<td>666</td>
<td>3</td>
<td>24.3</td>
<td>Google</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>268</td>
<td>2655</td>
<td>9.8</td>
<td>96.9</td>
<td>runPod</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>233</td>
<td>2647</td>
<td>8.5</td>
<td>96.6</td>
<td>runPod</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>230</td>
<td>2656</td>
<td>8.4</td>
<td>96.9</td>
<td>runPod</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>315</td>
<td>2535</td>
<td>11.5</td>
<td>92.5</td>
<td>M1</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>310</td>
<td>2541</td>
<td>11.3</td>
<td>92.7</td>
<td>M1</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>317</td>
<td>2565</td>
<td>11.5</td>
<td>93.6</td>
<td>M1</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>309</td>
<td>1102</td>
<td>11.3</td>
<td>40.2</td>
<td>runPod</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>295</td>
<td>1132</td>
<td>10.7</td>
<td>41.3</td>
<td>runPod</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>272</td>
<td>1129</td>
<td>9.9</td>
<td>41.2</td>
<td>runPod</td>
</tr>
<tr>
<td>starcoder2-7b-fp16</td>
<td>601</td>
<td>2594</td>
<td>21.9</td>
<td>94.7</td>
<td>runPod</td>
</tr>
<tr>
<td>starcoder2-7b-fp16</td>
<td>545</td>
<td>2594</td>
<td>19.9</td>
<td>94.6</td>
<td>runPod</td>
</tr>
<tr>
<td>starcoder2-7b-fp16</td>
<td>530</td>
<td>2575</td>
<td>19.3</td>
<td>93.9</td>
<td>runPod</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-10fn1" fn-type="other">
<p>Note: Table columns: Category, CWE. Category percentage with respect to the total, CWE percentage with respect to the total.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>For the first prompt, Gemini 2-0 Flash showed the lowest hallucination rate relative to the provided categories and CWEs, while Gemini 1.5 Flash performed worst in returning valid CWEs.</p>
<p>In <xref ref-type="table" rid="table-11">Table 11</xref>, the &#x201C;hallucinations&#x201D; from the second prompt are displayed, and they already include a prior analysis performed by a static-analysis tool that lists the CWE, possible category, and description of the potential vulnerability.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Quantity of hallucination per LLM in the second prompt</title>
</caption>
<table>
<colgroup>
<col align="center" width="45mm"/>
<col align="center" width="10mm"/>
<col align="center" width="13mm"/>
<col align="center" width="14mm"/>
<col align="center" width="16mm"/>
<col align="center" width="12mm"/> </colgroup>
<thead>
<tr>
<th>Model LLM</th>
<th>Cat.</th>
<th>CWE</th>
<th>Cat. (%)</th>
<th>CWE (%)</th>
<th>Env.</th>
</tr>
</thead>
<tbody>
<tr>
<td>deepseek-coder 6 7b instruct</td>
<td>0</td>
<td>203</td>
<td>0</td>
<td>7.4</td>
<td>M1</td>
</tr>
<tr>
<td>gemini 2 0 flash-001</td>
<td>136</td>
<td>477</td>
<td>5</td>
<td>17.4</td>
<td>Google</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The results for DeepSeek Coder 6-7B-Instruct are better than for Gemini 2-0-Flash 001. Based on the hallucination outcomes, this could be seen as a lack of corpus update regarding the CWEs; therefore, the LLMs have been evaluated on the <italic>vulnerability category</italic> rather than on specific CWEs for the first prompt.</p>
<p>For the second prompt, which includes the results of static code analysis, the category is fed as input, and the model is asked to confirm whether it is indeed a true positive (TP) or a false positive (FP) and whether it preserves the category or changes it in the output.
<list list-type="bullet">
<list-item>
<p>Actual Category &#x003D;&#x003D; Expected Category &#x0026;&#x0026; Is a Real Vulnerability &#x003D;&#x003D; Is a Vulnerability</p></list-item>
</list></p>
<p>Below are the overall results with all metrics broken down by category: Command injection, insecure cookie, LDAP injection, path traversal, SQL injection, trust boundary, weak encryption algorithm, weak hashing algorithm, weak randomness, XPath injection, and XSS.</p>
<p><bold>Overall Results</bold></p>
<p>In the <xref ref-type="fig" rid="fig-7">Fig. 7a</xref> bar chart summarizing the results obtained for all vulnerabilities across the four scenarios (Business critical, Non critical, Best effort, minimum effort) in contrast, <xref ref-type="table" rid="table-12">Table 12</xref> ranked by F1-S presents the overall metrics. In scenarios where detecting the most real vulnerabilities is critical, such as in high-assurance systems where Recall is the primary metric, FindSecBugs v1.4.6 delivered the best results. However, for scenarios that prioritize balanced detection, minimal manual effort, and more accurate findings, the best performance came from Gemini 2 together with FindSecBugs results. The metric McNemar reflects statistical disagreement or difference in predictions between tools.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Overall results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-7.tif"/>
</fig><table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Overall results</title>
</caption>
<table>
<colgroup>
<col align="center" width="28mm"/>
<col align="center" width="12mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
<th>McN</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.93</td>
<td>0.03</td>
<td>0.97</td>
<td>0.90</td>
<td>0.92</td>
<td>0.88</td>
<td>0.90</td>
<td>0.81</td>
<td>0.90</td>
<td>14.04</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.83</td>
<td>0.12</td>
<td>0.88</td>
<td>0.86</td>
<td>0.90</td>
<td>0.82</td>
<td>0.86</td>
<td>0.72</td>
<td>0.71</td>
<td>40.92</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>0.97</td>
<td>0.58</td>
<td>0.42</td>
<td>0.74</td>
<td>0.69</td>
<td>0.92</td>
<td>0.79</td>
<td>0.56</td>
<td>0.39</td>
<td>327.27</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>0.97</td>
<td>0.53</td>
<td>0.47</td>
<td>0.74</td>
<td>0.70</td>
<td>0.85</td>
<td>0.77</td>
<td>0.49</td>
<td>0.44</td>
<td>116.16</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.60</td>
<td>0.19</td>
<td>0.81</td>
<td>0.76</td>
<td>0.83</td>
<td>0.67</td>
<td>0.74</td>
<td>0.53</td>
<td>0.41</td>
<td>101.81</td>
</tr>
<tr>
<td>deepseek-coder-<break/>6-7b-inst &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.70</td>
<td>0.35</td>
<td>0.65</td>
<td>0.68</td>
<td>0.69</td>
<td>0.68</td>
<td>0.69</td>
<td>0.36</td>
<td>0.35</td>
<td>0.5</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.64</td>
<td>0.13</td>
<td>0.87</td>
<td>0.71</td>
<td>0.81</td>
<td>0.57</td>
<td>0.67</td>
<td>0.47</td>
<td>0.51</td>
<td>236.64</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.61</td>
<td>0.11</td>
<td>0.89</td>
<td>0.71</td>
<td>0.84</td>
<td>0.55</td>
<td>0.67</td>
<td>0.49</td>
<td>0.50</td>
<td>223.29</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.62</td>
<td>0.11</td>
<td>0.89</td>
<td>0.71</td>
<td>0.82</td>
<td>0.57</td>
<td>0.67</td>
<td>0.47</td>
<td>0.50</td>
<td>291.43</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.33</td>
<td>0.04</td>
<td>0.96</td>
<td>0.67</td>
<td>0.85</td>
<td>0.44</td>
<td>0.58</td>
<td>0.46</td>
<td>0.29</td>
<td>488.8</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.31</td>
<td>0.05</td>
<td>0.95</td>
<td>0.67</td>
<td>0.84</td>
<td>0.45</td>
<td>0.58</td>
<td>0.45</td>
<td>0.26</td>
<td>530.89</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.30</td>
<td>0.06</td>
<td>0.94</td>
<td>0.66</td>
<td>0.84</td>
<td>0.42</td>
<td>0.57</td>
<td>0.44</td>
<td>0.24</td>
<td>534.28</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.17</td>
<td>0.25</td>
<td>0.75</td>
<td>0.54</td>
<td>0.61</td>
<td>0.31</td>
<td>0.41</td>
<td>0.13</td>
<td>&#x2212;0.09</td>
<td>376.07</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.18</td>
<td>0.24</td>
<td>0.76</td>
<td>0.52</td>
<td>0.57</td>
<td>0.30</td>
<td>0.39</td>
<td>0.07</td>
<td>&#x2212;0.06</td>
<td>340.63</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.17</td>
<td>0.26</td>
<td>0.74</td>
<td>0.52</td>
<td>0.57</td>
<td>0.30</td>
<td>0.39</td>
<td>0.08</td>
<td>&#x2212;0.08</td>
<td>343.96</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.16</td>
<td>0.13</td>
<td>0.87</td>
<td>0.54</td>
<td>0.63</td>
<td>0.25</td>
<td>0.36</td>
<td>0.14</td>
<td>0.03</td>
<td>636.95</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.12</td>
<td>0.88</td>
<td>0.55</td>
<td>0.67</td>
<td>0.25</td>
<td>0.36</td>
<td>0.19</td>
<td>&#x2212;0.01</td>
<td>569.79</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.11</td>
<td>0.12</td>
<td>0.88</td>
<td>0.49</td>
<td>0.51</td>
<td>0.21</td>
<td>0.30</td>
<td>&#x2212;0.01</td>
<td>&#x2212;0.01</td>
<td>498.5</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.09</td>
<td>0.12</td>
<td>0.88</td>
<td>0.48</td>
<td>0.50</td>
<td>0.21</td>
<td>0.29</td>
<td>&#x2212;0.02</td>
<td>&#x2212;0.03</td>
<td>483.34</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.11</td>
<td>0.13</td>
<td>0.87</td>
<td>0.48</td>
<td>0.48</td>
<td>0.14</td>
<td>0.22</td>
<td>&#x2212;0.05</td>
<td>&#x2212;0.03</td>
<td>691.78</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.05</td>
<td>0.07</td>
<td>0.93</td>
<td>0.47</td>
<td>0.45</td>
<td>0.08</td>
<td>0.14</td>
<td>&#x2212;0.07</td>
<td>&#x2212;0.02</td>
<td>934.44</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.04</td>
<td>0.06</td>
<td>0.94</td>
<td>0.48</td>
<td>0.45</td>
<td>0.07</td>
<td>0.12</td>
<td>&#x2212;0.07</td>
<td>&#x2212;0.02</td>
<td>986.14</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.05</td>
<td>0.06</td>
<td>0.94</td>
<td>0.47</td>
<td>0.43</td>
<td>0.06</td>
<td>0.10</td>
<td>&#x2212;0.09</td>
<td>&#x2212;0.01</td>
<td>1050.63</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-12fn1" fn-type="other">
<p>Note: Table columns: Tool, Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness, McNemar; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
 
<p><bold>Command Injection</bold></p>
<p>In the <xref ref-type="fig" rid="fig-8">Fig. 8a</xref> bar chart summarizing the results obtained for Command Injection across the four scenarios (Business critical and Best effort) in contrast, <xref ref-type="table" rid="table-13">Table 13</xref> ranked by F1-S summarizes the key performance metrics for detecting command injection vulnerabilities. Gemini v1.5 was the top performer in the business-critical scenario, achieving the highest Recall, the most important metric when missing a real vulnerability is unacceptable. Additionally, it maintained strong Precision across other scenarios, making it suitable even when balancing detection quality and effort.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Command injection results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-8.tif"/>
</fig><table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Command injection results</title>
</caption>
<table>
<colgroup>
<col align="center" width="29mm"/>
<col align="center" width="12mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/>
<col align="center" width="8mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.99</td>
<td>0.02</td>
<td>0.98</td>
<td>0.99</td>
<td>0.98</td>
<td>0.99</td>
<td>0.99</td>
<td>0.98</td>
<td>0.98</td>
</tr>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.92</td>
<td>0.00</td>
<td>1.00</td>
<td>0.96</td>
<td>1.00</td>
<td>0.92</td>
<td>0.96</td>
<td>0.93</td>
<td>0.92</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.73</td>
<td>0.10</td>
<td>0.90</td>
<td>0.82</td>
<td>0.88</td>
<td>0.73</td>
<td>0.80</td>
<td>0.65</td>
<td>0.63</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.71</td>
<td>0.09</td>
<td>0.91</td>
<td>0.81</td>
<td>0.89</td>
<td>0.71</td>
<td>0.79</td>
<td>0.65</td>
<td>0.63</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.72</td>
<td>0.11</td>
<td>0.89</td>
<td>0.80</td>
<td>0.87</td>
<td>0.72</td>
<td>0.79</td>
<td>0.63</td>
<td>0.61</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-inst &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.70</td>
<td>0.33</td>
<td>0.67</td>
<td>0.69</td>
<td>0.68</td>
<td>0.70</td>
<td>0.69</td>
<td>0.37</td>
<td>0.37</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.89</td>
<td>0.11</td>
<td>0.56</td>
<td>0.53</td>
<td>1.00</td>
<td>0.69</td>
<td>0.53</td>
<td>0.11</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.89</td>
<td>0.11</td>
<td>0.56</td>
<td>0.53</td>
<td>1.00</td>
<td>0.69</td>
<td>0.53</td>
<td>0.11</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.42</td>
<td>0.38</td>
<td>0.62</td>
<td>0.52</td>
<td>0.52</td>
<td>0.42</td>
<td>0.47</td>
<td>0.04</td>
<td>0.04</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.37</td>
<td>0.39</td>
<td>0.61</td>
<td>0.49</td>
<td>0.49</td>
<td>0.37</td>
<td>0.42</td>
<td>&#x2212;0.02</td>
<td>&#x2212;0.02</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.30</td>
<td>0.22</td>
<td>0.78</td>
<td>0.54</td>
<td>0.58</td>
<td>0.30</td>
<td>0.40</td>
<td>0.11</td>
<td>0.09</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.29</td>
<td>0.26</td>
<td>0.74</td>
<td>0.52</td>
<td>0.54</td>
<td>0.29</td>
<td>0.38</td>
<td>0.05</td>
<td>0.04</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.28</td>
<td>0.28</td>
<td>0.72</td>
<td>0.50</td>
<td>0.50</td>
<td>0.28</td>
<td>0.36</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.28</td>
<td>0.38</td>
<td>0.62</td>
<td>0.45</td>
<td>0.42</td>
<td>0.28</td>
<td>0.33</td>
<td>&#x2212;0.12</td>
<td>&#x2212;0.11</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.25</td>
<td>0.28</td>
<td>0.72</td>
<td>0.49</td>
<td>0.48</td>
<td>0.25</td>
<td>0.33</td>
<td>&#x2212;0.03</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.23</td>
<td>0.34</td>
<td>0.66</td>
<td>0.44</td>
<td>0.40</td>
<td>0.23</td>
<td>0.29</td>
<td>&#x2212;0.14</td>
<td>&#x2212;0.11</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.12</td>
<td>0.21</td>
<td>0.79</td>
<td>0.45</td>
<td>0.37</td>
<td>0.12</td>
<td>0.18</td>
<td>&#x2212;0.16</td>
<td>&#x2212;0.09</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.17</td>
<td>0.83</td>
<td>0.46</td>
<td>0.36</td>
<td>0.10</td>
<td>0.15</td>
<td>&#x2212;0.16</td>
<td>&#x2212;0.07</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.08</td>
<td>0.13</td>
<td>0.87</td>
<td>0.47</td>
<td>0.38</td>
<td>0.08</td>
<td>0.13</td>
<td>&#x2212;0.13</td>
<td>&#x2212;0.05</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-13fn1" fn-type="other">
<p>Note: Table columns: Tool, Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>Insecure Cookie</bold></p>
<p>In the <xref ref-type="fig" rid="fig-9">Fig. 9a</xref> bar chart summarizing the results obtained for Insecure Cookie across the four scenarios (Business critical, Non critical, Best effort, minimum effort) in contrast, <xref ref-type="table" rid="table-14">Table 14</xref> ranked by F1-S presents the overall metrics. FindSecBugs v1.4.5 was the best choice in the business-critical scenario, among all the other scenarios providing strong precision and balance between minimum effort, best effort and non critical systems.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Insecure cookie results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-9.tif"/>
</fig><table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Insecure cookie results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="15mm"/>
<col align="center" width="12mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>Google</td>
<td>0.89</td>
<td>0.00</td>
<td>1.00</td>
<td>0.94</td>
<td>1.00</td>
<td>0.89</td>
<td>0.94</td>
<td>0.89</td>
<td>0.89</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.83</td>
<td>0.35</td>
<td>0.65</td>
<td>0.75</td>
<td>0.73</td>
<td>0.83</td>
<td>0.78</td>
<td>0.50</td>
<td>0.48</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.75</td>
<td>0.35</td>
<td>0.65</td>
<td>0.70</td>
<td>0.71</td>
<td>0.75</td>
<td>0.73</td>
<td>0.40</td>
<td>0.40</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.31</td>
<td>0.52</td>
<td>0.48</td>
<td>0.39</td>
<td>0.41</td>
<td>0.31</td>
<td>0.35</td>
<td>&#x2212;0.22</td>
<td>&#x2212;0.21</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.19</td>
<td>0.42</td>
<td>0.58</td>
<td>0.37</td>
<td>0.35</td>
<td>0.19</td>
<td>0.25</td>
<td>&#x2212;0.27</td>
<td>&#x2212;0.23</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.19</td>
<td>0.39</td>
<td>0.61</td>
<td>0.39</td>
<td>0.37</td>
<td>0.19</td>
<td>0.25</td>
<td>&#x2212;0.24</td>
<td>&#x2212;0.19</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.11</td>
<td>0.03</td>
<td>0.97</td>
<td>0.51</td>
<td>0.80</td>
<td>0.11</td>
<td>0.20</td>
<td>0.28</td>
<td>0.08</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.11</td>
<td>0.10</td>
<td>0.90</td>
<td>0.48</td>
<td>0.57</td>
<td>0.11</td>
<td>0.19</td>
<td>0.04</td>
<td>0.01</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.08</td>
<td>0.06</td>
<td>0.94</td>
<td>0.48</td>
<td>0.60</td>
<td>0.08</td>
<td>0.15</td>
<td>0.07</td>
<td>0.02</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.08</td>
<td>0.39</td>
<td>0.61</td>
<td>0.33</td>
<td>0.20</td>
<td>0.08</td>
<td>0.12</td>
<td>&#x2212;0.43</td>
<td>&#x2212;0.30</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.03</td>
<td>0.19</td>
<td>0.81</td>
<td>0.39</td>
<td>0.14</td>
<td>0.03</td>
<td>0.05</td>
<td>&#x2212;0.44</td>
<td>&#x2212;0.17</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.03</td>
<td>0.13</td>
<td>0.87</td>
<td>0.42</td>
<td>0.20</td>
<td>0.03</td>
<td>0.05</td>
<td>&#x2212;0.36</td>
<td>&#x2212;0.10</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.03</td>
<td>0.35</td>
<td>0.65</td>
<td>0.31</td>
<td>0.08</td>
<td>0.03</td>
<td>0.04</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.33</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.03</td>
<td>0.35</td>
<td>0.65</td>
<td>0.31</td>
<td>0.08</td>
<td>0.03</td>
<td>0.04</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.33</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.06</td>
<td>0.94</td>
<td>0.43</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.06</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.00</td>
<td>0.90</td>
<td>0.10</td>
<td>0.04</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.92</td>
<td>&#x2212;0.90</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.06</td>
<td>0.94</td>
<td>0.43</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.06</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.45</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.45</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.45</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.06</td>
<td>0.94</td>
<td>0.43</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.55</td>
<td>&#x2212;0.06</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-14fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>LDAP Injection</bold></p>
<p>In the <xref ref-type="fig" rid="fig-10">Fig. 10a</xref> bar chart summarizing the results obtained for LDAP Injection across the four scenarios (Business critical, Non critical, Best effort, minimum effort) in contrast, <xref ref-type="table" rid="table-15">Table 15</xref> ranked by F1-S presents the overall metrics. FindSecBugs v1.4.5 was the best performer in the business-critical scenario, where maximizing Recall is essential to avoid missing real vulnerabilities in high-risk systems. In contrast, Gemini 2.0 combined with SAST achieved superior results in the non-critical, best-effort, and minimum-effort scenarios, offering a stronger balance between detection accuracy and reduced manual effort, as aligned with the respective priorities of Informedness, F1-score, and Markedness.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>LDAP injection results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-10.tif"/>
</fig><table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>LDAP injection results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="15mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.96</td>
<td>0.00</td>
<td>1.00</td>
<td>0.98</td>
<td>1.00</td>
<td>0.96</td>
<td>0.98</td>
<td>0.97</td>
<td>0.96</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>1.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.98</td>
<td>0.96</td>
<td>1.00</td>
<td>0.98</td>
<td>0.96</td>
<td>0.97</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.85</td>
<td>0.06</td>
<td>0.94</td>
<td>0.90</td>
<td>0.92</td>
<td>0.85</td>
<td>0.88</td>
<td>0.80</td>
<td>0.79</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.85</td>
<td>0.13</td>
<td>0.88</td>
<td>0.86</td>
<td>0.85</td>
<td>0.85</td>
<td>0.85</td>
<td>0.73</td>
<td>0.73</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.78</td>
<td>0.06</td>
<td>0.94</td>
<td>0.86</td>
<td>0.91</td>
<td>0.78</td>
<td>0.84</td>
<td>0.75</td>
<td>0.72</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-inst &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.93</td>
<td>0.25</td>
<td>0.75</td>
<td>0.51</td>
<td>0.76</td>
<td>0.93</td>
<td>0.83</td>
<td>0.68</td>
<td>0.68</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.84</td>
<td>0.16</td>
<td>0.54</td>
<td>0.50</td>
<td>1.00</td>
<td>0.67</td>
<td>0.50</td>
<td>0.16</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.84</td>
<td>0.16</td>
<td>0.54</td>
<td>0.50</td>
<td>1.00</td>
<td>0.67</td>
<td>0.50</td>
<td>0.16</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.15</td>
<td>0.00</td>
<td>1.00</td>
<td>0.61</td>
<td>1.00</td>
<td>0.15</td>
<td>0.26</td>
<td>0.58</td>
<td>0.15</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.11</td>
<td>0.00</td>
<td>1.00</td>
<td>0.56</td>
<td>1.00</td>
<td>0.11</td>
<td>0.20</td>
<td>0.57</td>
<td>0.11</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.07</td>
<td>0.03</td>
<td>0.97</td>
<td>0.56</td>
<td>0.67</td>
<td>0.07</td>
<td>0.13</td>
<td>0.22</td>
<td>0.04</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.07</td>
<td>0.03</td>
<td>0.97</td>
<td>0.53</td>
<td>0.67</td>
<td>0.07</td>
<td>0.13</td>
<td>0.22</td>
<td>0.04</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.04</td>
<td>0.00</td>
<td>1.00</td>
<td>0.56</td>
<td>1.00</td>
<td>0.04</td>
<td>0.07</td>
<td>0.55</td>
<td>0.04</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.04</td>
<td>0.03</td>
<td>0.97</td>
<td>0.54</td>
<td>0.50</td>
<td>0.04</td>
<td>0.07</td>
<td>0.04</td>
<td>0.01</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.04</td>
<td>0.00</td>
<td>1.00</td>
<td>0.56</td>
<td>1.00</td>
<td>0.04</td>
<td>0.07</td>
<td>0.55</td>
<td>0.04</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.04</td>
<td>0.00</td>
<td>1.00</td>
<td>0.59</td>
<td>1.00</td>
<td>0.04</td>
<td>0.07</td>
<td>0.55</td>
<td>0.04</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.06</td>
<td>0.94</td>
<td>0.53</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.47</td>
<td>&#x2212;0.06</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.83</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.47</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.06</td>
<td>0.94</td>
<td>0.51</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.47</td>
<td>&#x2212;0.06</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.53</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.47</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.56</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.47</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.53</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.47</td>
<td>&#x2212;0.03</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-15fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>Path Traversal</bold></p>
<p>In the <xref ref-type="fig" rid="fig-11">Fig. 11a</xref> bar chart summarizing the results obtained for Path Traversal across the four scenarios (Business critical, Non critical, Best effort, minimum effort) in contrast, <xref ref-type="table" rid="table-16">Table 16</xref> ranked by F1-S presents the overall metrics. Across all evaluated scenarios, Gemini v1.5 and Gemini v2.0 combined with SAST achieved superior performance compared to other LLM-based approaches, demonstrating strong adaptability and effectiveness across varying levels of security requirements.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Path traversal results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-11.tif"/>
</fig><table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Path traversal results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>1.00</td>
<td>0.01</td>
<td>0.99</td>
<td>1.00</td>
<td>0.99</td>
<td>1.00</td>
<td>1.00</td>
<td>0.99</td>
<td>0.99</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.95</td>
<td>0.03</td>
<td>0.97</td>
<td>0.96</td>
<td>0.97</td>
<td>0.95</td>
<td>0.96</td>
<td>0.93</td>
<td>0.93</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.95</td>
<td>0.02</td>
<td>0.98</td>
<td>0.96</td>
<td>0.98</td>
<td>0.95</td>
<td>0.96</td>
<td>0.93</td>
<td>0.93</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.96</td>
<td>0.06</td>
<td>0.94</td>
<td>0.95</td>
<td>0.94</td>
<td>0.96</td>
<td>0.95</td>
<td>0.90</td>
<td>0.90</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.95</td>
<td>0.07</td>
<td>0.93</td>
<td>0.94</td>
<td>0.93</td>
<td>0.95</td>
<td>0.94</td>
<td>0.87</td>
<td>0.87</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>0.96</td>
<td>0.87</td>
<td>0.13</td>
<td>0.54</td>
<td>0.52</td>
<td>0.96</td>
<td>0.68</td>
<td>0.31</td>
<td>0.10</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.96</td>
<td>0.04</td>
<td>0.52</td>
<td>0.51</td>
<td>1.00</td>
<td>0.67</td>
<td>0.51</td>
<td>0.04</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-inst &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.42</td>
<td>0.55</td>
<td>0.45</td>
<td>0.44</td>
<td>0.43</td>
<td>0.42</td>
<td>0.43</td>
<td>&#x2212;0.13</td>
<td>&#x2212;0.13</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.42</td>
<td>0.63</td>
<td>0.37</td>
<td>0.40</td>
<td>0.40</td>
<td>0.42</td>
<td>0.41</td>
<td>&#x2212;0.21</td>
<td>&#x2212;0.21</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.39</td>
<td>0.60</td>
<td>0.40</td>
<td>0.40</td>
<td>0.39</td>
<td>0.39</td>
<td>0.39</td>
<td>&#x2212;0.21</td>
<td>&#x2212;0.21</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.29</td>
<td>0.58</td>
<td>0.42</td>
<td>0.35</td>
<td>0.33</td>
<td>0.29</td>
<td>0.31</td>
<td>&#x2212;0.30</td>
<td>&#x2212;0.29</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.09</td>
<td>0.19</td>
<td>0.81</td>
<td>0.46</td>
<td>0.32</td>
<td>0.09</td>
<td>0.14</td>
<td>&#x2212;0.20</td>
<td>&#x2212;0.10</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.07</td>
<td>0.18</td>
<td>0.82</td>
<td>0.45</td>
<td>0.27</td>
<td>0.07</td>
<td>0.11</td>
<td>&#x2212;0.25</td>
<td>&#x2212;0.11</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.06</td>
<td>0.19</td>
<td>0.81</td>
<td>0.44</td>
<td>0.24</td>
<td>0.06</td>
<td>0.10</td>
<td>&#x2212;0.29</td>
<td>&#x2212;0.13</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.02</td>
<td>0.00</td>
<td>1.00</td>
<td>0.51</td>
<td>1.00</td>
<td>0.02</td>
<td>0.04</td>
<td>0.51</td>
<td>0.02</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.02</td>
<td>0.00</td>
<td>1.00</td>
<td>0.51</td>
<td>1.00</td>
<td>0.02</td>
<td>0.03</td>
<td>0.51</td>
<td>0.01</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.02</td>
<td>0.01</td>
<td>0.99</td>
<td>0.50</td>
<td>0.50</td>
<td>0.02</td>
<td>0.03</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.01</td>
<td>0.06</td>
<td>0.94</td>
<td>0.48</td>
<td>0.11</td>
<td>0.01</td>
<td>0.01</td>
<td>&#x2212;0.40</td>
<td>&#x2212;0.05</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.09</td>
<td>0.91</td>
<td>0.46</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.52</td>
<td>&#x2212;0.09</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.50</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.50</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.50</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-16fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>SQL Injection</bold></p>
<p>In the <xref ref-type="fig" rid="fig-12">Fig. 12a</xref> bar chart summarizing the results obtained for SQL Injection across the four scenarios (Business critical, Non critical, Best effort, minimum effort) in contrast, <xref ref-type="table" rid="table-17">Table 17</xref> ranked by F1-S presents the overall metrics. While Gemini v2.0 outperformed in most scenarios, FindSecBugs was more stable and effective in the business-critical scenario, where high Recall is the top priority.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>SQL injection results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-12.tif"/>
</fig><table-wrap id="table-17">
<label>Table 17</label>
<caption>
<title>SQL injection results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="12mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.99</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.99</td>
<td>1.00</td>
<td>0.99</td>
<td>0.99</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.99</td>
<td>0.01</td>
<td>0.99</td>
<td>0.99</td>
<td>0.99</td>
<td>0.99</td>
<td>0.99</td>
<td>0.98</td>
<td>0.98</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.99</td>
<td>0.00</td>
<td>1.00</td>
<td>0.99</td>
<td>1.00</td>
<td>0.99</td>
<td>0.99</td>
<td>0.99</td>
<td>0.99</td>
</tr>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.94</td>
<td>0.00</td>
<td>1.00</td>
<td>0.97</td>
<td>1.00</td>
<td>0.94</td>
<td>0.97</td>
<td>0.94</td>
<td>0.94</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.70</td>
<td>0.03</td>
<td>0.97</td>
<td>0.82</td>
<td>0.96</td>
<td>0.70</td>
<td>0.81</td>
<td>0.69</td>
<td>0.66</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.69</td>
<td>0.06</td>
<td>0.94</td>
<td>0.80</td>
<td>0.93</td>
<td>0.69</td>
<td>0.79</td>
<td>0.65</td>
<td>0.63</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-inst &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.76</td>
<td>0.24</td>
<td>0.76</td>
<td>0.76</td>
<td>0.79</td>
<td>0.76</td>
<td>0.78</td>
<td>0.52</td>
<td>0.52</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.68</td>
<td>0.08</td>
<td>0.92</td>
<td>0.79</td>
<td>0.91</td>
<td>0.68</td>
<td>0.78</td>
<td>0.62</td>
<td>0.60</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.91</td>
<td>0.09</td>
<td>0.58</td>
<td>0.56</td>
<td>1.00</td>
<td>0.72</td>
<td>0.56</td>
<td>0.09</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.91</td>
<td>0.09</td>
<td>0.58</td>
<td>0.56</td>
<td>1.00</td>
<td>0.72</td>
<td>0.56</td>
<td>0.09</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.26</td>
<td>0.22</td>
<td>0.78</td>
<td>0.50</td>
<td>0.58</td>
<td>0.26</td>
<td>0.36</td>
<td>0.06</td>
<td>0.04</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.22</td>
<td>0.24</td>
<td>0.76</td>
<td>0.47</td>
<td>0.52</td>
<td>0.22</td>
<td>0.31</td>
<td>&#x2212;0.03</td>
<td>&#x2212;0.02</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.18</td>
<td>0.08</td>
<td>0.92</td>
<td>0.52</td>
<td>0.72</td>
<td>0.18</td>
<td>0.29</td>
<td>0.21</td>
<td>0.10</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.21</td>
<td>0.24</td>
<td>0.76</td>
<td>0.46</td>
<td>0.50</td>
<td>0.21</td>
<td>0.29</td>
<td>&#x2212;0.05</td>
<td>&#x2212;0.04</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.14</td>
<td>0.07</td>
<td>0.93</td>
<td>0.50</td>
<td>0.69</td>
<td>0.14</td>
<td>0.23</td>
<td>0.17</td>
<td>0.07</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.14</td>
<td>0.12</td>
<td>0.88</td>
<td>0.48</td>
<td>0.58</td>
<td>0.14</td>
<td>0.22</td>
<td>0.04</td>
<td>0.02</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.13</td>
<td>0.15</td>
<td>0.85</td>
<td>0.46</td>
<td>0.51</td>
<td>0.13</td>
<td>0.21</td>
<td>&#x2212;0.04</td>
<td>&#x2212;0.02</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.13</td>
<td>0.16</td>
<td>0.84</td>
<td>0.46</td>
<td>0.49</td>
<td>0.13</td>
<td>0.20</td>
<td>&#x2212;0.06</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.07</td>
<td>0.93</td>
<td>0.48</td>
<td>0.62</td>
<td>0.10</td>
<td>0.17</td>
<td>0.09</td>
<td>0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.09</td>
<td>0.91</td>
<td>0.47</td>
<td>0.56</td>
<td>0.10</td>
<td>0.17</td>
<td>0.02</td>
<td>0.01</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.08</td>
<td>0.06</td>
<td>0.94</td>
<td>0.47</td>
<td>0.60</td>
<td>0.08</td>
<td>0.14</td>
<td>0.06</td>
<td>0.02</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-17fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>Trust Boundary</bold></p>
<p>In the <xref ref-type="fig" rid="fig-13">Fig. 13a</xref> bar chart summarizing the results obtained for Trust Boundary across the four scenarios (Business critical, Non critical, Best effort, minimum effort), <xref ref-type="table" rid="table-18">Table 18</xref> ranked by F1-S presents the overall metrics. FindSecBugs achieved the highest Recall overall, making it the most effective tool in scenarios where detecting all real vulnerabilities is critical. However, Gemini v2.0 with SAST integration achieved top Recall in three scenarios while also offering better Precision, making it a strong alternative where a balance between detection and false positives is important.</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Trust boundary results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-13.tif"/>
</fig><table-wrap id="table-18">
<label>Table 18</label>
<caption>
<title>Trust boundary results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.90</td>
<td>0.00</td>
<td>1.00</td>
<td>0.94</td>
<td>1.00</td>
<td>0.90</td>
<td>0.95</td>
<td>0.84</td>
<td>0.90</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.81</td>
<td>0.19</td>
<td>0.72</td>
<td>0.70</td>
<td><bold>1.00</bold></td>
<td>0.83</td>
<td>0.70</td>
<td>0.19</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.81</td>
<td>0.19</td>
<td>0.72</td>
<td>0.70</td>
<td>1.00</td>
<td>0.83</td>
<td>0.70</td>
<td>0.19</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-inst &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.57</td>
<td>0.47</td>
<td>0.53</td>
<td>0.56</td>
<td>0.70</td>
<td>0.57</td>
<td>0.63</td>
<td>0.09</td>
<td>0.10</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.11</td>
<td>0.00</td>
<td>1.00</td>
<td>0.41</td>
<td>1.00</td>
<td>0.11</td>
<td>0.20</td>
<td>0.37</td>
<td>0.11</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.07</td>
<td>0.93</td>
<td>0.38</td>
<td>0.73</td>
<td>0.10</td>
<td>0.17</td>
<td>0.08</td>
<td>0.03</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.05</td>
<td>0.95</td>
<td>0.39</td>
<td>0.80</td>
<td>0.10</td>
<td>0.17</td>
<td>0.15</td>
<td>0.05</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.04</td>
<td>0.35</td>
<td>0.65</td>
<td>0.25</td>
<td>0.17</td>
<td>0.04</td>
<td>0.06</td>
<td>&#x2212;0.57</td>
<td>&#x2212;0.31</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.02</td>
<td>0.05</td>
<td>0.95</td>
<td>0.34</td>
<td>0.50</td>
<td>0.02</td>
<td>0.05</td>
<td>&#x2212;0.16</td>
<td>&#x2212;0.02</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.02</td>
<td>0.07</td>
<td>0.93</td>
<td>0.33</td>
<td>0.40</td>
<td>0.02</td>
<td>0.05</td>
<td>&#x2212;0.27</td>
<td>&#x2212;0.05</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.01</td>
<td>0.00</td>
<td>1.00</td>
<td>0.35</td>
<td>1.00</td>
<td>0.01</td>
<td>0.02</td>
<td>0.34</td>
<td>0.01</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.34</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.05</td>
<td>0.95</td>
<td>0.33</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.67</td>
<td>&#x2212;0.05</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</table-wrap-foot>
</table-wrap>
 
<p><bold>Weak Encryption Algorithm</bold></p>
<p>In the <xref ref-type="fig" rid="fig-14">Fig. 14a</xref> bar chart summarizing the results obtained for Weak Encryption Algorithm across the four scenarios (Business critical, Non critical, Best effort, minimum effort). <xref ref-type="table" rid="table-19">Table 19</xref> ranked by F1-S presents the overall metrics. FindSecBugs v1.13 delivered the best results, closely followed by Gemini v2.0 combined with FindSecBugs. For the remaining scenarios (including non-critical, best-effort, and minimum-effort) FindSecBugs v1.13.0 achieved the highest overall performance, with Gemini v2.0 &#x002B; SAST integration consistently ranking second.</p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>Weak encryption algorithm results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-14.tif"/>
</fig><table-wrap id="table-19">
<label>Table 19</label>
<caption>
<title>Weak encryption algorithm results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.96</td>
<td>0.04</td>
<td>0.96</td>
<td>0.96</td>
<td>0.96</td>
<td>0.96</td>
<td>0.96</td>
<td>0.92</td>
<td>0.92</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.46</td>
<td>0.54</td>
<td>0.78</td>
<td>0.71</td>
<td>1.00</td>
<td>0.83</td>
<td>0.71</td>
<td>0.54</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.77</td>
<td>0.11</td>
<td>0.89</td>
<td>0.83</td>
<td>0.88</td>
<td>0.77</td>
<td>0.82</td>
<td>0.66</td>
<td>0.66</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instr &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.82</td>
<td>0.29</td>
<td>0.71</td>
<td>0.76</td>
<td>0.76</td>
<td>0.82</td>
<td>0.79</td>
<td>0.53</td>
<td>0.52</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.51</td>
<td>0.15</td>
<td>0.85</td>
<td>0.67</td>
<td>0.80</td>
<td>0.51</td>
<td>0.62</td>
<td>0.40</td>
<td>0.36</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.50</td>
<td>0.22</td>
<td>0.78</td>
<td>0.63</td>
<td>0.72</td>
<td>0.50</td>
<td>0.59</td>
<td>0.31</td>
<td>0.28</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.45</td>
<td>0.19</td>
<td>0.81</td>
<td>0.62</td>
<td>0.73</td>
<td>0.45</td>
<td>0.55</td>
<td>0.29</td>
<td>0.26</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.32</td>
<td>0.04</td>
<td>0.96</td>
<td>0.62</td>
<td>0.89</td>
<td>0.32</td>
<td>0.47</td>
<td>0.45</td>
<td>0.28</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.28</td>
<td>0.00</td>
<td>1.00</td>
<td>0.62</td>
<td>1.00</td>
<td>0.28</td>
<td>0.44</td>
<td>0.56</td>
<td>0.28</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.27</td>
<td>0.03</td>
<td>0.97</td>
<td>0.60</td>
<td>0.90</td>
<td>0.27</td>
<td>0.41</td>
<td>0.44</td>
<td>0.23</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.09</td>
<td>0.09</td>
<td>0.91</td>
<td>0.48</td>
<td>0.52</td>
<td>0.09</td>
<td>0.16</td>
<td>&#x2212;0.01</td>
<td>0.00</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.09</td>
<td>0.07</td>
<td>0.93</td>
<td>0.49</td>
<td>0.60</td>
<td>0.09</td>
<td>0.16</td>
<td>0.08</td>
<td>0.02</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.06</td>
<td>0.14</td>
<td>0.86</td>
<td>0.44</td>
<td>0.33</td>
<td>0.06</td>
<td>0.10</td>
<td>&#x2212;0.22</td>
<td>&#x2212;0.08</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.05</td>
<td>0.06</td>
<td>0.94</td>
<td>0.47</td>
<td>0.50</td>
<td>0.05</td>
<td>0.10</td>
<td>&#x2212;0.03</td>
<td>&#x2212;0.01</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.05</td>
<td>0.09</td>
<td>0.91</td>
<td>0.46</td>
<td>0.39</td>
<td>0.05</td>
<td>0.09</td>
<td>&#x2212;0.15</td>
<td>&#x2212;0.04</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.05</td>
<td>0.16</td>
<td>0.84</td>
<td>0.43</td>
<td>0.28</td>
<td>0.05</td>
<td>0.09</td>
<td>&#x2212;0.28</td>
<td>&#x2212;0.10</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.05</td>
<td>0.12</td>
<td>0.88</td>
<td>0.44</td>
<td>0.30</td>
<td>0.05</td>
<td>0.08</td>
<td>&#x2212;0.25</td>
<td>&#x2212;0.07</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.04</td>
<td>0.10</td>
<td>0.90</td>
<td>0.44</td>
<td>0.29</td>
<td>0.04</td>
<td>0.07</td>
<td>&#x2212;0.25</td>
<td>&#x2212;0.06</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.02</td>
<td>0.00</td>
<td>1.00</td>
<td>0.48</td>
<td>1.00</td>
<td>0.02</td>
<td>0.05</td>
<td>0.48</td>
<td>0.02</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.01</td>
<td>0.00</td>
<td>1.00</td>
<td>0.48</td>
<td>1.00</td>
<td>0.01</td>
<td>0.02</td>
<td>0.47</td>
<td>0.01</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.01</td>
<td>0.99</td>
<td>0.47</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.53</td>
<td>&#x2212;0.01</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.03</td>
<td>0.97</td>
<td>0.46</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.53</td>
<td>&#x2212;0.03</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-19fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
 
<p><bold>Weak Hashing Algorithm</bold></p>
<p>In the <xref ref-type="fig" rid="fig-15">Fig. 15a</xref> bar chart summarizing the results obtained for Weak Hash Algorithm across the four scenarios (Business critical, Non critical, Best effort, minimum effort). <xref ref-type="table" rid="table-20">Table 20</xref> ranked by F1-S presents the overall metrics. Gemini v2.0 was the top performer in the business-critical scenario, closely followed by Gemini v2.0 combined with SAST. In all other scenarios, Gemini v2.0 consistently delivered the best results.</p>
<fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Weak hashing algorithm results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-15.tif"/>
</fig><table-wrap id="table-20">
<label>Table 20</label>
<caption>
<title>Weak hashing algorithm results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.82</td>
<td>0.02</td>
<td>0.98</td>
<td>0.89</td>
<td>0.98</td>
<td>0.82</td>
<td>0.89</td>
<td>0.80</td>
<td>0.80</td>
</tr>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.82</td>
<td>0.07</td>
<td>0.93</td>
<td>0.87</td>
<td>0.93</td>
<td>0.82</td>
<td>0.87</td>
<td>0.74</td>
<td>0.75</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>0.69</td>
<td>0.00</td>
<td>1.00</td>
<td>0.83</td>
<td>1.00</td>
<td>0.69</td>
<td>0.82</td>
<td>0.73</td>
<td>0.69</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>0.69</td>
<td>0.00</td>
<td>1.00</td>
<td>0.83</td>
<td>1.00</td>
<td>0.69</td>
<td>0.82</td>
<td>0.73</td>
<td>0.69</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instr &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.55</td>
<td>0.40</td>
<td>0.60</td>
<td>0.57</td>
<td>0.62</td>
<td>0.55</td>
<td>0.58</td>
<td>0.15</td>
<td>0.15</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.29</td>
<td>0.03</td>
<td>0.97</td>
<td>0.60</td>
<td>0.93</td>
<td>0.29</td>
<td>0.45</td>
<td>0.46</td>
<td>0.27</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.20</td>
<td>0.03</td>
<td>0.97</td>
<td>0.55</td>
<td>0.90</td>
<td>0.20</td>
<td>0.33</td>
<td>0.40</td>
<td>0.17</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.24</td>
<td>0.33</td>
<td>0.67</td>
<td>0.44</td>
<td>0.47</td>
<td>0.24</td>
<td>0.32</td>
<td>&#x2212;0.11</td>
<td>&#x2212;0.09</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.17</td>
<td>0.04</td>
<td>0.96</td>
<td>0.53</td>
<td>0.85</td>
<td>0.17</td>
<td>0.28</td>
<td>0.34</td>
<td>0.13</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.16</td>
<td>0.07</td>
<td>0.93</td>
<td>0.51</td>
<td>0.72</td>
<td>0.16</td>
<td>0.27</td>
<td>0.20</td>
<td>0.09</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.16</td>
<td>0.13</td>
<td>0.87</td>
<td>0.48</td>
<td>0.60</td>
<td>0.16</td>
<td>0.26</td>
<td>0.06</td>
<td>0.03</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.16</td>
<td>0.25</td>
<td>0.75</td>
<td>0.43</td>
<td>0.44</td>
<td>0.16</td>
<td>0.24</td>
<td>&#x2212;0.14</td>
<td>&#x2212;0.09</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.15</td>
<td>0.08</td>
<td>0.92</td>
<td>0.50</td>
<td>0.68</td>
<td>0.15</td>
<td>0.24</td>
<td>0.15</td>
<td>0.06</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.15</td>
<td>0.19</td>
<td>0.81</td>
<td>0.45</td>
<td>0.49</td>
<td>0.15</td>
<td>0.23</td>
<td>&#x2212;0.07</td>
<td>&#x2212;0.04</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.16</td>
<td>0.35</td>
<td>0.65</td>
<td>0.39</td>
<td>0.36</td>
<td>0.16</td>
<td>0.22</td>
<td>&#x2212;0.24</td>
<td>&#x2212;0.18</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.12</td>
<td>0.07</td>
<td>0.93</td>
<td>0.49</td>
<td>0.70</td>
<td>0.12</td>
<td>0.21</td>
<td>0.17</td>
<td>0.06</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.12</td>
<td>0.19</td>
<td>0.81</td>
<td>0.44</td>
<td>0.44</td>
<td>0.12</td>
<td>0.19</td>
<td>&#x2212;0.12</td>
<td>&#x2212;0.06</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.12</td>
<td>0.22</td>
<td>0.79</td>
<td>0.42</td>
<td>0.41</td>
<td>0.12</td>
<td>0.19</td>
<td>&#x2212;0.16</td>
<td>&#x2212;0.09</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.17</td>
<td>0.83</td>
<td>0.43</td>
<td>0.42</td>
<td>0.10</td>
<td>0.16</td>
<td>&#x2212;0.15</td>
<td>&#x2212;0.07</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.05</td>
<td>0.07</td>
<td>0.93</td>
<td>0.44</td>
<td>0.43</td>
<td>0.05</td>
<td>0.08</td>
<td>&#x2212;0.13</td>
<td>&#x2212;0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.04</td>
<td>0.05</td>
<td>0.95</td>
<td>0.45</td>
<td>0.50</td>
<td>0.04</td>
<td>0.07</td>
<td>&#x2212;0.05</td>
<td>&#x2212;0.01</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.03</td>
<td>0.00</td>
<td>1.00</td>
<td>0.47</td>
<td>1.00</td>
<td>0.03</td>
<td>0.06</td>
<td>0.46</td>
<td>0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.01</td>
<td>0.05</td>
<td>0.95</td>
<td>0.44</td>
<td>0.17</td>
<td>0.01</td>
<td>0.01</td>
<td>&#x2212;0.39</td>
<td>&#x2212;0.04</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-20fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
 
<p><bold>Weak Randomness</bold></p>
<p>In the <xref ref-type="fig" rid="fig-16">Fig. 16a</xref> bar chart summarizing the results obtained for Weak Randomness across the four scenarios (Business critical, Non critical, Best effort, minimum effort). <xref ref-type="table" rid="table-21">Table 21</xref> ranked by F1-S presents the overall metrics. FindSecBugs (both versions) consistently outperformed all other tools across all scenarios, demonstrating strong and reliable detection regardless of the use case.</p>
<fig id="fig-16">
<label>Figure 16</label>
<caption>
<title>Weak randomness results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-16.tif"/>
</fig><table-wrap id="table-21">
<label>Table 21</label>
<caption>
<title>Weak randomness results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.94</td>
<td>0.02</td>
<td>0.98</td>
<td>0.96</td>
<td>0.97</td>
<td>0.94</td>
<td>0.96</td>
<td>0.93</td>
<td>0.92</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.91</td>
<td>0.08</td>
<td>0.92</td>
<td>0.92</td>
<td>0.90</td>
<td>0.91</td>
<td>0.91</td>
<td>0.83</td>
<td>0.83</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.99</td>
<td>0.17</td>
<td>0.83</td>
<td>0.90</td>
<td>0.82</td>
<td>0.99</td>
<td>0.90</td>
<td>0.81</td>
<td>0.82</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.80</td>
<td>0.10</td>
<td>0.90</td>
<td>0.86</td>
<td>0.86</td>
<td>0.80</td>
<td>0.83</td>
<td>0.71</td>
<td>0.70</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.78</td>
<td>0.11</td>
<td>0.89</td>
<td>0.84</td>
<td>0.85</td>
<td>0.78</td>
<td>0.81</td>
<td>0.68</td>
<td>0.67</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.76</td>
<td>0.13</td>
<td>0.87</td>
<td>0.82</td>
<td>0.82</td>
<td>0.76</td>
<td>0.79</td>
<td>0.64</td>
<td>0.63</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instr &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.56</td>
<td>0.40</td>
<td>0.60</td>
<td>0.58</td>
<td>0.52</td>
<td>0.56</td>
<td>0.54</td>
<td>0.15</td>
<td>0.16</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.24</td>
<td>0.17</td>
<td>0.83</td>
<td>0.57</td>
<td>0.53</td>
<td>0.24</td>
<td>0.33</td>
<td>0.11</td>
<td>0.07</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.23</td>
<td>0.13</td>
<td>0.87</td>
<td>0.59</td>
<td>0.59</td>
<td>0.23</td>
<td>0.33</td>
<td>0.18</td>
<td>0.10</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.25</td>
<td>0.21</td>
<td>0.79</td>
<td>0.55</td>
<td>0.48</td>
<td>0.25</td>
<td>0.33</td>
<td>0.05</td>
<td>0.03</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.28</td>
<td>0.39</td>
<td>0.61</td>
<td>0.46</td>
<td>0.36</td>
<td>0.28</td>
<td>0.32</td>
<td>&#x2212;0.12</td>
<td>&#x2212;0.11</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.28</td>
<td>0.37</td>
<td>0.63</td>
<td>0.47</td>
<td>0.37</td>
<td>0.28</td>
<td>0.32</td>
<td>&#x2212;0.10</td>
<td>&#x2212;0.09</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.21</td>
<td>0.09</td>
<td>0.91</td>
<td>0.60</td>
<td>0.65</td>
<td>0.21</td>
<td>0.31</td>
<td>0.24</td>
<td>0.12</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.19</td>
<td>0.06</td>
<td>0.94</td>
<td>0.61</td>
<td>0.71</td>
<td>0.19</td>
<td>0.30</td>
<td>0.30</td>
<td>0.13</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.15</td>
<td>0.08</td>
<td>0.92</td>
<td>0.58</td>
<td>0.60</td>
<td>0.15</td>
<td>0.24</td>
<td>0.18</td>
<td>0.07</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.12</td>
<td>0.11</td>
<td>0.89</td>
<td>0.55</td>
<td>0.46</td>
<td>0.12</td>
<td>0.19</td>
<td>0.02</td>
<td>0.01</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.11</td>
<td>0.09</td>
<td>0.91</td>
<td>0.56</td>
<td>0.51</td>
<td>0.11</td>
<td>0.19</td>
<td>0.08</td>
<td>0.03</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.11</td>
<td>0.07</td>
<td>0.93</td>
<td>0.57</td>
<td>0.57</td>
<td>0.11</td>
<td>0.19</td>
<td>0.14</td>
<td>0.05</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.11</td>
<td>0.89</td>
<td>0.54</td>
<td>0.40</td>
<td>0.10</td>
<td>0.16</td>
<td>&#x2212;0.04</td>
<td>&#x2212;0.02</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.09</td>
<td>0.13</td>
<td>0.87</td>
<td>0.52</td>
<td>0.35</td>
<td>0.09</td>
<td>0.14</td>
<td>&#x2212;0.11</td>
<td>&#x2212;0.04</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.07</td>
<td>0.07</td>
<td>0.93</td>
<td>0.55</td>
<td>0.46</td>
<td>0.07</td>
<td>0.13</td>
<td>0.02</td>
<td>0.00</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-21fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>XPath Injection</bold></p>
<p>In the <xref ref-type="fig" rid="fig-17">Fig. 17A</xref> bar chart summarizing the results obtained for XPath Injection across the four scenarios (Business critical, Non critical, Best effort, minimum effort). <xref ref-type="table" rid="table-22">Table 22</xref> ranked by F1-S presents the overall metrics. Both Gemini v2.0 and Qwen 2.5 demonstrated superior performance across all evaluated scenarios. This consistency suggests that these models possess a strong understanding of XPath-related security patterns, enabling effective detection regardless of the specific use case.</p>
<fig id="fig-17">
<label>Figure 17</label>
<caption>
<title>XPath injection results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-17.tif"/>
</fig><table-wrap id="table-22">
<label>Table 22</label>
<caption>
<title>XPath injection results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>1.00</td>
<td>0.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.95</td>
<td>0.05</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>1.00</td>
<td>0.05</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>1.00</td>
<td>0.05</td>
<td>0.95</td>
<td>0.97</td>
<td>0.94</td>
<td>1.00</td>
<td>0.97</td>
<td>0.94</td>
<td>0.95</td>
</tr>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.87</td>
<td>0.00</td>
<td>1.00</td>
<td>0.94</td>
<td>1.00</td>
<td>0.87</td>
<td>0.93</td>
<td>0.91</td>
<td>0.87</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.93</td>
<td>0.10</td>
<td>0.90</td>
<td>0.91</td>
<td>0.88</td>
<td>0.93</td>
<td>0.90</td>
<td>0.82</td>
<td>0.83</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instr &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.80</td>
<td>0.25</td>
<td>0.75</td>
<td>0.77</td>
<td>0.71</td>
<td>0.80</td>
<td>0.75</td>
<td>0.54</td>
<td>0.55</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.87</td>
<td>0.45</td>
<td>0.55</td>
<td>0.69</td>
<td>0.59</td>
<td>0.87</td>
<td>0.70</td>
<td>0.44</td>
<td>0.42</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.95</td>
<td>0.05</td>
<td>0.46</td>
<td>0.44</td>
<td>1.00</td>
<td>0.61</td>
<td>0.44</td>
<td>0.05</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.47</td>
<td>0.25</td>
<td>0.75</td>
<td>0.63</td>
<td>0.58</td>
<td>0.47</td>
<td>0.52</td>
<td>0.24</td>
<td>0.22</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.20</td>
<td>0.15</td>
<td>0.85</td>
<td>0.57</td>
<td>0.50</td>
<td>0.20</td>
<td>0.29</td>
<td>0.09</td>
<td>0.05</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.20</td>
<td>0.20</td>
<td>0.80</td>
<td>0.54</td>
<td>0.43</td>
<td>0.20</td>
<td>0.27</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.20</td>
<td>0.20</td>
<td>0.80</td>
<td>0.54</td>
<td>0.43</td>
<td>0.20</td>
<td>0.27</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.20</td>
<td>0.25</td>
<td>0.75</td>
<td>0.51</td>
<td>0.38</td>
<td>0.20</td>
<td>0.26</td>
<td>&#x2212;0.07</td>
<td>&#x2212;0.05</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.13</td>
<td>0.05</td>
<td>0.95</td>
<td>0.60</td>
<td>0.67</td>
<td>0.13</td>
<td>0.22</td>
<td>0.26</td>
<td>0.08</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.07</td>
<td>0.00</td>
<td>1.00</td>
<td>0.60</td>
<td>1.00</td>
<td>0.07</td>
<td>0.13</td>
<td>0.59</td>
<td>0.07</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.15</td>
<td>0.85</td>
<td>0.49</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.47</td>
<td>&#x2212;0.15</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.05</td>
<td>0.95</td>
<td>0.54</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>&#x2212;0.44</td>
<td>&#x2212;0.05</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.57</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.57</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.57</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.57</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.00</td>
<td>0.00</td>
<td>1.00</td>
<td>0.57</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
<td>0.00</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-22fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>XSS (Cross-Site Scripting)</bold></p>
<p>In the <xref ref-type="fig" rid="fig-18">Fig. 18A</xref> bar chart summarizing the results obtained for XSS across the four scenarios (Business critical, Non critical, Best effort, minimum effort). <xref ref-type="table" rid="table-23">Table 23</xref> ranked by F1-S presents the overall metrics. Business-critical scenario, FindSecBugs delivered the highest Recall, making it the most effective at detecting all real issues when missing vulnerabilities is unacceptable. However, it exhibited lower Precision compared to Gemini v2.0 combined with SAST, which produced fewer false positives.</p>
<fig id="fig-18">
<label>Figure 18</label>
<caption>
<title>XSS results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-18.tif"/>
</fig><table-wrap id="table-23">
<label>Table 23</label>
<caption>
<title>XSS results</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="13mm"/>
<col align="center" width="13mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="5mm"/>
<col align="center" width="6mm"/>
<col align="center" width="6mm"/>
<col align="center" width="5mm"/>
<col align="center" width="7mm"/>
<col align="center" width="5mm"/> </colgroup>
<thead>
<tr>
<th>Tool</th>
<th>Type</th>
<th>Env.</th>
<th>TPR</th>
<th>FPR</th>
<th>TNR</th>
<th>Acc.</th>
<th>Pre.</th>
<th>Rec.</th>
<th>F1-S</th>
<th>MK</th>
<th>Inf.</th>
</tr>
</thead>
<tbody>
<tr>
<td>gemini-2-0-flash-001 &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1 &#x002B; G<sup>2</sup></td>
<td>0.98</td>
<td>0.19</td>
<td>0.81</td>
<td>0.90</td>
<td>0.86</td>
<td>0.98</td>
<td>0.92</td>
<td>0.84</td>
<td>0.79</td>
</tr>
<tr>
<td>gemini-2-0-flash-001</td>
<td>LLM</td>
<td>Google</td>
<td>0.90</td>
<td>0.24</td>
<td>0.76</td>
<td>0.84</td>
<td>0.82</td>
<td>0.90</td>
<td>0.86</td>
<td>0.69</td>
<td>0.66</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.89</td>
<td>0.22</td>
<td>0.78</td>
<td>0.84</td>
<td>0.83</td>
<td>0.89</td>
<td>0.86</td>
<td>0.69</td>
<td>0.68</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.86</td>
<td>0.19</td>
<td>0.81</td>
<td>0.84</td>
<td>0.84</td>
<td>0.86</td>
<td>0.85</td>
<td>0.68</td>
<td>0.68</td>
</tr>
<tr>
<td>qwen2-5-coder-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.88</td>
<td>0.21</td>
<td>0.79</td>
<td>0.84</td>
<td>0.83</td>
<td>0.88</td>
<td>0.85</td>
<td>0.68</td>
<td>0.67</td>
</tr>
<tr>
<td>SBwFindSecBugs v1.13.0</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.52</td>
<td>0.48</td>
<td>0.76</td>
<td>0.69</td>
<td>1.00</td>
<td>0.82</td>
<td>0.69</td>
<td>0.48</td>
</tr>
<tr>
<td>FBwFindSecBugs v1.4.6</td>
<td>SAST</td>
<td>M1</td>
<td>1.00</td>
<td>0.63</td>
<td>0.37</td>
<td>0.71</td>
<td>0.65</td>
<td>1.00</td>
<td>0.79</td>
<td>0.65</td>
<td>0.37</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instr &#x002B; S</td>
<td>S &#x002B; LLM<sup>1</sup></td>
<td>M1</td>
<td>0.79</td>
<td>0.30</td>
<td>0.70</td>
<td>0.75</td>
<td>0.76</td>
<td>0.79</td>
<td>0.78</td>
<td>0.50</td>
<td>0.50</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.60</td>
<td>0.06</td>
<td>0.94</td>
<td>0.76</td>
<td>0.93</td>
<td>0.60</td>
<td>0.73</td>
<td>0.59</td>
<td>0.54</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.60</td>
<td>0.07</td>
<td>0.93</td>
<td>0.75</td>
<td>0.91</td>
<td>0.60</td>
<td>0.72</td>
<td>0.57</td>
<td>0.53</td>
</tr>
<tr>
<td>phi3-3-8b-mini-128k-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.58</td>
<td>0.07</td>
<td>0.93</td>
<td>0.74</td>
<td>0.91</td>
<td>0.58</td>
<td>0.71</td>
<td>0.57</td>
<td>0.51</td>
</tr>
<tr>
<td>gemini-1-5-flash-002</td>
<td>LLM</td>
<td>Google</td>
<td>0.80</td>
<td>0.56</td>
<td>0.45</td>
<td>0.64</td>
<td>0.63</td>
<td>0.80</td>
<td>0.70</td>
<td>0.28</td>
<td>0.24</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.34</td>
<td>0.35</td>
<td>0.65</td>
<td>0.48</td>
<td>0.53</td>
<td>0.34</td>
<td>0.42</td>
<td>&#x2212;0.01</td>
<td>&#x2212;0.01</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.32</td>
<td>0.37</td>
<td>0.63</td>
<td>0.46</td>
<td>0.50</td>
<td>0.32</td>
<td>0.39</td>
<td>&#x2212;0.06</td>
<td>&#x2212;0.05</td>
</tr>
<tr>
<td>mistral-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.30</td>
<td>0.36</td>
<td>0.64</td>
<td>0.46</td>
<td>0.50</td>
<td>0.30</td>
<td>0.38</td>
<td>&#x2212;0.07</td>
<td>&#x2212;0.06</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.28</td>
<td>0.40</td>
<td>0.60</td>
<td>0.43</td>
<td>0.45</td>
<td>0.28</td>
<td>0.35</td>
<td>&#x2212;0.13</td>
<td>&#x2212;0.12</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.27</td>
<td>0.39</td>
<td>0.61</td>
<td>0.43</td>
<td>0.45</td>
<td>0.27</td>
<td>0.34</td>
<td>&#x2212;0.14</td>
<td>&#x2212;0.12</td>
</tr>
<tr>
<td>deepseek-coder-6-7b-instruct</td>
<td>LLM</td>
<td>M1</td>
<td>0.25</td>
<td>0.41</td>
<td>0.59</td>
<td>0.40</td>
<td>0.41</td>
<td>0.25</td>
<td>0.31</td>
<td>&#x2212;0.19</td>
<td>&#x2212;0.16</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.11</td>
<td>0.07</td>
<td>0.93</td>
<td>0.49</td>
<td>0.66</td>
<td>0.11</td>
<td>0.19</td>
<td>0.13</td>
<td>0.04</td>
</tr>
<tr>
<td>codellama-7b-instruct</td>
<td>LLM</td>
<td>runpod</td>
<td>0.10</td>
<td>0.05</td>
<td>0.95</td>
<td>0.49</td>
<td>0.71</td>
<td>0.10</td>
<td>0.18</td>
<td>0.19</td>
<td>0.05</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.06</td>
<td>0.04</td>
<td>0.96</td>
<td>0.47</td>
<td>0.61</td>
<td>0.06</td>
<td>0.10</td>
<td>0.07</td>
<td>0.01</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.04</td>
<td>0.05</td>
<td>0.95</td>
<td>0.45</td>
<td>0.45</td>
<td>0.04</td>
<td>0.07</td>
<td>&#x2212;0.09</td>
<td>&#x2212;0.02</td>
</tr>
<tr>
<td>starcoder2-7b</td>
<td>LLM</td>
<td>runpod</td>
<td>0.02</td>
<td>0.03</td>
<td>0.97</td>
<td>0.45</td>
<td>0.42</td>
<td>0.02</td>
<td>0.04</td>
<td>&#x2212;0.13</td>
<td>&#x2212;0.01</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-23fn1" fn-type="other">
<p>Note: Table columns: Tool,Type, Environment, TPR, FPR, TNR, Accuracy, Precision, Recall, F1 Score, Markedness, Informedness; <sup>1</sup>SAST &#x002B; LLM; <sup>2</sup>M1 &#x002B; Google; Main metric of Business critical, non-critical, best effort, minimum effort.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Analysis and Discussion</title>
<p>Technology is advancing rapidly, giving rise to a multitude of new technologies as well as new vulnerabilities. Cybersecurity must therefore seek an ally to keep pace with this rapid expansion. Is artificial intelligence truly a support for cybersecurity? What advantages and disadvantages does it bring?</p>
<p>In the context of the Secure Software Development Life Cycle (S-SDLC), one of the most critical activities is the review of source-code vulnerabilities, typically performed with Static Application Security Testing (SAST) tools. The principal drawback of these tools is the high rate of false positives. Consequently, it is essential to explore the potential of artificial intelligence in SAST. Leveraging AI could provide a substantial advantage by enabling earlier detection of vulnerabilities, thereby dramatically reducing cost and effort.</p>
<p>This study differs from other works mainly in its focus: it aims to evaluate the efficiency of large language models (LLMs), LLMs combined with SAST, or SAST alone on a specific programming language (Java), which is widely used in enterprise environments. The work explicitly addresses the removal of any external inferences, such as code references or comments that might alter the outcome itself. Additionally, it distinguishes itself in its evaluation methodology; it does not merxely concentrate on the best possible result but considers various real-world business contexts: Business-critical, Non-critical, Best-effort, and Minimum-effort [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>This work will address the following questions:</p>
<p><bold><italic>Do the Results of Large Language Models (LLMs) Exceed Those of Traditional Static Analysis Tools?</italic></bold></p>
<p>In the experiment, models with 7 B and 8 B parameters were used, along with an LLM containing more than 109 B parameters (Gemini). Although the performance of the models improves with increasing size, for &#x201C;Business-Critical&#x201D; software the best results were still achieved by the standalone SAST application. In contrast, for &#x201C;Non-Critical&#x201D;, &#x201C;Best-Effort&#x201D;, and &#x201C;Minimum-Effort&#x201D; scenarios, the combination of Gemini &#x002B; SAST yielded superior outcomes.</p>
<p>The vulnerabilities for which the traditional tool was outperformed in the &#x201C;Business-Critical&#x201D; category (<xref ref-type="fig" rid="fig-19">Fig. 19</xref>) are those related to Command Injection, Path Traversal, and SQL Injection; Gemini achieved higher precision than SAST in these cases. Conversely, in the &#x201C;Non-Critical&#x201D; (<xref ref-type="fig" rid="fig-20">Fig. 20</xref>), &#x201C;Best-Effort&#x201D; (<xref ref-type="fig" rid="fig-21">Fig. 21</xref>), and &#x201C;Minimum-Effort&#x201D; (<xref ref-type="fig" rid="fig-22">Fig. 22</xref>) scenarios, a significant advantage was observed for the LLM &#x002B; SAST combination.</p>
<fig id="fig-19">
<label>Figure 19</label>
<caption>
<title>Matrix of business critical results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-19.tif"/>
</fig><fig id="fig-20">
<label>Figure 20</label>
<caption>
<title>Matrix of non critical results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-20.tif"/>
</fig><fig id="fig-21">
<label>Figure 21</label>
<caption>
<title>Matrix of best effort results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-21.tif"/>
</fig><fig id="fig-22">
<label>Figure 22</label>
<caption>
<title>Matrix of minimum effort results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-22.tif"/>
</fig>
<p><bold><italic>Is There an Improvement When Combining LLMs with SAST Tools?</italic></bold></p>
<p><xref ref-type="fig" rid="fig-23">Fig. 23</xref> shows significant improvements are observed for many vulnerabilities, with a marked enhancement in trust-boundary issues. However, in some cases the combined approach worsens the results, such as for XPath Injection.</p>
<fig id="fig-23">
<label>Figure 23</label>
<caption>
<title>Comparison of local LLM and LLM as a service with SAST</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-23a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-23b.tif"/>
</fig>
<p>Thus, it can be concluded that the combination of LLMs and SAST yields a significant improvement, both for 7 B-parameter models and for 109 B-parameter models.</p>
<p><bold><italic>Are Reliable Results Obtained with Locally Hosted Models on a Typical Workstation?</italic></bold></p>
<p><xref ref-type="fig" rid="fig-24">Fig. 24</xref> shows that the 7-B-parameter Qwen 2.5 Coder achieves satisfactory performance (scores &#x003E; 0.8) on several vulnerability categories (including LDAP Injection, Path Traversal, and XPath Injection) across all tested scenarios (Business-Critical, Non-Critical, Best-Effort, and Minimum-Effort). However, for other types of vulnerability the performance is less encouraging. Consequently, the use of 7-B models is not recommended as a replacement for dedicated SAST tools.</p>
<fig id="fig-24">
<label>Figure 24</label>
<caption>
<title>Vulnerability comparison between LLMs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-24a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74566-fig-24b.tif"/>
</fig>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions and Future Work</title>
<p>We have worked to demonstrate the potential that language models (LLMs) can bring to cyber-security in the context of static code analysis (SAST), both as standalone solutions and in combination with other tools.</p>
<p>LLMs represent a technology with great promise for vulnerability detection and for reducing false positives. However, their behavior is not entirely predictable, as these models are trained on massive datasets and their output depends on the input context, without guaranteeing deterministic responses.</p>
<p>Despite their probabilistic nature, we have observed that it is possible to steer their outputs toward more deterministic behaviors through careful prompt engineering.</p>
<p>It would be valuable to explore models trained specifically on corpora focused on vulnerability detection. This could reduce model size, improve accuracy, and facilitate integration into resource-constrained enterprise environments. Because this field is rapidly evolving, we recommend in-depth research into LLM-based agents and techniques such as Retrieval-Augmented Generation (RAG) using MCP (Modal Context Protocol), which can enhance results without retraining the model, simply by augmenting its knowledge with external contextual information.</p>
<p>Among the main advantages of LLMs is their accuracy, sufficient to consider them a useful complement to traditional SAST. However, the primary limitations include the processing time required to analyze large volumes of code and the limited contextual capacity, which can hinder analysis of large classes or files.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Jos&#x00E9; Armando Santas Ciavatta; methodology, Jos&#x00E9; Armando Santas Ciavatta, Juan Ram&#x00F3;n Bermejo Higuera, software, Jos&#x00E9; Armando Santas Ciavatta; validation, Juan Ram&#x00F3;n Bermejo Higuera, Javier Bermejo Higuera, Juan Antonio Sicilia Montalvo, Tom&#x00E1;s Sureda Riera, Jes&#x00FA;s P&#x00E9;rez Melero; formal analysis, Juan Ram&#x00F3;n Bermejo Higuera, Javier Bermejo Higuera, Juan Antonio Sicilia Montalvo, Tom&#x00E1;s Sureda Riera, Jes&#x00FA;s P&#x00E9;rez Melero; investigation, Jos&#x00E9; Armando Santas Ciavatta, Juan Ram&#x00F3;n Bermejo Higuera; resources, Jos&#x00E9; Armando Santas Ciavatta, Juan Ram&#x00F3;n Bermejo Higuera; data curation, Jos&#x00E9; Armando Santas Ciavatta, Juan Ram&#x00F3;n Bermejo Higuera, Javier Bermejo Higuera, Juan Antonio Sicilia Montalvo, Tom&#x00E1;s Sureda Riera, Jes&#x00FA;s P&#x00E9;rez Melero; writing&#x2014;original draft preparation, Jos&#x00E9; Armando Santas Ciavatta; writing&#x2014;review and editing, Jos&#x00E9; Armando Santas Ciavatta, Juan Ram&#x00F3;n Bermejo Higuera; visualization, Juan Ram&#x00F3;n Bermejo Higuera, Javier Bermejo Higuera, Juan Antonio Sicilia Montalvo; supervision, Juan Ram&#x00F3;n Bermejo Higuera, Javier Bermejo Higuera, Juan Antonio Sicilia Montalvo, Tom&#x00E1;s Sureda Riera, Jes&#x00FA;s P&#x00E9;rez Melero; project administration, Juan Ram&#x00F3;n Bermejo Higuera, Javier Bermejo Higuera, Juan Antonio Sicilia Montalvo, Tom&#x00E1;s Sureda Riera, Jes&#x00FA;s P&#x00E9;rez Melero; funding acquisition, no funding. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data openly available in a public repository. The data that support the findings of this study are openly available in [Experiment-SAST-LLM-Tools] at <ext-link ext-link-type="uri" xlink:href="https://github.com/eltitopera/Experiment-SAST-LLM-Tools">https://github.com/eltitopera/Experiment-SAST-LLM-Tools</ext-link> (accessed on 20 November 2025).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Akhavani</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Ousat</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kharraz</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Open source, open threats? Investigating security challenges in open-source software</article-title>. <comment>arXiv:2506.12995. 2025</comment>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>CVEDetails</collab></person-group>. <article-title>CVEDetails&#x2014;2025: the most critical web application security risks [Internet]. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.cvedetails.com/vulnerabilities-by-types.php">https://www.cvedetails.com/vulnerabilities-by-types.php</ext-link>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Riskhan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ullah Sheikh</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Hossain</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zainol</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jhanjh</surname> <given-names>NZ</given-names></string-name></person-group>. <article-title>Major vulnerabilities of web application in real world scenarios and their prevention</article-title>. In: <conf-name>2025 International Conference on Intelligent and Cloud Computing (ICoICC); 2025 May 2&#x2013;3</conf-name>; <publisher-loc>Bhubaneswar, India. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icoicc64033.2025.11052016</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bettaieb</surname> <given-names>S</given-names></string-name>, <string-name><surname>Casino</surname> <given-names>F</given-names></string-name></person-group>. <article-title>A comprehensive analysis on software vulnerability detection datasets: trends, challenges, and road ahead</article-title>. <source>Int J Inf Secur</source>. <year>2024</year>;<volume>23</volume>(<issue>5</issue>):<fpage>3311</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10207-024-00888-y</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Gartner</collab></person-group>. <article-title>Worldwide IT spending forecast [Internet]. 2025 [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.gartner.com/en/newsroom/press-releases/2025-01-21-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2025">https://www.gartner.com/en/newsroom/press-releases/2025-01-21-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2025</ext-link>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Check Point Software</collab></person-group>. <article-title>Q1 2025 global cyber attack report [Internet]. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://blog.checkpoint.com/research/q1-2025-global-cyber-attack-report-from-check-point-software-an-almost-50-surge-in-cyber-threats-worldwide-with-a-rise-of-126-in-ransomware-attacks/">https://blog.checkpoint.com/research/q1-2025-global-cyber-attack-report-from-check-point-software-an-almost-50-surge-in-cyber-threats-worldwide-with-a-rise-of-126-in-ransomware-attacks/</ext-link>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Check Point Software</collab></person-group>. <article-title>Global cyber attacks surge 21% in Q2 2025: europe experiences the highest increase of all regions [Internet]. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://blog.checkpoint.com/research/global-cyber-attacks-surge-21-in-q2-2025-europe-experiences-the-highest-increase-of-all-regions/">https://blog.checkpoint.com/research/global-cyber-attacks-surge-21-in-q2-2025-europe-experiences-the-highest-increase-of-all-regions/</ext-link>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kiela</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Test scores of AI systems on various capabilities relative to human performance [Internet]. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://ourworldindata.org/grapher/test-scores-ai-capabilities-relative-human-performance?focus=Language&#x002B;understanding&#x007E;Predictive&#x002B;reasoning">https://ourworldindata.org/grapher/test-scores-ai-capabilities-relative-human-performance?focus=Language&#x002B;understanding&#x007E;Predictive&#x002B;reasoning</ext-link>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Singh</surname> <given-names>T</given-names></string-name></person-group>. <chapter-title>Artificial intelligence-driven cyberattacks</chapter-title>. In: <source>Cybersecurity, psychology and people hacking</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Palgrave Macmillan</publisher-name>; <year>2025</year>. p. <fpage>167</fpage>&#x2013;<lpage>88</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Ayodele</surname> <given-names>TO</given-names></string-name></person-group>. <chapter-title>Impact of AI-generated phishing attacks: a new cybersecurity threat</chapter-title>. In: <source>Intelligent computing</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>301</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-92605-1_19</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Harzevili</surname> <given-names>NS</given-names></string-name>, <string-name><surname>Belle</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>ZMJ</given-names></string-name>, <string-name><surname>Nagappan</surname> <given-names>N</given-names></string-name></person-group>. <article-title>A systematic literature review on automated software vulnerability detection using machine learning</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>57</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3699711</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>El Husseini</surname> <given-names>F</given-names></string-name>, <string-name><surname>Noura</surname> <given-names>H</given-names></string-name>, <string-name><surname>Salman</surname> <given-names>O</given-names></string-name>, <string-name><surname>Chehab</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Advanced machine learning approaches for zero-day attack detection: a review</article-title>. In: <conf-name>2024 8th Cyber Security in Networking Conference (CSNet); 2024 Dec 4&#x2013;6</conf-name>; <publisher-loc>Paris, France. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/csnet64211.2024.10851751</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Madupati</surname> <given-names>B</given-names></string-name></person-group>. <article-title>AI&#x2019;s impact on traditional software development</article-title>. <comment>arXiv:2502.18476. 2025</comment>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Santos</surname> <given-names>R</given-names></string-name>, <string-name><surname>Rizvi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cesarone</surname> <given-names>B</given-names></string-name>, <string-name><surname>Gunn</surname> <given-names>W</given-names></string-name>, <string-name><surname>McConnell</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Reducing software vulnerabilities using machine learning static application security testing</article-title>. In: <conf-name>2021 International Conference on Software Security and Assurance (ICSSA); 2021 Nov 10&#x2013;12</conf-name>; <publisher-loc>Altoona, PA, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>15</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icssa53632.2021.00016</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Odera</surname> <given-names>D</given-names></string-name>, <string-name><surname>Otieno</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ounza</surname> <given-names>JE</given-names></string-name></person-group>. <article-title>Security risks in the software development lifecycle: a review</article-title>. <source>World J Adv Eng Technol Sci</source>. <year>2023</year>;<volume>8</volume>(<issue>2</issue>):<fpage>230</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.30574/wjaets.2023.8.2.0101</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vald&#x00E9;s-Rodr&#x00ED;guez</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hochstetter-Diez</surname> <given-names>J</given-names></string-name>, <string-name><surname>Di&#x00E9;guez-Rebolledo</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bustamante-Mora</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cadena-Mart&#x00ED;nez</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Analysis of strategies for the integration of security practices in agile software development: a sustainable SME approach</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>8</issue>):<fpage>35204</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3372385</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vidyasagar</surname> <given-names>V</given-names></string-name></person-group>. <article-title>DevSecOps: integrating security into the DevOps lifecycle</article-title>. <source>Int J Mach Learn Res Cybersecur Artif Intell</source>. <year>2025</year>;<volume>16</volume>(<issue>1</issue>):<fpage>11</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.5281/zenodo.15483597</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Adriani</surname> <given-names>ZA</given-names></string-name>, <string-name><surname>Raharjo</surname> <given-names>T</given-names></string-name>, <string-name><surname>Trisnawaty</surname> <given-names>NW</given-names></string-name></person-group>. <article-title>Comprehensive examination of risk management practices throughout the software development life cycle (SDLC): a systematic literature review</article-title>. <source>Indones J Comput Sci</source>. <year>2024</year>;<volume>13</volume>(<issue>3</issue>):<fpage>3844</fpage>. doi:<pub-id pub-id-type="doi">10.33022/ijcs.v13i3.4016</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A comprehensive study on static application security testing (SAST) tools for Android</article-title>. <source>IEEE Trans Software Eng</source>. <year>2024</year>;<volume>50</volume>(<issue>12</issue>):<fpage>3385</fpage>&#x2013;<lpage>402</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tse.2024.3488041</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nguyen-Duc</surname> <given-names>A</given-names></string-name>, <string-name><surname>Do</surname> <given-names>MV</given-names></string-name>, <string-name><surname>Luong Hong</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Nguyen Khac</surname> <given-names>K</given-names></string-name>, <string-name><surname>Nguyen Quang</surname> <given-names>A</given-names></string-name></person-group>. <article-title>On the adoption of static analysis for software security assessment&#x2014;a case study of an open-source e-government project</article-title>. <source>Comput Secur</source>. <year>2021</year>;<volume>111</volume>(<issue>6</issue>):<fpage>102470</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2021.102470</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>B&#x00F6;hme</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bodden</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bultan</surname> <given-names>T</given-names></string-name>, <string-name><surname>Cadar</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Scanniello</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Software security analysis in 2030 and beyond: a research roadmap</article-title>. <source>ACM Trans Softw Eng Methodol</source>. <year>2025</year>;<volume>34</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3708533</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sarkar</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rakhra</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>V</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jairath</surname> <given-names>K</given-names></string-name>, <string-name><surname>Maan</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Comparing traditional vs agile methods for software development projects: a case study</article-title>. In: <conf-name>2024 7th International Conference on Contemporary Computing and Informatics (IC3I); 2024 Sep 18&#x2013;20</conf-name>; <publisher-loc>Greater Noida, India. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>47</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ic3i61595.2024.10829321</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>ITU</collab></person-group>. <article-title>Using the internet in 2024. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.itu.int/en/ITU-D/Statistics/pages/stat/default.aspx">https://www.itu.int/en/ITU-D/Statistics/pages/stat/default.aspx</ext-link>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>IBM</collab></person-group>. <article-title>Cost of a data breach report. 2025 [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.ibm.com/think/x-force/2025-cost-of-a-data-breach-navigating-ai">https://www.ibm.com/think/x-force/2025-cost-of-a-data-breach-navigating-ai</ext-link>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Antunes</surname> <given-names>N</given-names></string-name>, <string-name><surname>Vieira</surname> <given-names>M</given-names></string-name></person-group>. <article-title>On the metrics for benchmarking vulnerability detection tools</article-title>. In: <conf-name>2015 45th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN); 2015 Jun 22&#x2013;25</conf-name>; <publisher-loc>Rio de Janeiro, Brazil. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2015</year>. p. <fpage>283</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1109/DSN.2015.30</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>de Vicente Mohino</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bermejo Higuera</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bermejo Higuera</surname> <given-names>JR</given-names></string-name>, <string-name><surname>Sicilia Montalvo</surname> <given-names>JA</given-names></string-name></person-group>. <article-title>The application of a new secure software development life cycle (S-SDLC) with agile methodologies</article-title>. <source>Electronics</source>. <year>2019</year>;<volume>8</volume>(<issue>11</issue>):<fpage>1218</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics8111218</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Golovyrin</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Analysis of tools for static security testing of applications [Internet]. 2023 [cited 2025 Nov 19]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.researchgate.net/profile/Leonid-Golovyrin/publication/372134053_Analysis_of_tools_for_static_security_testing_of_applications/links/64a5dce2b9ed6874a5fc718a/Analysis-of-tools-for-static-security-testing-of-applications.pdf">https://www.researchgate.net/profile/Leonid-Golovyrin/publication/372134053_Analysis_of_tools_for_static_security_testing_of_applications/links/64a5dce2b9ed6874a5fc718a/Analysis-of-tools-for-static-security-testing-of-applications.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ferrara</surname> <given-names>P</given-names></string-name>, <string-name><surname>Olivieri</surname> <given-names>L</given-names></string-name>, <string-name><surname>Spoto</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Static privacy analysis by flow reconstruction of tainted data</article-title>. <source>Int J Soft Eng Knowl Eng</source>. <year>2021</year>;<volume>31</volume>(<issue>7</issue>):<fpage>973</fpage>&#x2013;<lpage>1016</lpage>. doi:<pub-id pub-id-type="doi">10.1142/s0218194021500303</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Benchmarking static analysis for PHP applications security</article-title>. <source>Entropy</source>. <year>2025</year>;<volume>27</volume>(<issue>9</issue>):<fpage>926</fpage>. doi:<pub-id pub-id-type="doi">10.3390/e27090926</pub-id>; <pub-id pub-id-type="pmid">41008052</pub-id></mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>D&#x00ED;az</surname> <given-names>G</given-names></string-name>, <string-name><surname>Bermejo</surname> <given-names>JR</given-names></string-name></person-group>. <article-title>Static analysis of source code security: assessment of tools against SAMATE tests</article-title>. <source>Inf Softw Technol</source>. <year>2013</year>;<volume>55</volume>(<issue>8</issue>):<fpage>1462</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.infsof.2013.02.005</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Understanding industry perspectives of static application security testing (SAST) evaluation</article-title>. <source>Proc ACM Softw Eng</source>. <year>2025</year>;<volume>2</volume>(<issue>FSE</issue>):<fpage>3033</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3729404</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Comparison and evaluation on static application security testing (SAST) tools for Java</article-title>. In: <conf-name>Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2023); 2023 Dec 3&#x2013;9</conf-name>; <publisher-loc>San Francisco, CA, USA. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery (ACM)</publisher-name>; <year>2023</year>. p. <fpage>921</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3611643.3616262</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ansgariusson</surname> <given-names>W</given-names></string-name>, <string-name><surname>St&#x00E5;hl</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Comparative analysis of static application security testing tools on real-world java vulnerabilities. 2025 [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://lup.lub.lu.se/student-papers/search/publication/9189955">https://lup.lub.lu.se/student-papers/search/publication/9189955</ext-link>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Higuera</surname> <given-names>B</given-names></string-name>, <string-name><surname>Higuera</surname> <given-names>JB</given-names></string-name>, <string-name><surname>Montalvo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Villalba</surname> <given-names>JC</given-names></string-name>, <string-name><surname>P&#x00E9;rez</surname> <given-names>JJN</given-names></string-name></person-group>. <article-title>Benchmarking approach to compare web applications static analysis tools detecting OWASP top ten security vulnerabilities</article-title>. <source>Comput Mater Contin</source>. <year>2020</year>;<volume>64</volume>(<issue>3</issue>):<fpage>1555</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2020.010885</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Sheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name></person-group>. <source>LLMs in software security: a survey of vulnerability detection techniques and insights</source>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>58</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3769082</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name></person-group>. <chapter-title>Automatic inspection of static application security testing (SAST) reports via large language model reasoning</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Barbosa</surname> <given-names>LS</given-names></string-name></person-group>, editors. <source>Artificial intelligence logic and applications</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>128</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-96-0354-1_11</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Keltek</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sani</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name></person-group>. <chapter-title>LSAST: enhancing cybersecurity through LLM-supported static application security testing</chapter-title>. In: <source>ICT systems security and privacy protection</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>166</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-92882-6_12</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>&#x00C7;etin</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ekmekcioglu</surname> <given-names>E</given-names></string-name>, <string-name><surname>Arief</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hernandez-Castro</surname> <given-names>J</given-names></string-name></person-group>. <article-title>An empirical evaluation of large language models in static code analysis for PHP vulnerability detection</article-title>. <source>J Univers Comput Sci</source>. <year>2024</year>;<volume>30</volume>(<issue>9</issue>):<fpage>1163</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.3897/jucs.134739</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Charoenwet</surname> <given-names>W</given-names></string-name>, <string-name><surname>Thongtanunam</surname> <given-names>P</given-names></string-name>, <string-name><surname>Pham</surname> <given-names>VT</given-names></string-name>, <string-name><surname>Treude</surname> <given-names>C</given-names></string-name></person-group>. <article-title>An empirical study of static analysis tools for secure code review</article-title>. In: <conf-name>Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2024); 2024 Jul 15&#x2013;19</conf-name>; <publisher-loc>Vienna, Austria. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery (ACM)</publisher-name>; <year>2024</year>. p. <fpage>691</fpage>&#x2013;<lpage>703</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3650212.3680313</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>NIST</collab></person-group>. <article-title>SARD acknowledgments and test suites descriptions. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.nist.gov/itl/ssd/software-quality-group/sard-acknowledgments-and-test-suites-descriptions">https://www.nist.gov/itl/ssd/software-quality-group/sard-acknowledgments-and-test-suites-descriptions</ext-link>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>OWASP</collab></person-group>. <article-title>Benchmark project. [cited 2025 Octo 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://owasp.org/www-project-benchmark/">https://owasp.org/www-project-benchmark/</ext-link>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Meta</collab></person-group>. <article-title>Introducing code llama. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/meta-llama/codellama">https://github.com/meta-llama/codellama</ext-link>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>DeepSeek</collab></person-group>. <article-title>DeepSeek coder. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://deepseekcoder.github.io/">https://deepseekcoder.github.io/</ext-link>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>PromptLayer</collab></person-group>. <article-title>Implementation details. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.promptlayer.com/models/deepseek-coder-67b-instruct">https://www.promptlayer.com/models/deepseek-coder-67b-instruct</ext-link>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Mistral</collab></person-group>. <article-title>Mistral 7b. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://mistral.ai/news/announcing-mistral-7b">https://mistral.ai/news/announcing-mistral-7b</ext-link>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Ollama</collab></person-group>. <article-title>Phi3. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://ollama.com/library/phi3">https://ollama.com/library/phi3</ext-link>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Ollama</collab></person-group>. <article-title>Qwen2.5-coder. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://ollama.com/library/qwen2.5-coder">https://ollama.com/library/qwen2.5-coder</ext-link>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Starcoder</collab></person-group>. <article-title>StarCoder 2. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/bigcode-project/starcoder2?tab=readme-ov-file">https://github.com/bigcode-project/starcoder2?tab=readme-ov-file</ext-link>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Google</collab></person-group>. <article-title>Gemini 2.0. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-0-flash">https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-0-flash</ext-link>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Google</collab></person-group>. <article-title>Gemini 1.5. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://developers.googleblog.com/en/gemini-15-flash-8b-is-now-generally-available-for-use/">https://developers.googleblog.com/en/gemini-15-flash-8b-is-now-generally-available-for-use/</ext-link>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Bigcode</collab></person-group>. <article-title>Big code models leaderboard. [cited 2025 Oct 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://huggingface.co/spaces/bigcode/bigcode-models-leaderboard">https://huggingface.co/spaces/bigcode/bigcode-models-leaderboard</ext-link>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Khare</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dutta</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Solko-Breslin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Alur</surname> <given-names>R</given-names></string-name>, <string-name><surname>Naik</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Understanding the effectiveness of large language models in detecting security vulnerabilities</article-title>. In: <conf-name>2025 IEEE Conference on Software Testing, Verification and Validation (ICST); 2025 Mar 31&#x2013;Apr 4</conf-name>; <publisher-loc>Napoli, Italy. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>312</fpage>&#x2013;<lpage>23</lpage>. </mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Das Purba</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ghosh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Radford</surname> <given-names>BJ</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Software vulnerability detection using large language models</article-title>. In: <conf-name>2023 IEEE 34th International Symposium on Software Reliability Engineering Workshops (ISSREW); 2023 Oct 9&#x2013;12</conf-name>; <publisher-loc>Florence, Italy. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>321</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ISSREW60843.2023.00058</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Patsakis</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Casino</surname> <given-names>F</given-names></string-name></person-group>. <chapter-title>Outside the comfort zone: analysing LLM capabilities in software vulnerability detection</chapter-title>. In: <source>Computer security&#x2014;ESORICS</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>271</fpage>&#x2013;<lpage>89</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-70879-4_14</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ni</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Multitask-based evaluation of open-source LLM on software vulnerability</article-title>. <source>IIEEE Trans Software Eng</source>. <year>2024</year>;<volume>50</volume>(<issue>11</issue>):<fpage>3071</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tse.2024.3470333</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shimmi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Saini</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Schaefer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Okhravi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Rahimi</surname> <given-names>M</given-names></string-name></person-group>. <chapter-title>Software vulnerability detection using LLM: does additional information help?</chapter-title> In: <article-title>2024 Annual Computer Security Applications Conference Workshops (ACSAC Workshops); 2024 Dec 9&#x2013;10</article-title>; <publisher-loc>Honolulu, HI, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>189</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/acsacw65225.2024.00031</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Almeida</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Prompt engineering: a comparative study of prompting techniques in AI language models</article-title>. In: <conf-name>2025 IEEE Integrated STEM Education Conference (ISEC); 2025 Mar 15</conf-name>; <publisher-loc>Princeton, NJ, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>87</fpage>&#x2013;<lpage>95</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ISEC64801.2025.11147384</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lo</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Large language model for vulnerability detection and repair: literature review and the road ahead</article-title>. <source>ACM Trans Softw Eng Methodol</source>. <year>2025</year>;<volume>34</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3708522</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sajadi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Le</surname> <given-names>B</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Damevski</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chatterjee</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Do LLMs consider security? An empirical study on responses to programming questions</article-title>. <source>Empir Softw Eng</source>. <year>2025</year>;<volume>30</volume>(<issue>4</issue>):<fpage>101</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10664-025-10658-6</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ayd&#x0131;n</surname> <given-names>D</given-names></string-name>, <string-name><surname>Bahtiyar</surname> <given-names>&#x015E;</given-names></string-name></person-group>. <article-title>Security vulnerabilities in AI-generated JavaScript: a comparative study of large language models</article-title>. In: <conf-name>2025 IEEE International Conference on Cyber Security and Resilience (CSR); 2025 Aug 4&#x2013;6</conf-name>; <publisher-loc>Chania, Crete, Greece. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>203</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1109/csr64739.2025.11130176</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jaoua</surname> <given-names>I</given-names></string-name>, <string-name><surname>Ben Sghaier</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sahraoui</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Combining large language models with static analyzers for code review generation</article-title>. In: <conf-name>2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR); 2025 Apr 28&#x2013;29</conf-name>; <publisher-loc>Ottawa, ON, Canada. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>157</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.1109/msr66628.2025.00038</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Munson</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>J</given-names></string-name>, <string-name><surname>C&#x00E1;rdenas</surname> <given-names>&#x00C1;A</given-names></string-name></person-group>. <article-title>With a little help from my (LLM) friends: enhancing static analysis with LLMs to detect software vulnerabilities</article-title>. In: <conf-name>2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code); 2025 May 3</conf-name>; <publisher-loc>Ottawa, ON, Canada. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>41</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/llm4code66737.2025.00008</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Effective vulnerable function identification based on CVE description empowered by large language models</article-title>. In: <conf-name>Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering (ASE 2024); 2024 Oct 27</conf-name>; <publisher-loc>Sacramento, CA, USA. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery (ACM)</publisher-name>; <year>2024</year>. p. <fpage>789</fpage>&#x2013;<lpage>99</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3691620.3695013</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hossain</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Mithun Kumar</surname> <given-names>PK</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Amsaad</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Malicious code detection using LLM</article-title>. In: <conf-name>NAECON 2024&#x2014;IEEE National Aerospace and Electronics Conference; 2024 Jul 15&#x2013;18</conf-name>; <publisher-loc>Dayton, OH, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>356</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1109/naecon61878.2024.10670668</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Blefari</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cosentino</surname> <given-names>C</given-names></string-name>, <string-name><surname>Furfaro</surname> <given-names>A</given-names></string-name>, <string-name><surname>Marozzo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Pironti</surname> <given-names>FA</given-names></string-name></person-group>. <article-title>SecFlow: an agentic LLM-based framework for modular cyberattack analysis and explainability</article-title>. In: <conf-name>Proceedings of the 2025 Generative Code Intelligence Workshop (GeCo IN); 2025 May 23</conf-name>; <publisher-loc>Bologna, Italy. Aachen, Germany</publisher-loc>: <publisher-name>CEUR Workshop Proceedings</publisher-name>; <year>2025</year>. p. <fpage>41</fpage>&#x2013;<lpage>58</lpage>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Belcastro</surname> <given-names>L</given-names></string-name>, <string-name><surname>Carlucci</surname> <given-names>C</given-names></string-name>, <string-name><surname>Cosentino</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li&#x00F2;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Marozzo</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Enhancing network security using knowledge graphs and large language models for explainable threat detection</article-title>. <source>Future Gener Comput Syst</source>. <year>2026</year>;<volume>176</volume>(<issue>7</issue>):<fpage>108160</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.future.2025.108160</pub-id>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Analysis of the effectiveness of large language model feature in source code defect detection</article-title>. In: <conf-name>2024 3rd International Conference on Artificial Intelligence and Computer Information Technology (AICIT); 2024 Sep 20&#x2013;22</conf-name>; <publisher-loc>Yichang, China. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>276</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1109/aicit62434.2024.10730232</pub-id>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Smaili</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mekkaoui</surname> <given-names>DE</given-names></string-name>, <string-name><surname>Midoun</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Talhaoui</surname> <given-names>MZ</given-names></string-name>, <string-name><surname>Hamidaoui</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A transformer-based framework for software vulnerability detection using attention-driven convolutional neural networks</article-title>. <source>Eng Appl Artif Intell</source>. <year>2025</year>;<volume>160</volume>(<issue>9</issue>):<fpage>111859</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2025.111859</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>
























