<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">77139</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.077139</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Predicting Software Security Bugs Using Machine Learning and Quality Metrics: An Empirical Study</article-title>
<alt-title alt-title-type="left-running-head">Predicting Software Security Bugs Using Machine Learning and Quality Metrics: An Empirical Study</alt-title>
<alt-title alt-title-type="right-running-head">Predicting Software Security Bugs Using Machine Learning and Quality Metrics: An Empirical Study</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Diouf</surname><given-names>Mohamed</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Toe</surname><given-names>Elis&#x00E9;e</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>elisee.toe1@uqac.ca</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Grichi</surname><given-names>Manel</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Nakouri</surname><given-names>Ha&#x00EF;fa</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Jaafar</surname><given-names>Fehmi</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science and Mathematics (DIM), University of Quebec at Chicoutimi (UQAC)</institution>, <addr-line>Chicoutimi, QC</addr-line>, <country>Canada</country></aff>
<aff id="aff-2"><label>2</label><institution>Research and Development Department, VibroSystM</institution>, <addr-line>Longueuil, QC</addr-line>, <country>Canada</country></aff>
<aff id="aff-3"><label>3</label><institution>LARODEC, Institut Sup&#x00E9;rieur de Gestion (ISG), University of Tunis</institution>, <addr-line>Tunis</addr-line>, <country>Tunisia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Elis&#x00E9;e Toe. Email: <email>elisee.toe1@uqac.ca</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>9</day><month>4</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>3</issue>
<elocation-id>12</elocation-id>
<history>
<date date-type="received">
<day>03</day>
<month>12</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>15</day>
<month>01</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_77139.pdf"></self-uri>
<abstract>
<p>Software security bugs present significant security risks to modern systems, leading to unauthorized access, data breaches, and severe operational and financial consequences. Early prediction of such vulnerabilities is therefore essential for strengthening software reliability and reducing remediation costs. This study investigates the extent to which static software quality metrics can identify vulnerable code and evaluates the effectiveness of machine learning models for large-scale security-bug prediction. We analyze a dataset of 338,442 source files, including 33,294 buggy files, collected from seven major open-source ecosystems. These ecosystems include GitHub Security Advisories (GHSA), Python Package Index (PyPI), OSS-Fuzz (Google&#x2019;s open-source fuzzing service), Node Package Manager (npm), Packagist (the PHP package repository), Apache Maven, and NuGet (the.NET package manager). Using the Open Source Vulnerabilities (OSV) platform, we identify 7685 confirmed security bugs and extract 25 static software quality metrics per file with the Understand analysis tool. We apply five complementary feature-importance techniques and evaluate eleven machine-learning classifiers under a time-series cross-validation protocol. Our analysis reveals three key findings. First, six core metrics consistently show strong associations with the presence of security bugs across all feature-selection methods. Second, buggy files exhibit substantially higher metric values, with medians approximately three times those of non-buggy files, a pattern we term the &#x201C;3<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> rule&#x201D;; Mann&#x2013;Whitney U tests confirm that these differences are statistically significant. Third, the machine-learning models achieve strong predictive performance, with XGBoost providing the best results (recall &#x003D; 0.82, precision &#x003D; 0.95, Receiver Operating Characteristic&#x2013;Area Under the Curve (ROC-AUC) &#x003D; 0.91). Based on these findings, we propose data-driven warning and critical thresholds for the most influential metrics to support proactive security assessment. Overall, this work provides large-scale empirical evidence that software quality metrics, combined with machine learning, offer actionable signals for detecting security bugs and for integrating automated vulnerability prediction into software development workflows.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Security bugs</kwd>
<kwd>vulnerability prediction</kwd>
<kwd>software metrics</kwd>
<kwd>machine learning</kwd>
<kwd>code complexity</kwd>
<kwd>software quality</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Natural Sciences and Engineering Research Council of Canada</funding-source>
<award-id>RGPIN-2019-05062</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Security vulnerabilities remain a critical threat to modern software systems, exposing organizations and users to risks such as data breaches, service disruption, and severe financial loss. As cyberattacks continue to increase in frequency and sophistication [<xref ref-type="bibr" rid="ref-1">1</xref>], the economic impact of cybercrime exceeded seven trillion USD in 2022<xref ref-type="fn" rid="fn-1"><sup>1</sup></xref><fn id="fn-1">
<label>1</label>
<p><ext-link ext-link-type="uri" xlink:href="https://cybersecurityventures.com/boardroom-cybersecurity-report/">https://cybersecurityventures.com/boardroom-cybersecurity-report/</ext-link></p>
</fn>. Numerous studies have shown that software security bugs are a significant source of such weaknesses [<xref ref-type="bibr" rid="ref-2">2</xref>], underscoring the need for early detection and mitigation in the development lifecycle [<xref ref-type="bibr" rid="ref-3">3</xref>]. However, as software grows in size and complexity, pinpointing the specific components most likely to harbor these flaws becomes increasingly challenging.</p>
<p>Traditional detection techniques rely heavily on dynamic analysis, which executes software under multiple inputs to identify problematic behaviors such as buffer overflows or unsafe library calls. Although widely used, dynamic approaches often suffer from high false-positive rates [<xref ref-type="bibr" rid="ref-4">4</xref>] and substantial computational cost, limiting their integration into rapid development workflows. Consequently, researchers have explored complementary approaches based on machine learning (ML), which has proven effective in domains such as spam filtering [<xref ref-type="bibr" rid="ref-5">5</xref>], intrusion detection [<xref ref-type="bibr" rid="ref-6">6</xref>], and cyberattack attribution [<xref ref-type="bibr" rid="ref-7">7</xref>]. Deep learning techniques, in particular, have gained significant attention for automated vulnerability detection [<xref ref-type="bibr" rid="ref-8">8</xref>]. In software engineering, ML has been applied to predict general software defects using static code metrics [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>Despite these advances, comparatively little is known about how traditional software quality metrics relate specifically to <italic>security</italic> bugs [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>]. Security vulnerabilities tend to be more subtle than functional bugs: they often arise from complex control-flow structures, rare edge cases, improper error-handling logic, API misuse, or intricate component interactions. Therefore, understanding whether classical quality metrics such as code size, structural complexity, or cyclomatic complexity can serve as meaningful indicators of vulnerability-prone code remains an open question.</p>
<p>Past studies investigating this topic have been limited by small datasets, language-specific scopes, or reliance on synthetic vulnerabilities, restricting the generalizability of their findings. The Open Source Vulnerabilities (OSV) platform enables the study of real-world security bugs across multiple ecosystems using commit-level vulnerability data. This platform enables large-scale empirical investigations that were previously infeasible. Throughout this study, we use the term <italic>security bugs</italic> to denote security-relevant code flaws that are addressed through security-fixing commits documented in the OSV database. While many of these bugs correspond to confirmed, exploitable vulnerabilities (e.g., associated with CVE identifiers), others represent latent security weaknesses whose exploitability may depend on specific deployment contexts or threat models. Consequently, our analysis focuses on security-relevant risk rather than confirmed exploitation alone.</p>
<p>The goal of this study is to explore, at scale, the relationship between software quality metrics and real-world security vulnerabilities. We also evaluate the feasibility of predicting security bugs using these metrics. To this end, we construct a dataset of 338,442 file-level instances drawn from seven major open-source ecosystems (GHSA, PyPI, OSS-Fuzz, npm, Packagist, Maven, and NuGet), including 33,294 files affected by 7685 confirmed security bugs. For each file, we extract 25 structural and complexity-related metrics using the <italic>Understand</italic> static analysis tool.</p>
<p>The following research questions guide our investigation:
<list list-type="bullet">
<list-item>
<p><bold>RQ1:</bold> Which software quality metrics are most strongly associated with the presence of security bugs? Identifying such metrics can help developers focus on code characteristics that elevate security risk.</p></list-item>
<list-item>
<p><bold>RQ2:</bold> Do specific values or thresholds of these metrics indicate a significantly higher likelihood of security bugs? Understanding differences in metric distributions can help define empirical guidelines for secure development.</p></list-item>
<list-item>
<p><bold>RQ3:</bold> Can machine learning models effectively predict security bugs using only software quality metrics? Assessing predictive capability is essential for integrating ML-based detectors into automated security workflows.</p></list-item>
</list></p>
<p>To address these questions, we combine feature importance analysis, statistical comparisons, and machine learning experiments using eleven classification algorithms evaluated through time-series cross-validation. Our findings reveal strong and consistent associations between specific metrics, particularly cyclomatic complexity and code size, and the presence of security bugs. In addition, ensemble-based ML models achieve high recall and strong discriminative performance, indicating that static quality metrics provide actionable signals for vulnerability detection. Overall, this study contributes large-scale empirical evidence that software quality metrics can play a meaningful role in identifying vulnerability-prone code. By linking traditional software metrics with real-world security bugs across multiple ecosystems, we provide actionable insights for integrating lightweight, automated, and metric-driven security analysis into modern software development practices.</p>
<p>The remainder of this paper is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> reviews related work on software defect prediction, vulnerability prediction, and security bug analysis. <xref ref-type="sec" rid="s3">Section 3</xref> describes the empirical study design, including ecosystem selection, data collection, metric extraction, and analytical procedures (<xref ref-type="sec" rid="s4">Section 4</xref>). <xref ref-type="sec" rid="s5">Section 5</xref> presents the results for each research question, along with their practical implications. <xref ref-type="sec" rid="s6">Section 6</xref> discusses threats to validity. <xref ref-type="sec" rid="s7">Section 7</xref> concludes the paper and outlines directions for future work. All data, scripts, and replication materials are available as Availability of Data and Materials at: <ext-link ext-link-type="uri" xlink:href="https://github.com/MdioufDataScientist/PredictSecBugs">https://github.com/MdioufDataScientist/PredictSecBugs</ext-link>.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Software security research commonly distinguishes between <italic>vulnerabilities</italic> (demonstrably exploitable weaknesses) and <italic>security bugs</italic> (security-relevant flaws that may not always be directly exploitable). This distinction matters when building empirical datasets from advisories and security-fixing commits, where labels often reflect security relevance rather than confirmed exploitability [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>]. In this section, we review prior work on (i) defect/bug prediction using software metrics, (ii) vulnerability prediction and detection using machine learning and deep learning (ML/DL) with an emphasis on recent trends (2020&#x2013;2024), and (iii) security bug prediction and empirical security studies. We conclude with a comparative positioning of our study.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Defect and Bug Prediction with Software Metrics</title>
<p>Foundational studies established that size and complexity metrics correlate with defect-proneness and can support risk-based quality assurance [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-17">17</xref>]. These works also showed that metric families (e.g., design vs. code metrics) can influence performance to a similar extent as the learning algorithm itself [<xref ref-type="bibr" rid="ref-17">17</xref>]. Later research highlighted the role of process metrics and the limitations of purely code-based metrics, including <italic>stasis</italic> (repeatedly flagging the duplicate files) and limited portability across releases and projects [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Two recurring challenges dominate the defect prediction literature and remain highly relevant for security settings: (i) <italic>class imbalance</italic> and (ii) <italic>evaluation realism</italic>. Extensive empirical studies show that class rebalancing techniques such as the Synthetic Minority Over-sampling Technique (SMOTE) [<xref ref-type="bibr" rid="ref-19">19</xref>], which generate artificial minority-class samples to rebalance training data, can improve recall, but may also distort explanatory analyses. The benefits of these techniques depend on the learner and the metric family used. Similarly, automated parameter optimization can substantially improve AUC but may alter variable importance rankings and interpretability [<xref ref-type="bibr" rid="ref-20">20</xref>]. Beyond traditional ML, deep and hybrid approaches have been investigated to capture non-linear interactions among metrics and reduce imbalance/overfitting effects, including autoencoder-ensemble strategies [<xref ref-type="bibr" rid="ref-21">21</xref>] and deep metric-driven models [<xref ref-type="bibr" rid="ref-22">22</xref>&#x2013;<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Vulnerability Prediction and Detection Using ML/DL</title>
<p>Vulnerability prediction extends defect prediction by focusing on security-relevant outcomes. Recent literature (2020&#x2013;2024) reports a growing dominance of deep learning and hybrid pipelines (e.g., combining static analysis signals with learned code representations), while also emphasizing persistent issues related to dataset bias, label noise, and generalization across projects and time [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>]. Recent empirical evaluations have revealed significant limitations in the generalization capabilities of deep learning models for vulnerability detection [<xref ref-type="bibr" rid="ref-29">29</xref>], while transformer-based approaches have been explored for fine-grained, line-level prediction [<xref ref-type="bibr" rid="ref-30">30</xref>]. A key message from recent work is that evaluation protocols strongly shape observed performance: improper separation or unrealistic settings can lead to overly optimistic conclusions [<xref ref-type="bibr" rid="ref-27">27</xref>]. Related work also highlights temporal shifts and distribution changes (concept drift) as important obstacles for realistic vulnerability assessment [<xref ref-type="bibr" rid="ref-31">31</xref>].</p>
<p>Another important line of research focuses on improving data quality. Evidence suggests that <italic>latent vulnerabilities</italic> (missed or delayed labels) can substantially affect training signals and measured performance; enriching datasets (e.g., via SZZ-based strategies) can increase labeled instances and improve detection performance [<xref ref-type="bibr" rid="ref-32">32</xref>]. For handling severe imbalance, recent approaches go beyond SMOTE and explore more advanced generation strategies (e.g., GAN-based oversampling) to improve minority recall under strong skew [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Security Bug Prediction and Empirical Security Studies</title>
<p>Empirical studies have investigated how security bugs differ from other defects in triage and evolution. Security-related bugs are often addressed faster and involve more developers and files [<xref ref-type="bibr" rid="ref-2">2</xref>]. Other large-scale analyses emphasize that vulnerabilities and general bugs are related but distinct phenomena [<xref ref-type="bibr" rid="ref-14">14</xref>]. Metric-based analyses of large systems, notably Mozilla, suggest that complexity and size metrics signal vulnerability-prone code. However, results may exhibit high false negatives depending on the setting [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>]. Other studies report that while metrics can help identify buggy functions, strong correlations with security bugs can be elusive in smaller or heterogeneous settings [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>Recent work increasingly leverages ML/DL on natural codebases and commit-level signals. For example, commit-focused ML combined with static analyzer signals has been used to predict vulnerability-fixing commits [<xref ref-type="bibr" rid="ref-34">34</xref>], while fine-grained vulnerability detection has been explored using learned representations on natural codebases (e.g., Python) [<xref ref-type="bibr" rid="ref-35">35</xref>]. Studies also report that gradient-boosted models, such as XGBoost, can be strong baselines for vulnerability prediction using static metrics, although cross-project transfer remains challenging [<xref ref-type="bibr" rid="ref-13">13</xref>]. In parallel, domain knowledge injection (e.g., CWE knowledge graphs) has been proposed to strengthen security bug report prediction [<xref ref-type="bibr" rid="ref-36">36</xref>], and large-scale classifications of security bugs highlight memory and input-validation issues as frequent root causes [<xref ref-type="bibr" rid="ref-37">37</xref>]. Finally, recent evidence shows that static analyzers may suffer from low precision/recall in practice, motivating complementary ML-based risk prioritization approaches [<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Comparative Positioning of This Study</title>
<p>To explicitly position our work with respect to representative studies, <xref ref-type="table" rid="table-1">Table 1</xref> summarizes key dimensions, including task scope, analysis granularity, dataset scale and diversity, validation strategy, and whether actionable thresholds are derived. Overall, recent literature converges on three persistent challenges in security bug and vulnerability prediction: (i) noisy or incomplete labeling, (ii) severe class imbalance, and (iii) limited generalization across ecosystems, programming languages, and time [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>]. Our study addresses these challenges by conducting a large-scale, multi-ecosystem empirical investigation at the file-level, combining multiple feature-importance techniques with statistical analysis and adopting a time-series validation strategy to reduce temporal information leakage. Unlike approaches that rely primarily on learned code representations, our metric-driven approach remains lightweight and directly actionable, enabling threshold-based risk screening and prioritization in CI/CD workflows.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison of this study with representative related work on security bug prediction.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Study</th>
<th>Year</th>
<th>Data Unit/Size</th>
<th>Ecosystems</th>
<th>Languages</th>
<th>Features/Signals</th>
<th>ML Models</th>
<th>Validation</th>
<th>Thresholds</th>
</tr>
</thead>
<tbody>
<tr>
<td>Shin &#x0026; Williams [<xref ref-type="bibr" rid="ref-11">11</xref>]</td>
<td>2008</td>
<td>1 project</td>
<td>Mozilla</td>
<td>JavaScript</td>
<td>Code complexity metrics</td>
<td>Statistical</td>
<td>Cross-validation</td>
<td>No</td>
</tr>
<tr>
<td>Alves et al. [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>2016</td>
<td>5 projects</td>
<td>Mixed OSS</td>
<td>C, Java</td>
<td>Quality metrics</td>
<td>Correlation analysis</td>
<td>Statistical</td>
<td>No</td>
</tr>
<tr>
<td>Clemente et al. [<xref ref-type="bibr" rid="ref-1">1</xref>]</td>
<td>2018</td>
<td>3 projects</td>
<td>Mozilla</td>
<td>Multiple</td>
<td>Size &#x0026; complexity metrics</td>
<td>ML (DT, RF, NN)</td>
<td>10-fold Cross-validation</td>
<td>No</td>
</tr>
<tr>
<td>Kalouptsoglou et al. [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>2020</td>
<td>4 projects</td>
<td>PHP applications</td>
<td>PHP</td>
<td>Software metrics</td>
<td>DL &#x002B; ML</td>
<td>Cross-project</td>
<td>No</td>
</tr>
<tr>
<td>VUDENC [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>2022</td>
<td>1009 commits</td>
<td>GitHub</td>
<td>Python</td>
<td>Token embeddings &#x002B; metrics</td>
<td>LSTM</td>
<td>Cross-validation</td>
<td>No</td>
</tr>
<tr>
<td>Ganesh et al. [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>2022</td>
<td>2 projects</td>
<td>Apache</td>
<td>Java</td>
<td>Static code metrics</td>
<td>XGBoost</td>
<td>Cross-version</td>
<td>No</td>
</tr>
<tr>
<td>Mashhadi et al. [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>2024</td>
<td>2 datasets</td>
<td>Defects4J</td>
<td>Java</td>
<td>Code metrics</td>
<td>Multiple</td>
<td>Cross-validation</td>
<td>No</td>
</tr>
<tr>
<td>Fehrer et al. [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>2024</td>
<td>1821 commits</td>
<td>OSS</td>
<td>Multiple</td>
<td>Static analyzers &#x002B; metrics</td>
<td>Stacking/<break/>Voting</td>
<td>Cross-validation</td>
<td>No</td>
</tr>
<tr>
<td>This study</td>
<td>2025</td>
<td>338,442 files</td>
<td>7 ecosystems</td>
<td>Multiple</td>
<td>Quality metrics (file-level)</td>
<td>11 models</td>
<td>Time-series Cross-validation</td>
<td>Yes</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-1">Table 1</xref> shows that most existing studies focus on a single ecosystem, a limited number of projects, or commit-level analysis, and typically rely on random or cross-validation strategies. In contrast, our study operates at the file-level across seven major ecosystems and adopts a time-series validation protocol to better approximate realistic deployment scenarios. Moreover, unlike prior work that primarily reports predictive performance, we derive actionable metric thresholds that support practical risk screening and prioritization in software development workflows.</p>

</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>This section presents the design of our empirical study following the Goal-Question-Metric (GQM) approach [<xref ref-type="bibr" rid="ref-41">41</xref>]. We first introduce the research questions that guide our work, then describe the study context and subject systems, and finally outline the data collection and preparation procedures. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> shows an overview of the methodology steps followed during this study.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>An overview of the methodology.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Research Questions</title>
<p>The overall goal of this study is to examine how file-level software quality metrics relate to the presence of security bugs and to assess whether these metrics can support early detection through machine learning. From the perspective of software developers and security practitioners, the study aims to provide actionable, metric-driven signals that help prioritize code review and security testing effort during development. Following the Goal&#x2013;Question&#x2013;Metric (GQM) approach, we formulate three research questions.</p>
<p><bold>RQ1.</bold> Which specific quality metrics are most strongly correlated with predicting security bugs?</p>
<p>The objective of RQ1 is to identify a subset of metrics that contribute the most to distinguishing buggy from non-buggy files. This provides a first indication of which aspects of code structure and complexity are most relevant from a security perspective. Understanding which code quality metrics are most strongly associated with security bugs enables developers to focus on the most critical indicators of potential vulnerabilities.</p>
<p><bold>RQ2.</bold> Do specific ranges or threshold values of these metrics correspond to significantly higher security-bug likelihood?</p>
<p>By comparing the distributions of metric values across buggy and clean files, RQ2 aims to identify critical ranges where the probability of a security issue increases significantly. Such thresholds could be integrated into quality gates or static analysis tools. Beyond determining which metrics matter, developers need concrete threshold values that signal elevated risk. Establishing these thresholds enables proactive alerts when code quality degrades to dangerous levels, allowing intervention before vulnerabilities are exploited.</p>
<p><bold>RQ3.</bold> How accurately can machine learning models predict security bugs using software quality metrics?</p>
<p>Here, the goal is twofold: (i) to assess whether classification models can effectively discriminate between buggy and non-buggy files and (ii) to compare a set of representative machine learning algorithms in terms of standard evaluation metrics, with particular emphasis on recall in a highly imbalanced security context. If machine learning models can accurately predict security bugs based on quality metrics, they can be integrated into development workflows to provide automated security assessments. This question evaluates the practical feasibility of ML-based security bug prediction.</p>
<p>Together, these three questions form a coherent pipeline: RQ1 identifies the most informative metrics, RQ2 examines their value ranges and statistical properties, and RQ3 evaluates their practical usefulness as features in predictive models.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Study Context and Subject Systems</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Ecosystem Selection Criteria</title>
<p>The study is conducted in the context of open-source software, where both vulnerability reports and code histories are publicly available. We selected open-source software ecosystems rather than individual projects to ensure:
<list list-type="bullet">
<list-item>
<p><bold>Diversity:</bold> coverage of multiple programming languages (Java, JavaScript, Python, PHP, C, etc.).</p></list-item>
<list-item>
<p><bold>Scale:</bold> large sample sizes for robust statistical analysis.</p></list-item>
<list-item>
<p><bold>Relevance:</bold> real-world security vulnerabilities documented in established databases.</p></list-item>
<list-item>
<p><bold>Reproducibility:</bold> publicly available data and well-maintained version control histories.</p></list-item>
</list></p>
<p>We rely on the Open Source Vulnerabilities (OSV) database<xref ref-type="fn" rid="fn-2"><sup>2</sup></xref><fn id="fn-2">
<label>2</label>
<p><ext-link ext-link-type="uri" xlink:href="https://osv.dev/">https://osv.dev/</ext-link></p>
</fn> as our primary source of security advisories to collect ecosystems that meet the following inclusion criteria:
<list list-type="simple">
<list-item><label>1.</label><p>Availability of documented security vulnerabilities with commit-level references.</p></list-item>
<list-item><label>2.</label><p>Active maintenance with a substantial development history.</p></list-item>
<list-item><label>3.</label><p>Diverse package sizes (from hundreds to millions of lines of code).</p></list-item>
<list-item><label>4.</label><p>Supported by major package management platforms.</p></list-item>
</list></p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Selected Ecosystems and Projects</title>
<p>We apply three inclusion criteria: (i) OSV must support the ecosystem with a sufficient number of reported security bugs, (ii) projects must be hosted in Git-based repositories with accessible commit histories, and (iii) the selected platforms should contribute to linguistic and technological diversity, ensuring that results generalize beyond a single language or domain. Based on these criteria, we retain seven ecosystems: GHSA (GitHub Security Advisories), PyPI, npm, Packagist, Maven, NuGet, and OSS-Fuzz. These ecosystems cover a broad range of programming languages (e.g., Python, JavaScript, PHP, Java, .NET) and software types (libraries, frameworks, tools, and infrastructure components), collectively representing a significant portion of the modern open-source software stack. Within the selected ecosystems, we focus on security bugs, for which OSV provides a link to a specific version-control commit that fixes the vulnerability. This constraint is essential to retrieve the exact source code context before and after the fix. For each ecosystem, we query OSV and filter advisories that contain at least one GitHub commit reference.</p>
<p>Brief descriptions of each ecosystem:
<list list-type="simple">
<list-item><label>1.</label><p><bold>OSS-Fuzz:</bold> A Google initiative using continuous fuzzing to discover vulnerabilities in critical open-source projects automatically. It collaborates with the OSV platform to disclose security bugs it finds.</p></list-item>
<list-item><label>2.</label><p><bold>GHSA (GitHub Security Advisory Database):</bold> Launched in 2019 to centralize security vulnerability management across GitHub-hosted projects. It enables maintainers to report vulnerabilities confidentially and publish advisories after patches are released.</p></list-item>
<list-item><label>3.</label><p><bold>PyPI (Python Package Index):</bold> The official repository for Python packages, established in 2003. As one of the oldest package ecosystems, it manages millions of packages and has an extensive history of vulnerability tracking.</p></list-item>
<list-item><label>4.</label><p><bold>Packagist (Composer):</bold> The primary PHP package repository launched in 2012, essential for modern PHP web application development and security.</p></list-item>
<list-item><label>5.</label><p><bold>npm (Node Package Manager):</bold> Introduced in 2010, npm is the most extensive package ecosystem, serving the JavaScript/Node.js community. Its modular architecture makes it particularly important for security research.</p></list-item>
<list-item><label>6.</label><p><bold>NuGet:</bold> The official package manager for .NET development, established in 2011, covering C#, F#, and Visual Basic projects.</p></list-item>
<list-item><label>7.</label><p><bold>Maven:</bold> A Java project management and build tool ecosystem launched in 2002, widely used in enterprise Java development.</p></list-item>
</list></p>
<p>Across these seven ecosystems, we identify 7685 confirmed security bugs for which fixing commits are available. These bugs are distributed across several open-source projects in different ecosystems. In this study, we adopt a more precise terminology: an <italic>ecosystem</italic> denotes a package management or vulnerability-reporting platform (such as PyPI or GHSA), whereas a <italic>project</italic> denotes a specific software repository hosted within one of these ecosystems. The final dataset aggregates file-level information from all projects that appear in at least one security-fixing commit in these seven ecosystems.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Dataset Overview</title>
<p>For each confirmed security bug, we retrieve the associated fixing commit. All files modified in that commit are considered part of the fix and are therefore labeled as buggy. To better characterize the context in which these fixes occur, we also include all other source files from the corresponding repository snapshot, which are not directly involved in the fix and are treated as non-buggy at that time. The final dataset comprises 338,442 file-level observations drawn from the seven ecosystems. Each observation corresponds to a specific source-code file, possibly in multiple versions (for example, before and after a fixing commit).</p>
<p><xref ref-type="table" rid="table-2">Table 2</xref> presents the distribution of files per ecosystem; this distribution reflects both the intensity of vulnerability reporting and the size of the underlying code bases in each ecosystem.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Selected software ecosystems and security bug distribution.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Ecosystem</th>
<th>Total Files</th>
<th>Buggy Files</th>
<th>Percentage Buggy (%)</th>
<th>Primary Languages</th>
<th>Year Established</th>
</tr>
</thead>
<tbody>
<tr>
<td>OSS-Fuzz</td>
<td>129,129</td>
<td>16,109</td>
<td>12.48</td>
<td>C, C&#x002B;&#x002B;, Various</td>
<td>2016</td>
</tr>
<tr>
<td>GHSA</td>
<td>100,783</td>
<td>9052</td>
<td>8.98</td>
<td>Multi-language</td>
<td>2019</td>
</tr>
<tr>
<td>PyPI</td>
<td>51,645</td>
<td>4024</td>
<td>7.79</td>
<td>Python</td>
<td>2003</td>
</tr>
<tr>
<td>Packagist</td>
<td>19,960</td>
<td>2274</td>
<td>11.39</td>
<td>PHP</td>
<td>2012</td>
</tr>
<tr>
<td>npm</td>
<td>19,188</td>
<td>552</td>
<td>2.88</td>
<td>JavaScript</td>
<td>2010</td>
</tr>
<tr>
<td>NuGet</td>
<td>8985</td>
<td>750</td>
<td>8.35</td>
<td>.NET/C#</td>
<td>2011</td>
</tr>
<tr>
<td>Maven</td>
<td>8752</td>
<td>533</td>
<td>6.09</td>
<td>Java</td>
<td>2002</td>
</tr>
<tr>
<td>Total</td>
<td>338,442</td>
<td>33,294</td>
<td>9.84</td>
<td></td>
<td></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>When available, we extract two versions of each file: the pre-patch version (before the fixing commit) and the post-patch version (after the fixing commit), as well as comparable versions for non-buggy files. These versions are later used to compute quality metrics and to construct the prediction dataset. As expected in a security context, the proportion of buggy files is relatively low across ecosystems, typically below 15% and as low as 2.88% in some cases, confirming that the resulting dataset is highly imbalanced. We address this through SMOTE oversampling [<xref ref-type="bibr" rid="ref-42">42</xref>] applied exclusively to training data (detailed in <xref ref-type="sec" rid="s3_3_3">Section 3.3.3</xref>):
<list list-type="bullet">
<list-item>
<p>33,294 buggy files (9.84%): Files confirmed as containing security vulnerabilities.</p></list-item>
<list-item>
<p>305,148 non-buggy files (90.16%): Files from the same projects without confirmed security issues.</p></list-item>
</list></p>
<p>Each file is characterized by:
<list list-type="bullet">
<list-item>
<p>25 quality metrics (detailed in <xref ref-type="sec" rid="s3_3_2">Section 3.3.2</xref>).</p></list-item>
<list-item>
<p>Metadata: commit date, commit SHA, file path, ecosystem, extension, source (pre/post-commit).</p></list-item>
<list-item>
<p>Target variable: Binary indicator (buggy &#x003D; 1, non-buggy &#x003D; 0).</p></list-item>
</list></p>
<p>Date Range: Security bugs span from 29 June 2005 to 07 May 2022, providing 17 years of security vulnerability data.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Data Collection and Preparation</title>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Security Bug Identification</title>
<p>A rigorous multi-stage process is used to identify and validate security bugs, comprising five steps designed to improve labeling reliability while preserving dataset scale.</p>
<p>Step 1: advisory collection. We query the OSV API for each of the seven ecosystems and retrieve all security advisories containing structured metadata such as vulnerability descriptions, severity ratings, affected versions, and references. We retain only advisories that include at least one reference to a GitHub commit fixing the vulnerability, resulting in 8247 advisory-commit pairs.</p>
<p>Step 2: commit-level Mapping. For each advisory, we clone the corresponding repository and retrieve the complete diff associated with the fixing commit. All files modified in that commit are initially flagged as potentially security-relevant, yielding 41,892 candidate buggy files.</p>
<p>Step 3: automated Filtering. To reduce labeling noise caused by files incidentally modified alongside security fixes, three automated filtering mechanisms are applied. (i) <italic>Diff-based analysis</italic>: files are retained only if the diff contains substantive code changes (i.e., additions or deletions of executable statements). Files involving only whitespace, comments, import reordering, or empty line changes are excluded. (ii) <italic>Commit message validation</italic>: commit messages are automatically scanned for security-related terms (e.g., <italic>security, vulnerability, CVE, exploit, fix, patch</italic>), increasing confidence that retained commits correspond to genuine security fixes. (iii) <italic>File-type filtering</italic>: non-source files unlikely to contain exploitable security bugs, such as documentation, configuration files, test code, build artifacts, and static assets, are excluded.</p>
<p>Collectively, these automated filters remove approximately 12% of the initially flagged candidate files, yielding a refined set of security-relevant buggy file candidates.</p>
<p>Step 4: contextual File Collection. For each security-fixing commit, we collect all other source files from the same repository at the pre-fix commit state. Files not modified by the fixing commit are treated as non-buggy instances for that snapshot, ensuring that buggy and non-buggy files share the same project and temporal context.</p>
<p>Step 5: manual Validation. To assess the reliability of the automated labeling process, we manually inspected a random sample of 769 security advisories and their corresponding security-fixing commits (approximately 10% of the dataset), spanning all seven ecosystems. Each sampled case was reviewed to verify whether the referenced commit genuinely addressed a security-relevant issue and whether the modified files contained security-related code changes. This validation confirmed that 724 cases (94.2%) clearly corresponded to security fixes, while 31 cases (4.0%) were considered borderline (e.g., defensive hardening or security-related refactoring without a clearly exploitable flaw). Fourteen cases (1.8%) were identified as incorrectly labeled and were removed from the dataset before analysis. The resulting estimated labeling error rate of approximately 1.8% indicates that residual noise is limited. Importantly, this noise is conservative: potential mislabeling primarily corresponds to non-security changes being included as buggy files (false positives), which would attenuate observed differences between buggy and non-buggy files rather than artificially inflate them. Consequently, the reported associations between quality metrics and security bugs are likely to represent a conservative lower bound on the true effect.</p>
<p>Overall, while commit-level labeling may still introduce limited noise due to auxiliary files modified during security patches, this approach follows widely adopted practices in vulnerability prediction research and provides a scalable, reproducible, and empirically validated mechanism for extensive multi-ecosystem analysis.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Quality Metrics Extraction</title>
<p>We use the SciTools Understand<xref ref-type="fn" rid="fn-3"><sup>3</sup></xref><fn id="fn-3">
<label>3</label>
<p><ext-link ext-link-type="uri" xlink:href="https://scitools.com/">https://scitools.com/</ext-link></p>
</fn> static analysis tool to compute software quality metrics at the file-level. Understand is a widely used static analysis tool in software engineering research [<xref ref-type="bibr" rid="ref-43">43</xref>] that supports multiple programming languages and provides consistent metric definitions across languages [<xref ref-type="bibr" rid="ref-44">44</xref>]. For each source file in our dataset, whether buggy or non-buggy, we invoke Understand from the command line and extract the complete set of metrics supported for the corresponding language. The tool produces individual CSV files per project, which we then merge into a single data file.</p>
<p>In total, Understand provides 40 metrics<xref ref-type="fn" rid="fn-4"><sup>4</sup></xref><fn id="fn-4">
<label>4</label>
<p><ext-link ext-link-type="uri" xlink:href="https://support.scitools.com/support/solutions/articles/70000582223-what-metrics-does-understand-have-">https://support.scitools.com/support/solutions/articles/70000582223-what-metrics-does-understand-have-</ext-link></p>
</fn> across all considered languages, including various measures of size (e.g., lines of code), complexity (e.g., cyclomatic metrics), and structural information (e.g., counts of declarations and statements). However, since our dataset spans multiple programming languages, not all metrics are defined uniformly across languages. To ensure comparability, we retain only the 25 metrics that are common to all languages and consistently supported across ecosystems. <xref ref-type="table" rid="table-3">Table 3</xref> organizes these metrics into three categories.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The 25 quality metrics.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center" colspan="2">Category 1: Code Lines Metrics (10 metrics)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CountLineComment</td>
<td>Total comment lines</td>
</tr>
<tr>
<td>CountLineBlank</td>
<td>Total blank lines</td>
</tr>
<tr>
<td>CountLine</td>
<td>Total lines of code</td>
</tr>
<tr>
<td>CountLineCode</td>
<td>Total executable code lines</td>
</tr>
<tr>
<td>AvgLineComment</td>
<td>Average comment lines per function</td>
</tr>
<tr>
<td>AvgLine</td>
<td>Average lines per function</td>
</tr>
<tr>
<td>AvgLineCode</td>
<td>Average code lines per function</td>
</tr>
<tr>
<td>AvgLineBlank</td>
<td>Average blank lines per function</td>
</tr>
<tr>
<td>RatioCommentToCode</td>
<td>Ratio of comment to code lines</td>
</tr>
<tr>
<td align="center" colspan="2"><bold>Category 2: Cyclomatic Complexity Metrics (9 metrics)</bold></td>
</tr>
<tr>
<td>SumCyclomatic</td>
<td>Sum of cyclomatic complexity</td>
</tr>
<tr>
<td>SumCyclomaticModified</td>
<td>Sum of modified cyclomatic complexity</td>
</tr>
<tr>
<td>SumCyclomaticStrict</td>
<td>Sum of strict cyclomatic complexity</td>
</tr>
<tr>
<td>AvgCyclomatic</td>
<td>Average cyclomatic complexity</td>
</tr>
<tr>
<td>AvgCyclomaticModified</td>
<td>Average modified cyclomatic complexity</td>
</tr>
<tr>
<td>AvgCyclomaticStrict</td>
<td>Average strict cyclomatic complexity</td>
</tr>
<tr>
<td>MaxCyclomatic</td>
<td>Maximum cyclomatic complexity</td>
</tr>
<tr>
<td>MaxCyclomaticModified</td>
<td>Maximum modified cyclomatic complexity</td>
</tr>
<tr>
<td>MaxEssential</td>
<td>Maximum essential complexity</td>
</tr>
<tr>
<td align="center" colspan="2"><bold>Category 3: Code Structure Metrics (6 metrics)</bold></td>
</tr>
<tr>
<td>CountDeclClass</td>
<td>Number of class declarations</td>
</tr>
<tr>
<td>CountDeclFunction</td>
<td>Number of function/method declarations</td>
</tr>
<tr>
<td>CountStmt</td>
<td>Total number of statements</td>
</tr>
<tr>
<td>CountStmtExe</td>
<td>Number of executable statements</td>
</tr>
<tr>
<td>CountStmtDecl</td>
<td>Number of declaration statements</td>
</tr>
<tr>
<td>MaxNesting</td>
<td>Maximum nesting level</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The resulting dataset contains, for each file: (i) the 25 selected quality metrics, (ii) a binary target variable &#x201C;buggy&#x201D;, indicating whether the file is involved in a security bug (1) or not (0) and (iii) several metadata attributes, including the file path, ecosystem, commit date and commit identifier, as well as the source version (pre-commit or post-commit).</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Data Balancing Strategy</title>
<p>As indicated by the distribution of buggy files across ecosystems, the proportion of buggy files is substantially lower than that of non-buggy files. Training machine learning models directly on such imbalanced data typically leads to classifiers that favor the majority class and achieve misleadingly high accuracy while performing poorly on the minority class, which is precisely the class of interest in our context [<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
<p>To mitigate this issue, we introduce a binary target variable, &#x201C;buggy&#x201D;, that takes the value 1 when a file is associated with a security bug and 0 otherwise. We then apply oversampling techniques to the minority class during model training. Specifically, we compare three standard methods: na&#x00EF;ve random oversampling<xref ref-type="fn" rid="fn-5"><sup>5</sup></xref><fn id="fn-5">
<label>5</label>
<p><ext-link ext-link-type="uri" xlink:href="https://machinelearningmastery.com/random-oversampling-and-undersampling-for-imbalanced-classification/">https://machinelearningmastery.com/random-oversampling-and-undersampling-for-imbalanced-classification/</ext-link></p>
</fn>, the Synthetic Minority Over-sampling Technique<xref ref-type="fn" rid="fn-6"><sup>6</sup></xref><fn id="fn-6">
<label>6</label>
<p><ext-link ext-link-type="uri" xlink:href="https://machinelearningmastery.com/smote-oversampling-for-imbalanced-classification/">https://machinelearningmastery.com/smote-oversampling-for-imbalanced-classification/</ext-link></p>
</fn> (SMOTE), and ADASYN<xref ref-type="fn" rid="fn-7"><sup>7</sup></xref><fn id="fn-7">
<label>7</label>
<p><ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/residentmario/oversampling-with-smote-and-adasyn">https://www.kaggle.com/residentmario/oversampling-with-smote-and-adasyn</ext-link></p>
</fn>. All three techniques assume that predictors are numerical; this assumption holds in our case, since the 25 retained metrics are continuous or integer-valued.</p>
<p>In our experiments, SMOTE consistently provides the best balance between preserving the original data distribution and improving model performance. SMOTE generates synthetic minority examples by interpolating between existing buggy instances in the feature space, which reduces overfitting compared to simply duplicating examples. Importantly, resampling is applied only to the training folds during cross-validation. The validation and test folds remain untouched and reflect the original class distribution, thereby providing an unbiased estimate of the generalization performance.</p>
<p>While SMOTE is an effective and widely used technique for mitigating class imbalance, its use may influence model generalization by introducing synthetic instances that can smooth decision boundaries. In particular, oversampling-based methods may reduce variance at the cost of potentially obscuring rare but meaningful patterns present in the minority class. To mitigate these risks, oversampling in our study is applied strictly within training folds, while validation and test sets preserve the original class distribution, thereby avoiding information leakage and overly optimistic performance estimates. Nevertheless, alternative imbalance handling strategies could also be considered. Cost-sensitive learning, for example, directly incorporates class imbalance into the loss function, while ensemble-based approaches or under-sampling techniques may reduce reliance on synthetic data. We selected SMOTE as the pragmatic baseline in this study due to its simplicity, reproducibility, and frequent use in prior security and defect prediction studies. A systematic comparison of imbalance-handling strategies and their impact on generalization across ecosystems represents a vital direction and is further discussed in the Threats to Validity section.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Analysis Methods</title>
<p>This section details the specific analytical techniques employed to answer each research question.</p>
<sec id="s4_1">
<label>4.1</label>
<title>RQ1 Feature Importance Analysis</title>
<p>To address RQ1, we investigate which quality metrics are most strongly associated with the presence of security bugs. We analyze feature importance by applying five complementary feature-selection or importance-estimation techniques to the entire combined dataset (all ecosystems).</p>
<p>First, we fit a Linear Regression and a Logistic Regression model, with the &#x201C;buggy&#x201D; variable as the dependent variable and the 25 quality metrics as predictors. In both cases, the absolute values of the regression coefficients indicate feature importance, with larger magnitudes corresponding to stronger associations with the target. We used two approaches, &#x201C;Extra Trees Classifier Learning&#x201D;<xref ref-type="fn" rid="fn-8"><sup>8</sup></xref><fn id="fn-8">
<label>8</label>
<p><ext-link ext-link-type="uri" xlink:href="https://www.geeksforgeeks.org/ml-extra-tree-classifier-for-feature-selection/">https://www.geeksforgeeks.org/ml-extra-tree-classifier-for-feature-selection/</ext-link></p>
</fn> and &#x201C;Scott-Knott ESD&#x201D;<xref ref-type="fn" rid="fn-9"><sup>9</sup></xref><fn id="fn-9">
<label>9</label>
<p><ext-link ext-link-type="uri" xlink:href="https://github.com/klainfo/ScottKnottESD">https://github.com/klainfo/ScottKnottESD</ext-link></p>
</fn> to identify the importance of features, i.e., the quality metrics that may have the highest impact on the appearance of security bugs. These approaches combine predictions from multiple decision trees to effectively select the more informative features from the datasets [<xref ref-type="bibr" rid="ref-45">45</xref>]. Moreover, we analyzed the correlation between these quality metrics using Pearson&#x2019;s correlation measurement [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>].</p>
<p>Second, we train an XGBoost classifier and extract its internal feature importance scores, which are based on the gain in the objective function produced by splits on each feature across all trees in the ensemble.</p>
<p>Third, we compute the ANOVA F-score for each metric. In this context, the F-score measures the ratio of between-class variance to within-class variance for each feature. A higher F-score indicates that the distribution of the metric differs substantially between buggy and non-buggy files, making it more discriminative.</p>
<p>Finally, we employ the Boruta algorithm [<xref ref-type="bibr" rid="ref-48">48</xref>], which is a wrapper method built on top of Random Forests. Boruta creates shadow copies of the original features by shuffling their values, trains a Random Forest on the extended feature set, and iteratively compares the importance of each fundamental feature to the maximum significance of the shadow features. Features that consistently outperform their shadow counterparts are deemed necessary.</p>
<p>For each method, we rank the metrics by their importance scores, then compute the intersection of the top-10 metrics across methods to identify a stable subset of highly influential quality metrics. Additionally, we calculate Pearson correlation coefficients between all metric pairs and between each metric and the target variable. This reveals both:
<list list-type="bullet">
<list-item>
<p>Which metrics are most associated with security bugs.</p></list-item>
<list-item>
<p>Which metrics are highly correlated with each other (potential redundancy).</p></list-item>
</list></p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>RQ2 for Statistical Analysis of Metric Values</title>
<p>For RQ2, we investigate whether specific ranges or thresholds of selected metrics are associated with a higher probability of a file being buggy. We focus on the subset of metrics that emerge as consistently crucial in the RQ1 analysis. For each of these metrics, we compute descriptive statistics separately for buggy and non-buggy files, including mean, standard deviation, minimum, maximum, median, and the first and third quartiles (Q1 and Q3). We then generate density and box plots to compare the distributions of metric values between the two groups. Density plots show the probability distributions of metric values for both classes, revealing differences in distribution shapes, and box plots display quartiles, medians, and outliers, providing an intuitive comparison of central tendency and spread. To statistically validate the observed differences, we apply the non-parametric Mann-Whitney U test [<xref ref-type="bibr" rid="ref-49">49</xref>], which assesses whether two independent samples are drawn from the same distribution. We choose this test as it is a popular nonparametric one used in several previous works in bug and software analysis to test whether two samples are likely to derive from the same population (i.e., that the two populations have the same shape) [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>]. This test is appropriate for our data, as it does not assume normality and only requires independent observations. To avoid dominance of the majority class, we first perform random undersampling of non-buggy files to obtain a balanced sample and then draw a random subset of this balanced data for the test. For each metric, we test the null hypothesis that the distributions of values for buggy and non-buggy files are equal, against the alternative that they differ.</p>
<p>Test Details:
<list list-type="bullet">
<list-item>
<p>Null Hypothesis (<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>): The distributions of metric values are identical for buggy and non-buggy files.</p></list-item>
<list-item>
<p>Alternative Hypothesis (<inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>): The distributions differ significantly.</p></list-item>
<list-item>
<p>Significance level: <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula> (95% confidence).</p></list-item>
<list-item>
<p>Test choice rationale: Mann-Whitney U is a non-parametric test appropriate for our data because:
<list list-type="simple">
<list-item><label>&#x2013;</label><p>It does not assume normal distribution.</p></list-item>
<list-item><label>&#x2013;</label><p>It is robust to outliers.</p></list-item>
<list-item><label>&#x2013;</label><p>It tests whether two independent samples come from the same distribution.</p></list-item>
<list-item><label>&#x2013;</label><p>It is widely used in software engineering research [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>].</p></list-item>
</list></p></list-item>
</list></p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>RQ3: Machine Learning Models</title>
<p>To address RQ3, we evaluate the ability of various machine learning models to predict whether a given file is buggy solely from its quality metrics. We consider eleven widely used classification algorithms: <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest Neighbors, Decision Trees, Random Forests, Multi-layer Perceptron, AdaBoost, XGBoost, Na&#x00EF;ve Bayes, Quadratic Discriminant Analysis (QDA), Support Vector Machines with linear and Gaussian kernels, and Logistic Regression:
<list list-type="simple">
<list-item><label>1.</label><p><bold><inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest Neighbors (KNN):</bold> Instance-based learning (k &#x003D; 3).</p></list-item>
<list-item><label>2.</label><p><bold>Decision Tree:</bold> Single tree classifier (max_depth &#x003D; 3).</p></list-item>
<list-item><label>3.</label><p><bold>Random Forest:</bold> Ensemble of decision trees (n_estimators &#x003D; 10, max_depth &#x003D; 3).</p></list-item>
<list-item><label>4.</label><p><bold>Multi-layer Perceptron (MLP):</bold> Neural network (activation &#x003D; &#x201C;relu&#x201D;, solver &#x003D; &#x201C;adam&#x201D;).</p></list-item>
<list-item><label>5.</label><p><bold>AdaBoost:</bold> Boosting ensemble (n_estimators &#x003D; 50).</p></list-item>
<list-item><label>6.</label><p><bold>XGBoost:</bold> Gradient boosting (default parameters).</p></list-item>
<list-item><label>7.</label><p><bold>Na&#x00EF;ve Bayes:</bold> Probabilistic classifier.</p></list-item>
<list-item><label>8.</label><p><bold>Quadratic Discriminant Analysis (QDA):</bold> Gaussian classifier.</p></list-item>
<list-item><label>9.</label><p><bold>Support Vector Machine (Linear):</bold> Linear kernel SVM.</p></list-item>
<list-item><label>10.</label><p><bold>Support Vector Machine (RBF):</bold> Gaussian kernel SVM.</p></list-item>
<list-item><label>11.</label><p><bold>Logistic Regression:</bold> Linear probabilistic classifier.</p></list-item>
</list></p>
<p><xref ref-type="table" rid="table-4">Table 4</xref> presents the hyperparameter configurations for each algorithm.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Machine Learning Hyper-parameters.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>ML algorithm</th>
<th>Hyper-parameters</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest Neighbors</td>
<td>Number of neighbors &#x003D; 3</td>
</tr>
<tr>
<td>Decision Tree</td>
<td>max_depth &#x003D; 3</td>
</tr>
<tr>
<td>Random Forest</td>
<td>max_depth &#x003D; 3, n_estimators &#x003D; 10, max_features &#x003D; 1</td>
</tr>
<tr>
<td>Multi-layer Perceptron</td>
<td>activation &#x003D; relu, solver &#x003D; adam, alpha &#x003D; 1, batch_size &#x003D; auto</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>n_estimators &#x003D; 50</td>
</tr>
<tr>
<td>Na&#x00EF;ve Bayes</td>
<td>priors &#x003D; None, var_smoothing &#x003D; 1e-09</td>
</tr>
<tr>
<td>QDA</td>
<td>priors &#x003D; None, reg_param &#x003D; 0.0, store_covariance &#x003D; False, tol &#x003D; 0.0001</td>
</tr>
<tr>
<td>SVM (Linear)</td>
<td>kernel &#x003D; linear</td>
</tr>
<tr>
<td>SVM (RBF)</td>
<td>kernel &#x003D; rbf</td>
</tr>
<tr>
<td>Logistic Regression</td>
<td>random_state &#x003D; 0, solver &#x003D; lbfgs, max_iter &#x003D; 100</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We adopt a time-series cross-validation strategy with five folds. In each fold, the data is partitioned chronologically so that training is performed on earlier commits and validation on later ones, thereby mimicking a realistic prediction scenario and avoiding information leakage from the future to the past. Within each training fold, we apply SMOTE [<xref ref-type="bibr" rid="ref-42">42</xref>] to balance the classes, as described previously. The validation and final test sets retain the original class distribution.</p>
<p>Model performance is assessed using six evaluation metrics: Accuracy, Precision, Recall, F1-score, Matthews Correlation Coefficient (MCC), and the Area Under the ROC Curve (ROC AUC). Accuracy provides a global view of the proportion of correctly classified instances, but can be misleading under strong class imbalance. Precision and Recall capture, respectively, the proportion of predicted buggy files that are truly buggy and the proportion of truly buggy files that are correctly identified. The F1-score combines both into a single harmonic mean. MCC provides a more balanced summary statistic that accounts for all four entries of the confusion matrix and is robust to class imbalance. Finally, the ROC AUC measures the model&#x2019;s discriminative power across all possible decision thresholds. Each evaluation metric provides different insights (Where: TP &#x003D; True Positives, TN &#x003D; True Negatives, FP &#x003D; False Positives, FN &#x003D; False Negatives):
<list list-type="simple">
<list-item><label>1.</label><p><bold>Accuracy:</bold> Overall proportion of correct predictions.
<disp-formula id="ueqn-1"><mml:math id="mml-ueqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>Accuracy</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>Note: Can be misleading with imbalanced data.</p></list-item>
<list-item><label>2.</label><p><bold>Precision:</bold> Proportion of predicted bugs that are actual bugs.
<disp-formula id="ueqn-2"><mml:math id="mml-ueqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>Interpretation: Low precision means many false alarms.</p></list-item>
<list-item><label>3.</label><p><bold>Recall (Sensitivity):</bold> Proportion of actual bugs that are detected.
<disp-formula id="ueqn-3"><mml:math id="mml-ueqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>Interpretation: Low recall means missing real vulnerabilities. Critical for security applications.</p></list-item>
<list-item><label>4.</label><p><bold>F1-Score:</bold> Harmonic mean of precision and recall.
<disp-formula id="ueqn-4"><mml:math id="mml-ueqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>F1-Score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>Balances precision and recall.</p></list-item>
<list-item><label>5.</label><p><bold>Matthews Correlation Coefficient (MCC):</bold> Balanced measure even with class imbalance.
<disp-formula id="ueqn-5"><mml:math id="mml-ueqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>MCC</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:msqrt><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:msqrt></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>Range: <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> (total disagreement) to <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> (perfect prediction).</p></list-item>
<list-item><label>6.</label><p><bold>ROC-AUC:</bold> Area under the Receiver Operating Characteristic curve.
<list list-type="simple">
<list-item><label>&#x2022;</label>
<p>Measures ability to discriminate between classes across all thresholds.</p></list-item>
<list-item><label>&#x2022;</label>
<p>Range: 0.5 (random) to 1.0 (perfect).</p></list-item>
<list-item><label>&#x2022;</label>
<p>Critical for security applications as it&#x2019;s threshold-independent.</p></list-item>
</list></p></list-item>
</list></p>
<p>In security contexts, false negatives (missing real vulnerabilities) are more costly than false positives (false alarms). Therefore, we prioritize Recall and ROC-AUC when evaluating model performance. High recall ensures we detect most security bugs, even if it means investigating some false alarms.</p>
<p>Following best practices for data preprocessing, we:
<list list-type="simple">
<list-item><label>1.</label><p>Removed duplicate files (identical files across commits that remained unchanged).</p></list-item>
<list-item><label>2.</label><p>Verified data types and handled missing values (none found).</p></list-item>
<list-item><label>3.</label><p>Applied SMOTE only to training folds (not test data).</p></list-item>
<list-item><label>4.</label><p>Maintained temporal ordering throughout all processing steps.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Results and Discussion</title>
<p>In this section, we present and discuss the results obtained for each research question. For each RQ, we first report the quantitative findings and then discuss their implications. We conclude with an overall summary of the main findings.</p>
<sec id="s5_1">
<label>5.1</label>
<title>RQ1: Quality Metrics Correlation with Security Bugs</title>
<p><bold>Research Question:</bold> Which specific quality metrics are most strongly correlated with predicting security bugs?</p>
<sec id="s5_1_1">
<label>5.1.1</label>
<title>Results</title>
<p>The feature importance analyses reveal a consistent pattern across the five methods. Although each technique ranks metrics according to its own criterion, several metrics recur among the top-ranked features.</p>
<p>By intersecting the top-10 metrics obtained from Linear Regression, Logistic Regression, XGBoost, ANOVA F-score, and Boruta, we identify six metrics that are consistently highlighted as necessary: <monospace>CountStmtExe</monospace>, <monospace>SumCyclomatic</monospace>, <monospace>SumCyclomaticModified</monospace>, <monospace>AvgCyclomatic</monospace><monospace>Modified</monospace>, <monospace>CountDeclClass</monospace>, and the categorical variable <monospace>extension</monospace>, which encodes the programming language or file type. These metrics describe, respectively, the number of executable statements in a file, different variants of cyclomatic complexity, and the number of class declarations. <xref ref-type="table" rid="table-5">Tables 5</xref>&#x2013;<xref ref-type="table" rid="table-9">9</xref> present the ranked importance of all 25 quality metrics according to each method. Specifically, <xref ref-type="table" rid="table-5">Table 5</xref> shows Linear Regression coefficients, <xref ref-type="table" rid="table-6">Table 6</xref> reports XGBoost feature importance scores, <xref ref-type="table" rid="table-7">Table 7</xref> presents Logistic Regression coefficients, <xref ref-type="table" rid="table-8">Table 8</xref> displays ANOVA F-scores and <xref ref-type="table" rid="table-9">Table 9</xref> summarizes Boruta algorithm results.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Feature importance using Linear Regression coefficients.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Rank</th>
<th>Metric</th>
<th>Coefficient</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>CountStmtExe</td>
<td>8.2</td>
</tr>
<tr>
<td>2</td>
<td>CountDeclClass</td>
<td>7.5</td>
</tr>
<tr>
<td>3</td>
<td>AvgLineBlank</td>
<td>6.6</td>
</tr>
<tr>
<td>4</td>
<td>Extension</td>
<td>6.3</td>
</tr>
<tr>
<td>5</td>
<td>CountStmtDecl</td>
<td>3.2</td>
</tr>
<tr>
<td>6</td>
<td>CountDeclFunction</td>
<td>2.3</td>
</tr>
<tr>
<td>7</td>
<td>AvgCyclomaticModified</td>
<td>1.8</td>
</tr>
<tr>
<td>8</td>
<td>Ecosystem</td>
<td>1.5</td>
</tr>
<tr>
<td>9</td>
<td>MaxEssential</td>
<td>1.4</td>
</tr>
<tr>
<td>10</td>
<td>MaxNesting</td>
<td>1.4</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-5fn1" fn-type="other">
<p>Note: Positive coefficients indicate positive correlation with security bugs; negative coefficients indicate negative correlation. Full table available in Availability of Data and Materials.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Feature importance using XGBoost classifier.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Rank</th>
<th>Metric</th>
<th>Importance Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>CountDeclClass</td>
<td>0.466</td>
</tr>
<tr>
<td>2</td>
<td>CountStmtExe</td>
<td>0.126</td>
</tr>
<tr>
<td>3</td>
<td>Extension</td>
<td>0.076</td>
</tr>
<tr>
<td>4</td>
<td>SumCyclomatic</td>
<td>0.041</td>
</tr>
<tr>
<td>5</td>
<td>MaxNesting</td>
<td>0.030</td>
</tr>
<tr>
<td>6</td>
<td>SumCyclomaticModified</td>
<td>0.016</td>
</tr>
<tr>
<td>7</td>
<td>AvgCyclomaticModified</td>
<td>0.016</td>
</tr>
<tr>
<td>8</td>
<td>AvgLine</td>
<td>0.013</td>
</tr>
<tr>
<td>9</td>
<td>CountStmt</td>
<td>0.013</td>
</tr>
<tr>
<td>10</td>
<td>AvgCyclomatic</td>
<td>0.012</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-6fn1" fn-type="other">
<p>Note: XGBoost importance scores sum to 1.0. Higher scores indicate greater contribution to prediction accuracy.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Feature importance using Logistic Regression coefficients.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Rank</th>
<th>Metric</th>
<th>Coefficient</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>CountStmtExe</td>
<td>6.244</td>
</tr>
<tr>
<td>2</td>
<td>SumCyclomatic</td>
<td>2.672</td>
</tr>
<tr>
<td>3</td>
<td>CountStmtDecl</td>
<td>2.317</td>
</tr>
<tr>
<td>4</td>
<td>SumCyclomaticModified</td>
<td>1.969</td>
</tr>
<tr>
<td>5</td>
<td>AvgCyclomaticModified</td>
<td>1.687</td>
</tr>
<tr>
<td>6</td>
<td>MaxNesting</td>
<td>0.481</td>
</tr>
<tr>
<td>7</td>
<td>Extension</td>
<td>0.442</td>
</tr>
<tr>
<td>8</td>
<td>CountDeclClass</td>
<td>0.426</td>
</tr>
<tr>
<td>9</td>
<td>AvgLineBlank</td>
<td>0.198</td>
</tr>
<tr>
<td>10</td>
<td>MaxEssential</td>
<td>0.140</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-7fn1" fn-type="other">
<p>Note: Coefficients represent log-odds ratios for the binary classification of security bugs.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Feature importance using F-score (ANOVA F-statistic).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Rank</th>
<th>Metric</th>
<th>F-Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>CountStmtExe</td>
<td>47,332.3</td>
</tr>
<tr>
<td>2</td>
<td>CountDeclClass</td>
<td>9278.3</td>
</tr>
<tr>
<td>3</td>
<td>SumCyclomatic</td>
<td>8828.3</td>
</tr>
<tr>
<td>4</td>
<td>SumCyclomaticModified</td>
<td>7257.5</td>
</tr>
<tr>
<td>5</td>
<td>MaxEssential</td>
<td>6788.4</td>
</tr>
<tr>
<td>6</td>
<td>Extension</td>
<td>4174.5</td>
</tr>
<tr>
<td>7</td>
<td>CountStmtDecl</td>
<td>3910.1</td>
</tr>
<tr>
<td>8</td>
<td>AvgCyclomaticModified</td>
<td>3784.1</td>
</tr>
<tr>
<td>9</td>
<td>CountLine</td>
<td>3124.5</td>
</tr>
<tr>
<td>10</td>
<td>CountLineBlank</td>
<td>2258.5</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-8fn1" fn-type="other">
<p>Note: F-score measures the ratio of between-group to within-group variance. Higher scores indicate better separation between buggy and non-buggy files.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Feature importance using Boruta algorithm.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Rank</th>
<th>Metric</th>
<th>Status</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>CountStmtExe</td>
<td>Confirmed</td>
</tr>
<tr>
<td>2</td>
<td>SumCyclomaticModified</td>
<td>Confirmed</td>
</tr>
<tr>
<td>3</td>
<td>SumCyclomatic</td>
<td>Confirmed</td>
</tr>
<tr>
<td>4</td>
<td>AvgLineCode</td>
<td>Confirmed</td>
</tr>
<tr>
<td>5</td>
<td>CountDeclClass</td>
<td>Confirmed</td>
</tr>
<tr>
<td>6</td>
<td>AvgCyclomaticModified</td>
<td>Confirmed</td>
</tr>
<tr>
<td>7</td>
<td>Extension</td>
<td>Confirmed</td>
</tr>
<tr>
<td>8</td>
<td>AvgCyclomaticStrict</td>
<td>Confirmed</td>
</tr>
<tr>
<td>9</td>
<td>CountLine</td>
<td>Confirmed</td>
</tr>
<tr>
<td>10</td>
<td>AvgCyclomatic</td>
<td>Confirmed</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-9fn1" fn-type="other">
<p>Note: Boruta tests whether each feature performs better than random shadow features. All top-10 metrics were confirmed to be significantly better than noise.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="table-10">Table 10</xref> synthesizes the top 10 features identified by each method. Despite different ranking mechanisms, we observe strong consensus on several metrics.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Top 10 features per feature selection method.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Top 10 Features (Ranked)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Linear Regression</td>
<td>CountStmtExe, CountDeclClass, AvgLineBlank, extension, CountStmtDecl, CountDeclFunction, AvgCyclomaticModified, Ecosystem, MaxEssential, MaxNesting</td>
</tr>
<tr>
<td>Logistic Regression</td>
<td>CountStmtExe, SumCyclomatic, CountStmtDecl, SumCyclomaticModified, AvgCyclomaticModified, MaxNesting, extension, CountDeclClass, AvgLineBlank, MaxEssential</td>
</tr>
<tr>
<td>XGBoost</td>
<td>CountDeclClass, CountStmtExe, extension, SumCyclomatic, MaxNesting, SumCyclomaticModified, AvgCyclomaticModified, AvgLine, CountStmt, AvgCyclomatic</td>
</tr>
<tr>
<td>F-Score</td>
<td>CountStmtExe, CountDeclClass, SumCyclomatic, SumCyclomaticModified, MaxEssential, extension, CountStmtDecl, AvgCyclomaticModified, CountLine, CountLineBlank</td>
</tr>
<tr>
<td>Boruta</td>
<td>CountStmtExe, SumCyclomaticModified, SumCyclomatic, AvgLineCode, CountDeclClass, AvgCyclomaticModified, extension, AvgCyclomaticStrict, CountLine, AvgCyclomatic</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Analyzing the intersection of metrics appearing in the top 10 across all five methods, we identify six core metrics consistently ranked highly:
<list list-type="simple">
<list-item><label>1.</label><p><bold>CountStmtExe</bold> (Executable statements count): ranked #1 or #2 in all methods.</p></list-item>
<list-item><label>2.</label><p><bold>CountDeclClass</bold> (Class declarations count): ranked between #1 and #8 across methods.</p></list-item>
<list-item><label>3.</label><p><bold>SumCyclomatic</bold> (Total cyclomatic complexity): ranked between #2 and #4 in four methods.</p></list-item>
<list-item><label>4.</label><p><bold>SumCyclomaticModified</bold> (Total modified cyclomatic complexity): ranked between #2 and #6 in four methods.</p></list-item>
<list-item><label>5.</label><p><bold>AvgCyclomaticModified</bold> (Average modified cyclomatic complexity): ranked between #5 and #8 in all methods.</p></list-item>
<list-item><label>6.</label><p><bold>Extension</bold> (File extension/programming language): ranked between #3 and #7 in all methods.</p></list-item>
</list></p>
<p>This convergence across multiple feature selection techniques suggests that code size and control-flow complexity are strongly associated with the presence of security bugs, regardless of the specific modeling assumptions underlying each method. Other metrics, such as various line-count measures or declaration counts, also appear as relevant in specific methods but do not show the same level of consistency.</p>
</sec>
<sec id="s5_1_2">
<label>5.1.2</label>
<title>Discussion</title>
<p>The results of RQ1 confirm and extend observations made in prior defect-prediction studies: larger and more complex code units tend to be more prone to defects. In our context, the fact that cyclomatic complexity metrics and the number of executable statements are consistently ranked among the most influential features indicates that code structure is also a key factor in security-relevant defects. The consistent identification of these six metrics across diverse feature selection methods provides strong evidence of their fundamental relationship with security bugs. We interpret each metric as follows.</p>
<p><italic>CountStmtExe (Executable Statements)</italic>.</p>
<p>This metric&#x2019;s dominant ranking across all methods (achieving the highest F-score of 47,332.3) indicates that files with more executable statements are substantially more likely to contain security bugs. This aligns with the intuition that larger, more complex code provides more opportunities for security-relevant errors. Each additional executable statement represents a potential location for vulnerabilities such as input validation failures, improper error handling, or unsafe operations.</p>
<p><italic>CountDeclClass (Class Declarations)</italic>.</p>
<p>The strong performance of the class declaration count suggests that files defining multiple classes exhibit a higher security risk. This may reflect:
<list list-type="bullet">
<list-item>
<p>architectural complexity in files mixing multiple class definitions,</p></list-item>
<list-item>
<p>&#x201C;god classes&#x201D; or utility files that handle diverse responsibilities,</p></list-item>
<list-item>
<p>poor separation of concerns, making security boundaries harder to maintain.</p></list-item>
</list></p>
<p><italic>Cyclomatic Complexity Metrics (SumCyclomatic, SumCyclomaticModified, AvgCyclomaticModified)</italic>.</p>
<p>The prominence of multiple cyclomatic complexity variants confirms established software engineering principles: complex control flow increases the likelihood of vulnerabilities. High cyclomatic complexity indicates:
<list list-type="bullet">
<list-item>
<p>multiple execution paths requiring comprehensive security validation,</p></list-item>
<list-item>
<p>difficulty in reasoning about all possible states and transitions,</p></list-item>
<list-item>
<p>increased cognitive load during code review, making vulnerabilities easier to overlook,</p></list-item>
<list-item>
<p>greater surface area for logic errors that can be exploited.</p></list-item>
</list></p>
<p>This finding is consistent with McCabe&#x2019;s<xref ref-type="fn" rid="fn-10"><sup>10</sup></xref><fn id="fn-10">
<label>10</label>
<p><ext-link ext-link-type="uri" xlink:href="http://www.mccabe.com/pdf/MoreComplexEqualsLessSecure-McCabe.pdf">http://www.mccabe.com/pdf/MoreComplexEqualsLessSecure-McCabe.pdf</ext-link></p>
</fn> research, which demonstrates that complexity directly correlates with security weaknesses.</p>
<p><italic>Extension (Programming Language)</italic>.</p>
<p>The significance of the file extension indicates that security bug prevalence varies by programming language. This likely reflects:
<list list-type="bullet">
<list-item>
<p>Language-specific vulnerability patterns (e.g., memory safety issues in C/C&#x002B;&#x002B;, injection attacks in PHP),</p></list-item>
<list-item>
<p>Different security practices and tool support across language communities,</p></list-item>
<list-item>
<p>Inherent language characteristics (memory management, type systems, etc.).</p></list-item>
</list></p>
<p>The inclusion of <monospace>extension</monospace> among the essential features suggests that programming language characteristics modulate how complexity and size translate into security risk. For instance, a given level of cyclomatic complexity may pose greater security challenges in languages that heavily use dynamic features or reflection. Conversely, statically typed languages with strict structural constraints may better tolerate the same complexity level. While a detailed analysis of language-specific effects is beyond the scope of this paper, the importance of <monospace>extension</monospace> indicates an interesting direction for future work.</p>
<p>The reason for employing multiple feature selection methods is that each technique has different assumptions and biases:
<list list-type="bullet">
<list-item>
<p>Linear/Logistic Regression: assumes linear relationships, sensitive to feature scaling.</p></list-item>
<list-item>
<p>XGBoost: captures non-linear interactions, ranks by prediction contribution.</p></list-item>
<list-item>
<p>F-Score: univariate analysis, evaluates each feature independently.</p></list-item>
<list-item>
<p>Boruta: accounts for feature interactions via Random Forest, robust to noise.</p></list-item>
</list></p>
<p>Developers and security practitioners should prioritize monitoring these six metrics during code review and continuous integration:
<list list-type="simple">
<list-item><label>1.</label><p>Flag files with high executable statement counts for security review.</p></list-item>
<list-item><label>2.</label><p>Scrutinize files defining multiple classes for architectural security issues.</p></list-item>
<list-item><label>3.</label><p>Enforce cyclomatic complexity thresholds in coding standards.</p></list-item>
<list-item><label>4.</label><p>Apply language-specific security analysis tools appropriate to file extensions.</p></list-item>
</list></p>
<p>While these metrics show strong correlation, correlation does not imply causation. High complexity may not directly <italic>cause</italic> vulnerabilities but rather indicates code characteristics that make vulnerabilities more likely or harder to prevent. Additionally, these metrics are most useful in combination; no single metric should be used in isolation for security assessment. Overall, RQ1 provides empirical evidence that a relatively small subset of quality metrics accounts for most of the variance in security bugs. This subset forms the basis for the more fine-grained analyses in RQ2 and for the feature set used in the predictive models of RQ3.</p>
</sec>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>RQ2: Critical Thresholds for Quality Metrics</title>
<p><bold>Research Question:</bold> Do specific values or thresholds of quality metrics indicate a higher probability of security bug occurrence?</p>
<sec id="s5_2_1">
<label>5.2.1</label>
<title>Results</title>
<p>To address RQ2, we analyze the distributions of the five most consistently important metrics identified in RQ1 (excluding file extension, as it is categorical). We conducted a comprehensive statistical analysis comparing the distributions of quality metrics between buggy and non-buggy files. Across all metrics, we observe that the median values for buggy files are substantially higher than those for non-buggy files. For example, the median number of executable statements (<monospace>CountStmtExe</monospace>) is 16 for non-buggy files and 54 for buggy files, that is, more than a factor of three. A similar pattern holds for <monospace>SumCyclomatic</monospace> and <monospace>SumCyclomaticModified</monospace>. Density plots and box plots visually confirm that the distributions of buggy and clean files differ markedly, with buggy files exhibiting heavier tails and greater concentration in higher-value ranges. <xref ref-type="table" rid="table-11">Table 11</xref> presents detailed statistical measurements for five key metrics, computed separately for files without security bugs (305,148 files) and files with security bugs (33,294 files).</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Statistical measurements for key quality metrics.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Measurement</th>
<th>Non-Buggy Files</th>
<th>Buggy Files</th>
<th>Ratio (Buggy/Non-Buggy)</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="8">CountStmtExe</td>
<td>Count</td>
<td>305,149</td>
<td>33,294</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>Mean</td>
<td>96.02</td>
<td>192.41</td>
<td>2.00<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Std Dev</td>
<td>646.47</td>
<td>1045.52</td>
<td>1.62<inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Min</td>
<td>0.0</td>
<td>1.0</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>25% (Q1)</td>
<td>0.0</td>
<td>16.0</td>
<td><inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>

<td>Median</td>
<td>16.0</td>
<td>54.0</td>
<td>3.38<inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>75% (Q3)</td>
<td>70.0</td>
<td>130.0</td>
<td>1.86<inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Max</td>
<td>88,818.0</td>
<td>72,885.0</td>
<td>0.82<inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="8">SumCyclomatic</td>
<td>Count</td>
<td>305,149</td>
<td>33,294</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>Mean</td>
<td>41.98</td>
<td>91.78</td>
<td>2.19<inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Std Dev</td>
<td>523.73</td>
<td>603.12</td>
<td>1.15<inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Min</td>
<td>0.0</td>
<td>1.0</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>25% (Q1)</td>
<td>1.0</td>
<td>8.0</td>
<td>8.00<inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Median</td>
<td>8.0</td>
<td>24.0</td>
<td>3.00<inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>75% (Q3)</td>
<td>29.0</td>
<td>56.9</td>
<td>1.96<inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Max</td>
<td>121,890.0</td>
<td>22,705.0</td>
<td>0.19<inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="8">SumCyclomaticModified</td>
<td>Count</td>
<td>305,149</td>
<td>33,294</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>Mean</td>
<td>40.49</td>
<td>88.34</td>
<td>2.18<inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Std Dev</td>
<td>481.85</td>
<td>587.07</td>
<td>1.22<inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Min</td>
<td>0.0</td>
<td>1.0</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>25% (Q1)</td>
<td>1.0</td>
<td>8.0</td>
<td>8.00<inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Median</td>
<td>8.0</td>
<td>23.0</td>
<td>2.88<inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>75% (Q3)</td>
<td>28.0</td>
<td>55.0</td>
<td>1.96<inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Max</td>
<td>110,753.0</td>
<td>21,483.0</td>
<td>0.19<inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="8">AvgCyclomaticModified</td>
<td>Count</td>
<td>305,149</td>
<td>33,294</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>Mean</td>
<td>4.23</td>
<td>15.25</td>
<td>3.61<inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Std Dev</td>
<td>140.97</td>
<td>425.27</td>
<td>3.02<inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Min</td>
<td>0.0</td>
<td>1.0</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>25% (Q1)</td>
<td>1.0</td>
<td>1.0</td>
<td>1.00<inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Median</td>
<td>1.0</td>
<td>3.0</td>
<td>3.00<inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>75% (Q3)</td>
<td>3.0</td>
<td>5.0</td>
<td>1.67<inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Max</td>
<td>21,181.0</td>
<td>21,161.0</td>
<td>1.00<inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="8">CountDeclClass</td>
<td>Count</td>
<td>305,149</td>
<td>33,294</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>Mean</td>
<td>1.14</td>
<td>2.78</td>
<td>2.44<inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Std Dev</td>
<td>4.60</td>
<td>6.09</td>
<td>1.32<inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Min</td>
<td>0.0</td>
<td>1.0</td>
<td>&#x2013;</td>
</tr>
<tr>

<td>25% (Q1)</td>
<td>0.0</td>
<td>1.2</td>
<td><inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>

<td>Median</td>
<td>0.0</td>
<td>2.0</td>
<td><inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>

<td>75% (Q3)</td>
<td>1.0</td>
<td>2.6</td>
<td>2.60<inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>

<td>Max</td>
<td>335.0</td>
<td>333.0</td>
<td>0.99<inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To assess whether a single ecosystem drives the observed distribution shifts, we replicated the RQ2 analysis across individual ecosystems. For each ecosystem, we compared buggy and non-buggy files using the same non-parametric statistical framework and descriptive statistics. Across all ecosystems, buggy files consistently exhibit higher values for the most influential metrics, particularly executable statement count and cyclomatic complexity measures. While absolute metric values differ due to language conventions and project characteristics, the relative differences between buggy and non-buggy files remain stable. This consistency indicates that the relationship between structural code properties and security bugs is not confined to a specific language or platform. <xref ref-type="table" rid="table-12">Table 12</xref> shows that the median ratios between buggy and non-buggy files exceed one across all ecosystems for the most discriminative metrics. The magnitude of these ratios varies depending on ecosystem characteristics. For instance, ecosystems that rely on low-level or strongly typed languages (e.g., OSS-Fuzz and NuGet) exhibit more pronounced ratios. In contrast, dynamic-language ecosystems such as npm show more moderate differences. These variations reflect language-specific development practices rather than contradictions of the general trend.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Per-ecosystem median ratios for key metrics (buggy vs. non-buggy).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Ecosystem</th>
<th>Primary language</th>
<th>CountStmtExe</th>
<th>SumCyclomatic</th>
<th>AvgCyclomaticModified</th>
</tr>
</thead>
<tbody>
<tr>
<td>GHSA</td>
<td>Mixed</td>
<td>2.65<inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>2.30<inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>2.00<inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Maven</td>
<td>Java</td>
<td>2.68<inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>2.67<inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>2.00<inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>NuGet</td>
<td>.NET (C#)</td>
<td>125.00<inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>25.50<inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>4.00<inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>OSS-Fuzz</td>
<td>C/C&#x002B;&#x002B;</td>
<td>15.67<inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>6.25<inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>4.00<inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Packagist</td>
<td>PHP</td>
<td>2.39<inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>2.00<inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>2.00<inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>PyPI</td>
<td>Python</td>
<td>2.21<inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>1.94<inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>2.00<inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>npm</td>
<td>JavaScript</td>
<td>1.63<inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>1.46<inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>1.00<inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-12fn1" fn-type="other">
<p>Note: Median ratios greater than one across all ecosystems indicate that buggy files are consistently larger and more complex than non-buggy files, despite ecosystem-specific variations in magnitude.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The exceptionally high median ratio observed for NuGet (125<inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula>) in <xref ref-type="table" rid="table-12">Table 12</xref> is primarily driven by a near-zero denominator effect rather than by an intrinsically stronger separation between buggy and non-buggy files. In the NuGet ecosystem, the median value of <monospace>CountDeclClass</monospace> for non-buggy files is zero, while the corresponding median for buggy files is one. As a result, even a minimal absolute difference produces a mathematically large multiplicative ratio. This reflects a simple, interpretable structural characteristic of .NET projects: files that declare at least one class are substantially more likely to be involved in security-fixing commits than files that declare none. Importantly, this pattern is consistent with observations in other ecosystems for <monospace>CountDeclClass</monospace>, albeit with less extreme ratios due to higher non-buggy medians. We therefore caution against interpreting the 125<inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> ratio as a universal amplification effect and emphasize that absolute metric values and threshold-based indicators provide a more robust basis for cross-ecosystem comparison.</p>

<p>It is important to note that ecosystem-specific ratios (<xref ref-type="table" rid="table-12">Table 12</xref>) describe relative distributional shifts within individual ecosystems, whereas effect sizes (<xref ref-type="table" rid="table-13">Table 13</xref>) quantify the global strength of these differences across the full dataset.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Mann-Whitney U test results with effect sizes (RQ2).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Metric</th>
<th>U-statistic</th>
<th><italic>p</italic>-value</th>
<th>Effect size (<inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>)</th>
<th>Magnitude</th>
</tr>
</thead>
<tbody>
<tr>
<td>CountStmtExe</td>
<td><inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mn>6.87</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>9</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>16</mml:mn></mml:mrow></mml:math></inline-formula></td>
<td>0.35</td>
<td>Medium</td>
</tr>
<tr>
<td>SumCyclomatic</td>
<td><inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mn>6.89</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>9</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>16</mml:mn></mml:mrow></mml:math></inline-formula></td>
<td>0.36</td>
<td>Medium</td>
</tr>
<tr>
<td>SumCyclomaticModified</td>
<td><inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mn>6.88</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>9</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>16</mml:mn></mml:mrow></mml:math></inline-formula></td>
<td>0.36</td>
<td>Medium</td>
</tr>
<tr>
<td>AvgCyclomaticModified</td>
<td><inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mn>7.10</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>9</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>16</mml:mn></mml:mrow></mml:math></inline-formula></td>
<td>0.40</td>
<td>Medium</td>
</tr>
<tr>
<td>CountDeclClass</td>
<td><inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mn>8.32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>9</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>16</mml:mn></mml:mrow></mml:math></inline-formula></td>
<td>0.64</td>
<td>Large</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-13fn1" fn-type="other">
<p>Note: Effect size is reported as rank-biserial correlation (<inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>). Guidelines: <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>&#x2248;</mml:mo><mml:mn>0.1</mml:mn></mml:math></inline-formula> (small), <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mo>&#x2248;</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula> (medium), <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mo>&#x2265;</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula> (large).</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>To further assess the generalizability of our findings, we examined ecosystem-specific patterns using two complementary perspectives. First, distribution-based analyses confirmed that key metrics consistently differentiate buggy from non-buggy files across all ecosystems. Second, we assessed the robustness of the global prediction model when applied to ecosystem-specific subsets, reflecting a realistic deployment scenario in which a single trained model is used for heterogeneous projects.</p>
<p>While predictive performance may vary due to ecosystem-specific characteristics such as language diversity and class imbalance, the model maintains robust performance across ecosystems. These observations indicate that the identified relationships between software quality metrics and security bugs generalize beyond individual languages or platforms, and suggest that ecosystem-aware calibration may further enhance practical effectiveness.</p>
<p><bold>Key observation:</bold> Across all five metrics, the median values for buggy files are consistently between 2.88<inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> and 3.38<inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> higher than for non-buggy files (excluding <monospace>CountDeclClass</monospace> where the ratio is undefined due to zero median for non-buggy files).</p>
<p><xref ref-type="fig" rid="fig-2">Figs. 2</xref>&#x2013;<xref ref-type="fig" rid="fig-6">6</xref> present density plots comparing the distribution of metric values between buggy (red curve) and non-buggy (green curve) files.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Class-level complexity and security bugs (CountDeclClass).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Aggregate control-flow complexity and security bugs (SumCyclomatic).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-3.tif"/>
</fig><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Modified cyclomatic complexity and security bugs (SumCyclomaticModified).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Average function complexity and security bugs (AvgCyclomaticModified).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-5.tif"/>
</fig><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Executable code size and security bugs (CountStmtExe).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-6.tif"/>
</fig>
<p>The density plot reveals:
<list list-type="bullet">
<list-item>
<p>Non-buggy files: strong concentration at 0-1 class declarations, with a sharp peak near zero.</p></list-item>
<list-item>
<p>Buggy files: broader distribution extending to higher values, peak around 1-2 classes.</p></list-item>
</list></p>
<p>This suggests that files defining multiple classes show substantially elevated security risk.</p>
<p>The distribution for <monospace>SumCyclomatic</monospace> (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>) shows:
<list list-type="bullet">
<list-item>
<p>Non-buggy files: heavily right-skewed, most files under 50 total complexity.</p></list-item>
<list-item>
<p>Buggy files: similar shape but shifted right, extending to higher complexity values.</p></list-item>
</list></p>
<p>There is a clear separation: buggy files consistently exhibit higher cyclomatic complexity sums.</p>
<p>The pattern for <monospace>SumCyclomaticModified</monospace> (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>) is similar to that of <monospace>SumCyclomatic</monospace>:
<list list-type="bullet">
<list-item>
<p>Both distributions are right-skewed;</p></list-item>
<list-item>
<p>Buggy files show a notable shift toward higher complexity values;</p></list-item>
<list-item>
<p>There is substantial overlap but distinct peaks.</p></list-item>
</list></p>
<p>For <monospace>AvgCyclomaticModified</monospace> (<xref ref-type="fig" rid="fig-5">Fig. 5</xref>), we observe a distinctive pattern:
<list list-type="bullet">
<list-item>
<p>Non-buggy files: extreme concentration at very low values (near 1).</p></list-item>
<list-item>
<p>Buggy files: broader distribution with heavier tail toward higher complexity.The median difference (1.0 vs. 3.0) indicates that buggy files have a complexity that is 3 times the average.</p></list-item>
</list></p>
<p>For <monospace>CountStmtExe</monospace> (<xref ref-type="fig" rid="fig-6">Fig. 6</xref>), we observe the most pronounced separation:
<list list-type="bullet">
<list-item>
<p>Non-buggy files: sharp peak near zero, rapid decline.</p></list-item>
<list-item>
<p>Buggy files: broader distribution, substantial mass in the 50&#x2013;200 range.</p></list-item>
</list></p>
<p>The median difference (16 vs. 54 executable statements) is substantial.</p>
<p>Box plots provide a complementary view. <xref ref-type="fig" rid="fig-7">Figs. 7</xref>&#x2013;<xref ref-type="fig" rid="fig-11">11</xref> present box plots for the five metrics, showing quartiles and outliers for both classes.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Executable statement count as a security risk indicator (CountStmtExe).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Control-flow complexity distribution in buggy vs. non-buggy files (SumCyclomatic).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-8.tif"/>
</fig><fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Class declaration count and security-prone code (CountDeclClass).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-9.tif"/>
</fig><fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Average function complexity in security-prone files (AvgCyclomaticModified).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-10.tif"/>
</fig><fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Aggregate modified cyclomatic complexity and security bugs (SumCyclomaticModified).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_77139-fig-11.tif"/>
</fig>
<p>Key observations for <monospace>CountStmtExe</monospace> (<xref ref-type="fig" rid="fig-7">Fig. 7</xref>) include:
<list list-type="bullet">
<list-item>
<p>Q1 (25%): non-buggy &#x003D; 0, buggy &#x003D; 16 (substantial baseline difference).</p></list-item>
<list-item>
<p>Median: non-buggy &#x003D; 16, buggy &#x003D; 54 (3.38<inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> difference).</p></list-item>
<list-item>
<p>Q3 (75%): non-buggy &#x003D; 70, buggy &#x003D; 130 (1.86<inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> difference).</p></list-item>
</list></p>
<p>Both groups have extreme outliers, but buggy files show higher typical values.</p>
<p>For <monospace>SumCyclomatic</monospace> (<xref ref-type="fig" rid="fig-8">Fig. 8</xref>), interquartile ranges show clear separation, with a median difference of 3.00<inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> (8 vs. 24). Buggy files consistently exhibit higher complexity across all quartiles.</p>
<p>For <monospace>CountDeclClass</monospace> (<xref ref-type="fig" rid="fig-9">Fig. 9</xref>), we observe a distinctive pattern:
<list list-type="bullet">
<list-item>
<p>Non-buggy files: median of 0 (most files define no classes).</p></list-item>
<list-item>
<p>Buggy files: median of 2.0 (typical buggy file defines multiple classes).</p></list-item>
</list></p>
<p>This suggests a clear threshold: files defining two or more classes warrant security scrutiny.</p>
<p>For <monospace>AvgCyclomaticModified</monospace> (<xref ref-type="fig" rid="fig-10">Fig. 10</xref>):
<list list-type="bullet">
<list-item>
<p>Non-buggy files: compact distribution, most files have average complexity near 1.</p></list-item>
<list-item>
<p>Buggy files: wider distribution, median of 3.</p></list-item>
</list></p>
<p>This suggests that average function complexity <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mo>&#x2265;</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> indicates elevated risk.</p>
<p>For <monospace>SumCyclomaticModified</monospace> (<xref ref-type="fig" rid="fig-11">Fig. 11</xref>), the pattern is similar to <monospace>SumCyclomatic</monospace>, with clear separation in medians and quartiles. The 2.88<inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> median ratio confirms a substantial difference between buggy and non-buggy files.</p>
<p>To confirm that the observed differences between buggy and non-buggy files are statistically significant rather than artifacts of random sampling, we conducted Mann-Whitney U tests for each metric. All examined metrics yield highly significant <italic>p</italic>-values, far below the conventional significance threshold of 0.05, leading us to reject the null hypothesis that buggy and non-buggy files follow identical distributions. While statistical significance confirms the existence of differences, effect size analysis provides a more practical estimate of their magnitude. As reported in <xref ref-type="table" rid="table-13">Table 13</xref>, medium-to-large effect sizes are observed for the most influential metrics, indicating that the detected differences are not only statistically significant but also meaningful in practice. In particular, <italic>CountDeclClass</italic> exhibits a large effect size, highlighting the importance of structural class-level complexity as a strong indicator of security-prone code. Importantly, these distributional differences remain consistent across ecosystems. Although absolute metric values vary depending on language conventions and project characteristics, the relative shifts between buggy and non-buggy files persist across all ecosystems, supporting the generalizability of our findings beyond a single platform or language.</p>

<p><bold>Test interpretation:</bold>
<list list-type="bullet">
<list-item>
<p>Null Hypothesis (<inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>): metric distributions are identical for buggy and non-buggy files.</p></list-item>
<list-item>
<p>Alternative Hypothesis (<inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>): distributions differ significantly.</p></list-item>
<list-item>
<p>Result: all <italic>p</italic>-values <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>, providing strong evidence to reject <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</p></list-item>
</list></p>
<p>All five metrics show statistically significant differences between buggy and non-buggy files, confirming that the observed patterns are not due to chance. Based on the median and Q3 values from buggy files, we propose threshold guidelines for security risk assessment, summarized in <xref ref-type="table" rid="table-14">Table 14</xref>.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Proposed threshold guidelines for security risk assessment.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Warning Threshold</th>
<th>Critical Threshold</th>
<th>Rationale</th>
</tr>
</thead>
<tbody>
<tr>
<td>CountStmtExe</td>
<td><inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>50</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>130</mml:mn></mml:math></inline-formula></td>
<td>Median and Q3 for buggy files</td>
</tr>
<tr>
<td>SumCyclomatic</td>
<td><inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>24</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>57</mml:mn></mml:math></inline-formula></td>
<td>Median and Q3 for buggy files</td>
</tr>
<tr>
<td>SumCyclomaticModified</td>
<td><inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>23</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>55</mml:mn></mml:math></inline-formula></td>
<td>Median and Q3 for buggy files</td>
</tr>
<tr>
<td>AvgCyclomaticModified</td>
<td><inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>5</mml:mn></mml:math></inline-formula></td>
<td>Median and Q3 for buggy files</td>
</tr>
<tr>
<td>CountDeclClass</td>
<td><inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>2</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mo>&#x003E;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>3</mml:mn></mml:math></inline-formula></td>
<td>Median and Q3 for buggy files</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Threshold application:</bold>
<list list-type="bullet">
<list-item>
<p><bold>Warning Threshold:</bold> if the file exceeds the median of buggy files, schedule for security review.</p></list-item>
<list-item>
<p><bold>Critical Threshold:</bold> if the file exceeds the third quartile (Q3) of buggy files, prioritize for immediate security analysis.</p></list-item>
</list></p>
</sec>
<sec id="s5_2_2">
<label>5.2.2</label>
<title>Discussion: Towards Practical Threshold Validation</title>
<p>While the proposed thresholds are derived from statistically significant distributional differences, their primary objective is not to establish universal or prescriptive rules, but to provide empirically grounded risk indicators that can be adapted to diverse development contexts. In this sense, validation should be understood not only in terms of numerical performance metrics but also in terms of interpretability, stability across ecosystems, and practical applicability.</p>
<p>First, the thresholds are grounded in consistent patterns observed across multiple ecosystems and programming languages, suggesting that they capture a robust empirical signal rather than project-specific artifacts. Second, their formulation is intentionally simple and transparent (e.g., based on relative differences, such as the observed &#x201C;3<inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> rule&#x201D;), which facilitates reproducibility and independent re-evaluation across datasets or organizational settings. This is further supported by the per-ecosystem median ratios reported in <xref ref-type="table" rid="table-12">Table 12</xref>, which remain above one across ecosystems for the most discriminative metrics. Rather than relying on one-size-fits-all prescriptions, our approach provides empirically derived baseline thresholds that can be recalibrated locally (e.g., by language, project history, or domain constraints).</p>

<p>From a practical perspective, these thresholds are best validated through incremental adoption. Organizations should initially deploy them as non-blocking indicators for exploratory analysis, security awareness, or review prioritization. Subsequently, teams can adjust thresholds based on project history, language characteristics, and observed false-positive rates. In addition, combining multiple metrics (e.g., size and complexity) can improve specificity and reduce reliance on any single indicator.</p>
<p>Finally, while our study provides empirical evidence supporting the usefulness of such thresholds at scale, prospective validation within active development pipelines and organization-specific calibration remain essential steps for operational deployment. We therefore view the proposed thresholds as a reproducible empirical baseline rather than definitive security criteria, enabling informed adaptation rather than one-size-fits-all enforcement.</p>
<p>The statistical analyses in RQ2 provide strong evidence that files involved in security bugs tend to exhibit significantly larger and more complex code structures than clean files. In particular, the fact that median values for buggy files are often around three times higher than those for non-buggy files suggests that high complexity and size may be interpreted as risk factors for security bugs (i.e., security-relevant code flaws), rather than as direct evidence of exploitability. From a practical standpoint, these results support the idea of defining empirical thresholds for specific metrics above which the risk of security bugs increases significantly. For example, values of <monospace>CountStmtExe</monospace> or <monospace>SumCyclomatic</monospace> exceeding the upper quartile observed in buggy files could serve as triggers for additional code review or security testing. Such thresholds should not be interpreted as strict rules, but rather as indicators that a file deserves closer scrutiny. Organizations can operationalize these findings by:
<list list-type="simple">
<list-item><label>1.</label><p><bold>Automated code review triggers:</bold> configure continuous integration systems to flag files exceeding warning thresholds for mandatory security review before merge.</p></list-item>
<list-item><label>2.</label><p><bold>Technical debt prioritization:</bold> when refactoring, prioritize files exceeding critical thresholds, as they represent both maintainability and security liabilities.</p></list-item>
<list-item><label>3.</label><p><bold>Security testing focus:</bold> concentrate security testing resources (fuzzing, penetration testing, formal verification) on files exceeding thresholds, maximizing return on security investment.</p></list-item>
</list></p>
<p>Our &#x201C;3<inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> rule&#x201D; finding aligns with and extends previous research [<xref ref-type="bibr" rid="ref-52">52</xref>]:
<list list-type="bullet">
<list-item>
<p>Shin and Williams [<xref ref-type="bibr" rid="ref-11">11</xref>] showed that complexity metrics differentiate vulnerable code, but did not quantify specific thresholds.</p></list-item>
<list-item>
<p>Alenezi and Zarour [<xref ref-type="bibr" rid="ref-53">53</xref>] established that higher complexity correlates with security weaknesses.</p></list-item>
</list></p>
<p>Our contribution is to provide specific quantitative threshold guidelines derived from a large-scale, multi-ecosystem empirical analysis.</p>
<p>While we propose general thresholds based on our dataset, organizations should consider calibrating them to their specific context:
<list list-type="bullet">
<list-item>
<p><bold>Language-specific:</bold> different languages exhibit different baseline complexity profiles.</p></list-item>
<list-item>
<p><bold>Project-specific:</bold> legacy systems may have different baselines than greenfield projects.</p></list-item>
<list-item>
<p><bold>Domain-specific:</bold> security-critical domains (e.g., cryptography or authentication) may warrant stricter thresholds.</p></list-item>
</list></p>
<p>Some limitations and cautions we note are:
<list list-type="simple">
<list-item><label>1.</label><p><bold>Thresholds are guidelines, not absolutes.</bold> Exceeding a threshold indicates elevated risk but does not guarantee the presence of a vulnerability. Conversely, files below thresholds may still contain security bugs, particularly those arising from subtle logic or design errors.</p></list-item>
<list-item><label>2.</label><p><bold>Avoidance of metric gaming.</bold> Developers should not artificially manipulate metrics (e.g., splitting functions solely to reduce cyclomatic complexity) without genuine improvements to code clarity and security. Metric improvements should result from meaningful refactoring.</p></list-item>
<list-item><label>3.</label><p><bold>Ecological fallacy.</bold> Thresholds are derived from population-level statistics; individual files or projects may legitimately deviate without indicating increased security risk.</p></list-item>
<list-item><label>4.</label><p><bold>Evolution of threats.</bold> As attack techniques evolve, the relationship between metrics and security bugs may change. Periodic re-evaluation against updated vulnerability data is therefore recommended.</p></list-item>
</list></p>
<p>These thresholds represent an initial empirical baseline. Future work should:
<list list-type="bullet">
<list-item>
<p>develop language-specific threshold recommendations,</p></list-item>
<list-item>
<p>investigate threshold combinations (e.g., high complexity <italic>and</italic> high executable statement count),</p></list-item>
<list-item>
<p>validate thresholds prospectively on new projects,</p></list-item>
<list-item>
<p>explore dynamic threshold adjustment based on project history.</p></list-item>
</list></p>
<p>It is important to emphasize that high complexity and size do not cause security bugs; they correlate with underlying development and maintenance challenges such as reduced modularity, rushed development, or insufficient testing. Nevertheless, these metrics offer a lightweight, interpretable proxy that can be easily computed and monitored throughout the development process. RQ2 thus complements RQ1 by showing not only which metrics matter, but also how their values differ in the presence of security bugs.</p>
</sec>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>RQ3: Security Bug Prediction Performance</title>
<p><bold>Research Question:</bold> To what extent can machine learning models predict security bugs based on software quality metrics?</p>
<sec id="s5_3_1">
<label>5.3.1</label>
<title>Results</title>
<p>The classification results for the eleven machine learning models correspond to the average performance over the five folds of the time-series cross-validation, using the combined dataset from all ecosystems. At first glance, most models achieve high accuracy, often above 0.90. In contrast, given the strong class imbalance, accuracy alone is not informative. We therefore focus on Recall, F1-score, MCC, and ROC-AUC, which provide a more nuanced view of performance on the minority class. <xref ref-type="table" rid="table-15">Table 15</xref> presents the average performance across all folds for all ecosystems combined.</p>
<table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Classification performance across all algorithms.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Algorithm</th>
<th>Accuracy</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-Score</th>
<th>MCC</th>
<th>ROC-AUC</th>
</tr>
</thead>
<tbody>
<tr>
<td>XGBoost</td>
<td>0.98</td>
<td>0.95</td>
<td>0.82</td>
<td>0.88</td>
<td>0.87</td>
<td>0.91</td>
</tr>
<tr>
<td>Random Forest</td>
<td>0.97</td>
<td>0.89</td>
<td>0.82</td>
<td>0.85</td>
<td>0.84</td>
<td>0.90</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>0.97</td>
<td>0.90</td>
<td>0.79</td>
<td>0.84</td>
<td>0.83</td>
<td>0.89</td>
</tr>
<tr>
<td>Multi-layer Perceptron</td>
<td>0.97</td>
<td>0.90</td>
<td>0.77</td>
<td>0.83</td>
<td>0.82</td>
<td>0.88</td>
</tr>
<tr>
<td>Decision Tree</td>
<td>0.96</td>
<td>0.78</td>
<td>0.83</td>
<td>0.80</td>
<td>0.78</td>
<td>0.90</td>
</tr>
<tr>
<td><inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest Neighbors</td>
<td>0.94</td>
<td>0.77</td>
<td>0.48</td>
<td>0.59</td>
<td>0.58</td>
<td>0.73</td>
</tr>
<tr>
<td>Logistic Regression</td>
<td>0.90</td>
<td>0.31</td>
<td>0.03</td>
<td>0.05</td>
<td>0.06</td>
<td>0.51</td>
</tr>
<tr>
<td>Linear SVM</td>
<td>0.86</td>
<td>0.35</td>
<td>0.30</td>
<td>0.52</td>
<td>0.25</td>
<td>0.71</td>
</tr>
<tr>
<td>Gaussian SVM</td>
<td>0.87</td>
<td>0.29</td>
<td>0.36</td>
<td>0.47</td>
<td>0.22</td>
<td>0.69</td>
</tr>
<tr>
<td>Na&#x00EF;ve Bayes</td>
<td>0.84</td>
<td>0.16</td>
<td>0.15</td>
<td>0.15</td>
<td>0.07</td>
<td>0.53</td>
</tr>
<tr>
<td>QDA</td>
<td>0.62</td>
<td>0.15</td>
<td>0.61</td>
<td>0.24</td>
<td>0.15</td>
<td>0.62</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Among the evaluated models, tree-based methods and ensemble approaches such as XGBoost, Random Forests, Decision Trees, and AdaBoost clearly stand out. XGBoost achieves a Recall of about 0.82, a high F1-score, an MCC of about 0.87, and an ROC-AUC of about 0.91. Random Forests and AdaBoost achieve similar performance, with Recall values of 0.79&#x2013;0.82 and ROC-AUC values of 0.90. Decision Trees also perform well, with high Recall and ROC-AUC, although their MCC is slightly lower than that of the ensemble models. In contrast, models such as Na&#x00EF;ve Bayes, QDA, and Logistic Regression perform significantly worse in terms of Recall and MCC. Na&#x00EF;ve Bayes, for instance, achieves a low Recall, indicating that it fails to capture the complex decision boundaries needed to distinguish buggy from non-buggy files. The Support Vector Machine models exhibit intermediate behavior, with relatively high accuracy but modest Recall and F1-scores.</p>
<p>While descriptive metrics indicate that XGBoost achieves the best overall performance, we further examined whether these differences are statistically significant across models. To this end, we applied the Friedman test followed by a Nemenyi post-hoc analysis, as recommended for comparing multiple classifiers over repeated cross-validation folds (<xref ref-type="table" rid="table-16">Table 16</xref>).</p>
<table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Nemenyi post-hoc test ranking for ROC-AUC performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Algorithm</th>
<th>Mean Rank</th>
<th>Statistical Group</th>
</tr>
</thead>
<tbody>
<tr>
<td>XGBoost</td>
<td>1.2</td>
<td>A</td>
</tr>
<tr>
<td>Random Forest</td>
<td>1.8</td>
<td>A</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>2.4</td>
<td>A</td>
</tr>
<tr>
<td>Multi-layer Perceptron (MLP)</td>
<td>3.1</td>
<td>B</td>
</tr>
<tr>
<td>Decision Tree</td>
<td>3.7</td>
<td>B</td>
</tr>
<tr>
<td><inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest Neighbors (KNN)</td>
<td>5.9</td>
<td>C</td>
</tr>
<tr>
<td>Linear SVM</td>
<td>6.8</td>
<td>D</td>
</tr>
<tr>
<td>RBF SVM</td>
<td>7.2</td>
<td>D</td>
</tr>
<tr>
<td>Logistic Regression</td>
<td>8.5</td>
<td>E</td>
</tr>
<tr>
<td>Na&#x00EF;ve Bayes</td>
<td>9.1</td>
<td>E</td>
</tr>
<tr>
<td>Quadratic Discriminant Analysis (QDA)</td>
<td>10.3</td>
<td>F</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-16fn1" fn-type="other">
<p>Note: Lower ranks indicate better performance. Models sharing the same letter are not significantly different at <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The Friedman test rejected the null hypothesis that all classifiers perform equally (<inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msup><mml:mi>&#x03C7;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>52.7</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>d</mml:mi><mml:mi>f</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>), confirming the presence of statistically significant performance differences among the evaluated algorithms. The subsequent Nemenyi post-hoc test revealed that XGBoost significantly outperforms most non-ensemble models, including Logistic Regression, Na&#x00EF;ve Bayes, <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mi>k</mml:mi></mml:math></inline-formula>-Nearest Neighbors, Support Vector Machines, and single Decision Trees (<inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>).</p>
<p>No statistically significant difference was observed between XGBoost, Random Forest, and AdaBoost. This result indicates that ensemble tree-based methods, whether based on boosting or bagging, form a statistically homogeneous group with consistently superior performance for security bug prediction using static quality metrics. These findings confirm that the observed advantage of XGBoost is not due to random variation across folds but rather reflects a robust, reproducible performance pattern. From a practical standpoint, these results suggest that the choice among ensemble tree-based models may be guided by secondary considerations such as training time, interpretability, or deployment constraints, rather than raw predictive performance alone.</p>
</sec>
<sec id="s5_3_2">
<label>5.3.2</label>
<title>Discussion</title>
<p>The results for RQ3 demonstrate that machine learning models can effectively leverage software quality metrics to predict files likely to contain security bugs. In particular, tree-based ensemble methods such as XGBoost and Random Forests offer a favorable trade-off between high Recall, strong discriminative power (as reflected in ROC-AUC), and robust overall correlation with the correct labels (MCC).</p>
<p>In a security context, missing an actual vulnerability (false negative) is generally more costly than raising a false alarm (false positive). For this reason, Recall is a critical metric. The observed Recall values around 0.8 for the best models indicate that a substantial fraction of buggy files can be detected based on quality metrics alone. At the same time, the relatively high Precision and MCC scores suggest that the models do not simply label most files as buggy.</p>
<p>The relatively poor performance of generative models such as Na&#x00EF;ve Bayes or QDA is expected given the complex, non-linear interactions between metrics observed in RQ1 and RQ2. These models rely on strong distributional assumptions (for example, conditional independence or Gaussianity) that are unlikely to hold in our data. In contrast, ensemble methods are better suited to capturing heterogeneous, high-order patterns.</p>
<p>Direct comparisons between our machine learning approach and traditional static analysis or typical static application security testing (SAST) tools such as SpotBugs, SonarQube, or Infer are inherently challenging, as these tools operate at different granularities and pursue distinct objectives. Static analyzers typically rely on predefined rules to detect specific vulnerability patterns and report line-level warnings. In contrast, our approach produces file-level risk predictions based on structural and complexity metrics. We emphasize that this comparison is intended to be conceptual rather than quantitative, given the fundamentally different objectives and outputs of machine-learning-based prediction and rule-based static analysis.</p>
<p>Rather than replacing static analyzers, our method is best viewed as complementary. Empirical studies report that static analysis tools often achieve moderate recall for real-world vulnerabilities, with performance strongly dependent on the vulnerability category and configuration, and frequently suffer from high false-positive rates that limit adoption in practice. In contrast, our model achieves high recall while maintaining substantial precision, making it well-suited as an upstream prioritization mechanism. In practical security workflows, the proposed approach can serve as a high-recall pre-filter that identifies security-prone files early in the development lifecycle. These files can then be subjected to more expensive static or dynamic analyses or prioritized for manual review. Such a layered strategy leverages the strengths of both paradigms: the broad coverage of metric-based machine learning and the precise diagnostics of rule-based static analyzers.</p>
<p>Collectively, the results of RQ3 indicate that quality-metric-based prediction can be a viable component of a broader security assurance process. For example, the models could be integrated into continuous integration pipelines to flag potentially vulnerable files early, thereby directing developers&#x2019; and security analysts&#x2019; attention to the most critical parts of the codebase. Possible production deployment and integration strategies include:
<list list-type="bullet">
<list-item>
<p><bold>Pre-commit hooks:</bold> flag files exceeding thresholds or predicted as risky.</p></list-item>
<list-item>
<p><bold>CI/CD gates:</bold> block merges for high-risk predictions pending security review, using tiered risk levels to mitigate alert fatigue and false positives.</p></list-item>
<list-item>
<p><bold>Prioritized review:</bold> rank files by prediction probability for efficient allocation of review and testing resources.</p></list-item>
<list-item>
<p><bold>Feedback loop:</bold> incorporate confirmed vulnerabilities to retrain and recalibrate models periodically.</p></list-item>
</list></p>
<p>Organizations can adapt prediction thresholds based on risk appetite:
<list list-type="bullet">
<list-item>
<p><bold>Conservative (high Recall):</bold> lower threshold, more false alarms, fewer missed bugs.</p></list-item>
<list-item>
<p><bold>Balanced (default):</bold> threshold optimizing F1-score.</p></list-item>
<list-item>
<p><bold>Aggressive (high Precision):</bold> higher threshold, fewer false alarms, more missed bugs.</p></list-item>
</list></p>
<p>Our validation strategy, using time-series cross-validation, is critical for realistic performance estimation. Our reported metrics reflect the expected real-world performance of predicting future security bugs from historical patterns. Traditional random <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mi>k</mml:mi></mml:math></inline-formula>-fold cross-validation would overestimate performance by:
<list list-type="bullet">
<list-item>
<p>training on &#x201C;future&#x201D; data to predict &#x201C;past&#x201D; bugs,</p></list-item>
<list-item>
<p>violating temporal causality,</p></list-item>
<list-item>
<p>failing to reflect real deployment scenarios.</p></list-item>
</list></p>
<p>The current limitations of our study include:
<list list-type="simple">
<list-item><label>1.</label><p><bold>Static metrics only:</bold> we use code-level metrics, not runtime behavior or security-specific patterns.</p></list-item>
<list-item><label>2.</label><p><bold>Ecosystem aggregation:</bold> results combine diverse languages/frameworks; language-specific models might perform better.</p></list-item>
<list-item><label>3.</label><p><bold>Binary classification:</bold> we predict presence/absence, not vulnerability type or severity.</p></list-item>
<list-item><label>4.</label><p><bold>Historical data:</bold> trained on past bugs; new vulnerability classes may not be detected.</p></list-item>
</list></p>
<p>Future enhancements should focus on:
<list list-type="bullet">
<list-item>
<p><bold>Deep learning:</bold> exploring code embeddings (e.g., CodeBERT) combined with metrics.</p></list-item>
<list-item>
<p><bold>Process metrics:</bold> incorporating developer experience, code churn, and review quality.</p></list-item>
<list-item>
<p><bold>Semantic analysis:</bold> adding abstract syntax tree features and control flow graphs.</p></list-item>
<list-item>
<p><bold>Multi-task learning:</bold> simultaneously predicting bug type and severity.</p></list-item>
<list-item>
<p><bold>Active learning:</bold> prioritizing manual review of uncertain predictions to improve models.</p></list-item>
<list-item>
<p><bold>Cross-project prediction:</bold> developing models that generalize across projects with limited training data.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Summary of Findings</title>
<p>The three research questions addressed in this study provide complementary perspectives on the relationship between software quality metrics and security bugs. RQ1 shows that a small subset of metrics related to size and cyclomatic complexity is consistently associated with the presence of vulnerabilities. RQ2 demonstrates that the distributions of these metrics differ significantly between buggy and non-buggy files, with buggy files exhibiting substantially higher median and upper-quartile values. RQ3 establishes that machine learning models, particularly tree-based ensemble methods, can leverage these metrics to achieve high Recall and discriminative performance in predicting security bugs. <xref ref-type="table" rid="table-17">Table 17</xref> summarizes how each research question is addressed:</p>
<table-wrap id="table-17">
<label>Table 17</label>
<caption>
<title>Summary of how each research question is addressed.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>RQ</th>
<th>Data Used</th>
<th>Primary Methods</th>
<th>Key Outputs</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">RQ1</td>
<td>All 338,442 files</td>
<td>Five feature importance techniques</td>
<td>Ranked metrics</td>
</tr>
<tr>

<td>25 metrics</td>
<td>Pearson correlation</td>
<td>Consensus set of six core metrics</td>
</tr>
<tr>
<td rowspan="3">RQ2</td>
<td>All 338,442 files</td>
<td>Descriptive statistics</td>
<td>&#x201C;3<inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> rule&#x201D;</td>
</tr>
<tr>

<td>Five core</td>
<td>Visualizations</td>
<td>Warning and critical thresholds for</td>
</tr>
<tr>

<td>metrics</td>
<td>Mann-Whitney U test</td>
<td>key metrics</td>
</tr>
<tr>
<td rowspan="3">RQ3</td>
<td>All 338,442 files</td>
<td>Eleven ML algorithms</td>
<td>Performance comparison</td>
</tr>
<tr>

<td>25 metrics</td>
<td>Time-series cross-validation</td>
<td>Optimal model identified</td>
</tr>
<tr>
<td/>

<td>Six evaluation metrics</td>
<td>(XGBoost)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Overall, these findings suggest that static quality metrics, which are relatively cheap to compute, can play a meaningful role in supporting security assurance activities. They can help identify code regions that warrant additional review, serve as features for automated prediction models, and inform the definition of metric-based thresholds in secure development practices.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Synthesis and Implications for Research and Practice</title>
<p>Our findings align with a substantial body of prior work showing that code size and complexity are associated with defect proneness and security risk. Previous studies in vulnerability prediction have reported similar trends, though often at a smaller scale or with inconsistent conclusions [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>]. By leveraging a large, multi-ecosystem dataset and multiple complementary analysis techniques, our study confirms and quantifies these relationships in a security-specific context.</p>
<p>In contrast to studies reporting weak or inconclusive correlations between complexity metrics and vulnerabilities, such as Alves et al. [<xref ref-type="bibr" rid="ref-12">12</xref>], our results reveal consistent, statistically significant patterns across ecosystems. This discrepancy is likely attributable to differences in dataset size, labeling precision, and analysis methodology. In particular, our use of multiple feature importance methods and effect size analysis reduces dependence on any single modeling assumption.</p>
<p>Our findings also refine prior observations. While earlier work emphasized specific metrics such as nesting depth or function-level complexity, we find that aggregate file-level metrics, especially executable statement count and cyclomatic complexity, are more robust indicators across diverse languages and ecosystems. This suggests that overall structural burden, rather than isolated local complexity, plays a key role in security risk.</p>
<p>From a practical perspective, our study provides actionable guidance for practitioners by proposing empirically grounded threshold ranges and demonstrating that lightweight machine learning models can achieve high recall. These results support integrating metric-based risk indicators as early-warning mechanisms in CI/CD pipelines, where they can guide review prioritization and security testing efforts. In practice, such integration should be incremental and non-disruptive. To mitigate alert fatigue and false positives, prediction scores and threshold exceedance can be used to define relative risk levels rather than complex blocking rules. In this way, the proposed approach complements existing security practices by directing expert attention to higher-risk code regions without replacing established review or analysis processes.</p>
<p>For researchers, our work highlights the importance of multi-ecosystem datasets, transparent labeling strategies, and effect-size reporting in vulnerability-prediction studies. The identified core metrics provide a focused baseline for future research, while the remaining limitations motivate further exploration of semantic and process-level factors.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Threats to Validity</title>
<p>Before concluding, we discuss the main threats to validity that may affect the interpretation and generalization of our findings. This section outlines potential threats to the validity of our empirical study and the measures adopted to mitigate them. Following standard guidelines in empirical software engineering, we consider internal, external, construct, and conclusion validity and conclude with a note on reproducibility.</p>
<sec id="s6_1">
<label>6.1</label>
<title>Internal Validity</title>
<p>Internal validity concerns whether the observed relationships between quality metrics and security bugs are genuine associations rather than artifacts of data collection or analysis. We address the following concerns:</p>
<p><bold>Labeling accuracy:</bold> Our labeling strategy assumes that files modified in security-fixing commits are security-relevant. This may introduce noise, as some files may be changed for non-security reasons (e.g., refactoring or formatting). To mitigate this risk, we applied automated filtering to exclude non-source files and cosmetic-only changes (<xref ref-type="sec" rid="s3_3_1">Section 3.3.1</xref>) and conducted manual validation on a random sample of commits. This inspection confirmed that 94.2% of the sampled cases clearly corresponded to security fixes, while 4.0% were borderline and 1.8% were incorrectly labeled and removed before analysis. This low estimated residual noise rate is therefore unlikely to affect the observed large-scale statistical patterns materially.</p>
<p><bold>Confounding factors:</bold> The observed associations between complexity and security bugs may be influenced by confounding variables such as developer experience, code age, or review practices. While we cannot fully disentangle causal mechanisms in an observational study, our objective is predictive and risk-oriented rather than causal. The consistency of results across ecosystems and validation strategies supports their robustness for prioritization purposes.</p>
<p><bold>Temporal effects:</bold> Our dataset spans multiple years, during which development practices and security awareness have evolved. We mitigate temporal confounding by adopting time-series cross-validation, ensuring that models are constantly evaluated on future data relative to training data.</p>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>External Validity</title>
<p>External validity relates to the generalizability of our findings beyond the studied context. Our dataset covers seven major open-source ecosystems, a wide range of programming languages, and over 338,000 file-level observations, which strengthens generalizability to large, actively maintained open-source projects. Nevertheless, the results may not directly transfer to closed-source, safety-critical, or highly regulated domains (e.g., avionics or medical software), where development constraints differ substantially.</p>
<p>In addition, reliance on the OSV database may introduce reporting biases, as well-maintained projects are more likely to disclose vulnerabilities. While these limits representativeness, they reflect real-world vulnerability-discovery processes. We therefore encourage organizations to validate and calibrate thresholds using their own historical security data.</p>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Construct Validity</title>
<p>Construct validity assesses whether the selected measurements accurately represent the studied concepts:</p>
<p><bold>Security bug definition:</bold> In this study, the term <italic>security bug</italic> refers to security-relevant code flaws that were subsequently fixed through security-related commits. While many correspond to documented vulnerabilities, others represent latent weaknesses whose exploitability depends on context. This broader definition aligns with proactive security assessment but differs from studies focusing exclusively on confirmed exploits. As highlighted in recent analyses of vulnerability datasets [<xref ref-type="bibr" rid="ref-54">54</xref>,<xref ref-type="bibr" rid="ref-55">55</xref>], the quality and consistency of labeling for such data remain ongoing challenges in the field.</p>
<p><bold>Metric coverage:</bold> The selected static metrics capture code size and structural complexity, which are widely used proxies for maintainability and cognitive load. However, they do not capture semantic properties such as authentication logic correctness, data-flow violations, or API misuse. As such, our approach complements rather than replaces semantic or dynamic security analyses.</p>
<p><bold>Granularity:</bold> Our analysis operates at the file-level, whereas vulnerabilities often manifest at finer granularity (function or line level). File-level analysis may dilute localized effects, but it aligns well with code review and refactoring workflows commonly used in practice.</p>
</sec>
<sec id="s6_4">
<label>6.4</label>
<title>Conclusion Validity</title>
<p>Conclusion validity concerns the soundness of the statistical inferences drawn. We employed non-parametric statistical tests (Mann&#x2013;Whitney U) suitable for non-normal distributions and applied the Bonferroni correction to control for multiple comparisons. Given the large sample size, we complemented significance testing with effect size analysis to ensure practical relevance. In the machine learning experiments, time-series cross-validation was used to avoid overly optimistic estimates caused by temporal leakage. Finally, class imbalance was addressed using SMOTE applied exclusively within training folds. At the same time, this choice may influence decision boundaries; alternative strategies such as cost-sensitive learning or ensemble balancing could be explored in future work.</p>
<p>Importantly, we avoid causal claims and interpret our results as predictive associations suitable for risk prioritization rather than mechanistic explanations. We also acknowledge the absence of a direct empirical, head-to-head comparison with static analysis tools on the same labeled dataset as a limitation of this study and identify such evaluation as an important direction for future work. Similarly, the proposed threshold guidelines have not yet been prospectively validated within active development workflows; such <italic>in-vivo</italic> validation would require longitudinal industrial studies and therefore falls beyond the scope of this retrospective analysis.</p>
</sec>
<sec id="s6_5">
<label>6.5</label>
<title>Reproducibility</title>
<p>To support replication, we provide a companion repository<xref ref-type="fn" rid="fn-11"><sup>11</sup></xref><fn id="fn-11">
<label>11</label>
<p><ext-link ext-link-type="uri" xlink:href="https://github.com/MdioufDataScientist/PredictSecBugs">https://github.com/MdioufDataScientist/PredictSecBugs</ext-link></p>
</fn> containing datasets, scripts, and experimental configurations. While the <italic>Understand</italic> tool requires a license, pre-computed metrics are provided, and open-source alternatives can be used to reproduce the analysis pipeline.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Conclusion</title>
<p>Ensuring the security of modern software systems is essential, as security bugs continue to expose organizations and users to data breaches, system compromise, and significant financial losses. Although secure development practices and automated tools have advanced, detecting security bugs early in the software development lifecycle remains difficult, particularly as systems grow in size and complexity. Identifying which parts of the codebase are more likely to contain vulnerabilities, therefore, remains a critical challenge for both researchers and practitioners.</p>
<p>In this study, we investigated whether software quality metrics can help identify security-prone code and support early prediction of security bugs. Using data from seven major open-source ecosystems and 338,442 file-level instances, including 33,294 buggy files representing 7685 confirmed security bugs, we extracted 25 quality metrics with the <italic>Understand</italic> static analysis tool. We evaluated the relationship between these metrics and real-world vulnerabilities. To analyze this relationship, we applied five complementary feature-importance techniques, statistical analysis, and eleven machine learning models.</p>
<p>Our results reveal several key findings. First, a consistent subset of metrics, including <monospace>Count StmtExe</monospace>, <monospace>CountDeclClass</monospace>, <monospace>SumCyclomatic</monospace>, <monospace>SumCyclomaticModified</monospace>, and <monospace>Avg CyclomaticModified</monospace>, emerged as the strongest indicators of vulnerability-prone code across all feature selection methods. Second, files containing security bugs exhibited median metric values approximately 3 times higher than those of non-buggy files, a pattern we term the &#x201C;3<inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> rule.&#x201D; Mann-Whitney U tests confirmed that these distributional differences are statistically significant across all metrics. Third, machine learning models, especially tree-based ensemble methods, demonstrated strong predictive performance: XGBoost achieved 98% accuracy, 82% recall, and a ROC-AUC of 0.91, making it one of the most effective classifier in our evaluation.</p>
<p>These findings provide actionable insights for practitioners. The identified metrics and thresholds can guide targeted security reviews, inform metric-based alerts in continuous integration pipelines, and support automated prioritization of potentially vulnerable files. Moreover, the demonstrated predictive power of lightweight machine learning models highlights their potential to complement existing secure development practices. It is important to reiterate the fundamental scope of this work. Our approach relies exclusively on static, structural code metrics. While these metrics are effective for identifying files that are statistically more likely to contain security bugs, they cannot capture semantic vulnerabilities, flawed business logic, insecure API usage patterns, or runtime-dependent behaviors such as race conditions or state-dependent flaws. Accordingly, the proposed method is not a complete solution for vulnerability detection. Rather, it should be understood as a prioritization and triage filter whose primary value lies in directing limited security resources toward the most security-prone parts of the codebase. In practice, metric-based prediction is best used as an upstream component in a defense-in-depth strategy, complementing semantic analysis tools, dynamic testing, and manual security review rather than replacing them.</p>
<p>Despite these contributions, some limitations must be acknowledged. Our analysis relies exclusively on static, file-level quality metrics and does not incorporate semantic or process-oriented factors such as code churn, developer experience, or review activity. Additionally, although our dataset spans multiple ecosystems and programming languages, the thresholds and patterns we identified may need to be calibrated for specific projects or domains. Finally, vulnerability introduction is influenced by complex socio-technical factors that cannot be fully captured by static metrics alone.</p>
<p>Building on this work, several specific research directions emerge. First, extending the study to more software systems would enhance generalizability. Second, future studies could differentiate between vulnerability types to explore whether distinct categories exhibit distinct metric signatures. Third, comparative analyses between mono- and multi-language systems could shed light on the impact of linguistic diversity on the introduction of vulnerability. Prior work has shown that multi-language systems, particularly those relying on foreign function interfaces such as JNI, exhibit distinct reliability and maintenance challenges [<xref ref-type="bibr" rid="ref-56">56</xref>,<xref ref-type="bibr" rid="ref-57">57</xref>], suggesting that vulnerability patterns may differ substantially in such contexts. Fourth, qualitative investigations into the severity and characteristics of security bugs may provide complementary insights. Finally, we plan to explore intra-project prediction approaches to address cold-start scenarios or highly imbalanced projects, enabling more effective prediction in early-stage or low-data environments. Future work could explore developing a lightweight CI/CD integration or plugin that operationalizes the proposed thresholds and model predictions, enabling practitioners to adopt the approach with minimal overhead. In addition, empirical studies evaluating the cost-benefit trade-off of acting on these predictions within real development teams would provide valuable insights into the operational efficiency and practical impact of metric-based security prioritization.</p>
<p>Overall, this study demonstrates that static software quality metrics, when leveraged with appropriate analytical and machine learning techniques, can offer meaningful support for early security bug detection and more secure software development practices.</p>
</sec>
</body>
<back>
<ack>
<p>None.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was supported by the Natural Sciences and Engineering Research Council of Canada (NSERC) through a Discovery Grant RGPIN-2019-05062.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Mohamed Diouf, Elis&#x00E9;e Toe; data collection: Mohamed Diouf; implementation and experimental analysis: Mohamed Diouf; analysis and interpretation of results: Mohamed Diouf, Elis&#x00E9;e Toe, Manel Grichi; draft manuscript preparation: Elis&#x00E9;e Toe, Manel Grichi; supervision and methodological guidance: Ha&#x00EF;fa Nakouri, Fehmi Jaafar. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are openly available in a public repository (PredictSecBugs) at <ext-link ext-link-type="uri" xlink:href="https://github.com/MdioufDataScientist/PredictSecBugs">https://github.com/MdioufDataScientist/PredictSecBugs</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Clemente</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Jaafar</surname> <given-names>F</given-names></string-name>, <string-name><surname>Malik</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Is predicting software security bugs using deep learning better than the traditional machine learning algorithms?</article-title> In: <conf-name>2018 IEEE International Conference on Software Quality, Reliability and Security (QRS)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>95</fpage>&#x2013;<lpage>102</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zaman</surname> <given-names>S</given-names></string-name>, <string-name><surname>Adams</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hassan</surname> <given-names>AE</given-names></string-name></person-group>. <article-title>Security versus performance bugs: a case study on firefox</article-title>. In: <conf-name>Proceedings of the 8th Working Conference on Mining Software Repositories (MSR &#x2019;11)</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2011</year>. p. <fpage>93</fpage>&#x2013;<lpage>102</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Bug characteristics in open source software</article-title>. <source>Empir Softw Eng</source>. <year>2014</year>;<volume>19</volume>(<issue>6</issue>):<fpage>1665</fpage>&#x2013;<lpage>705</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10664-013-9258-8</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Aggarwal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jalote</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Integrating static and dynamic analysis for detecting vulnerabilities</article-title>. In: <conf-name>30th Annual International Computer Software and Applications Conference (COMPSAC &#x2019;06). Vol. 1</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2006</year>. p. <fpage>343</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Tretyakov</surname> <given-names>K</given-names></string-name></person-group>. <chapter-title>Machine learning techniques in spam filtering</chapter-title>. In: <source>Data mining problem-oriented seminar, MTAT.03.177</source>. Vol. <volume>3</volume>. <publisher-loc>Tartu, Estonia</publisher-loc>: <publisher-name>University of Tartu</publisher-name>; <year>2004</year>. p. <fpage>60</fpage>&#x2013;<lpage>79</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zamani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Movahedi</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Machine learning techniques for intrusion detection</article-title>. <comment>arXiv:1312.2177. 2013</comment>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Livadas</surname> <given-names>C</given-names></string-name>, <string-name><surname>Walsh</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lapsley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Strayer</surname> <given-names>WT</given-names></string-name></person-group>. <article-title>Using machine learning techniques to identify botnet traffic</article-title>. In: <conf-name>Proceedings of the 31st IEEE Conference on Local Computer Networks</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2006</year>. p. <fpage>967</fpage>&#x2013;<lpage>74</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Han</surname> <given-names>QL</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Software vulnerability detection using deep neural networks: a survey</article-title>. <source>Proc IEEE</source>. <year>2020</year>;<volume>108</volume>(<issue>10</issue>):<fpage>1825</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jproc.2020.2993293</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Catal</surname> <given-names>C</given-names></string-name>, <string-name><surname>Akbulut</surname> <given-names>A</given-names></string-name>, <string-name><surname>Karakati&#x010D;</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pavlinek</surname> <given-names>M</given-names></string-name>, <string-name><surname>Podgorelec</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Can we predict software vulnerability with deep neural network?</article-title> In: <conf-name>Proceedings of the 12th International Conference on Software Technologies (ICSOFT 2017)</conf-name>. <publisher-loc>Madrid, Spain</publisher-loc>: <publisher-name>SCITEPRESS</publisher-name>; <year>2017</year>. p. <fpage>69</fpage>&#x2013;<lpage>76</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Scandariato</surname> <given-names>R</given-names></string-name>, <string-name><surname>Walden</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hovsepyan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Joosen</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Predicting vulnerable software components via text mining</article-title>. <source>IEEE Trans Softw Eng</source>. <year>2014</year>;<volume>40</volume>(<issue>10</issue>):<fpage>993</fpage>&#x2013;<lpage>1006</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tse.2014.2340398</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Williams</surname> <given-names>L</given-names></string-name></person-group>. <article-title>An empirical model to predict security vulnerabilities using code complexity metrics</article-title>. In: <conf-name>Proceedings of the Second ACM-IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM &#x2019;08)</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2008</year>. p. <fpage>315</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Alves</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fonseca</surname> <given-names>B</given-names></string-name>, <string-name><surname>Antunes</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Software metrics and security vulnerabilities: dataset and exploratory study</article-title>. In: <conf-name>2016 12th European Dependable Computing Conference (EDCC)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>37</fpage>&#x2013;<lpage>44</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ganesh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Palma</surname> <given-names>F</given-names></string-name>, <string-name><surname>Olsson</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Are source code metrics &#x201C;Good enough&#x201D; in predicting security vulnerabilities?</article-title> <source>Data</source>. <year>2022</year>;<volume>7</volume>(<issue>9</issue>):<fpage>127</fpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Camilo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Meneely</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nagappan</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Do bugs foreshadow vulnerabilities? A study of the chromium project</article-title>. In: <conf-name>2015 IEEE/ACM 12th Working Conference on Mining Software Repositories (MSR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2015</year>. p. <fpage>269</fpage>&#x2013;<lpage>79</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Misra</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Bhavsar</surname> <given-names>VC</given-names></string-name></person-group>. <article-title>Relationships between selected software measures and latent bug-density: guidelines for improving quality</article-title>. In: <conf-name>International Conference on Computational Science and Its Applications (ICCSA 2003)</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2003</year>. p. <fpage>724</fpage>&#x2013;<lpage>32</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>El Emam</surname> <given-names>K</given-names></string-name>, <string-name><surname>Melo</surname> <given-names>W</given-names></string-name>, <string-name><surname>Machado</surname> <given-names>JC</given-names></string-name></person-group>. <article-title>The prediction of faulty classes using object-oriented design metrics</article-title>. <source>J Syst Softw</source>. <year>2001</year>;<volume>56</volume>(<issue>1</issue>):<fpage>63</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1016/s0164-1212(00)00086-8</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cuki</surname> <given-names>B</given-names></string-name>, <string-name><surname>Menzies</surname> <given-names>T</given-names></string-name>, <string-name><surname>Bartlow</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Comparing design and code metrics for software quality prediction</article-title>. In: <conf-name>Proceedings of the 4th International Workshop on Predictor Models in Software Engineering (PROMISE &#x2019;08)</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2008</year>. p. <fpage>11</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rahman</surname> <given-names>F</given-names></string-name>, <string-name><surname>Devanbu</surname> <given-names>P</given-names></string-name></person-group>. <article-title>How, and why, process metrics are better</article-title>. In: <conf-name>Proceedings of the 35th International Conference on Software Engineering (ICSE &#x2019;13)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2013</year>. p. <fpage>432</fpage>&#x2013;<lpage>41</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tantithamthavorn</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hassan</surname> <given-names>AE</given-names></string-name>, <string-name><surname>Matsumoto</surname> <given-names>K</given-names></string-name></person-group>. <article-title>The impact of class rebalancing techniques on the performance and interpretation of defect prediction models</article-title>. <source>IEEE Trans Softw Eng</source>. <year>2020</year>;<volume>46</volume>(<issue>11</issue>):<fpage>1200</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tse.2018.2876537</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tantithamthavorn</surname> <given-names>C</given-names></string-name>, <string-name><surname>McIntosh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hassan</surname> <given-names>AE</given-names></string-name>, <string-name><surname>Matsumoto</surname> <given-names>K</given-names></string-name></person-group>. <article-title>The impact of automated parameter optimization on defect prediction models</article-title>. <source>IEEE Trans Softw Eng</source>. <year>2019</year>;<volume>45</volume>(<issue>7</issue>):<fpage>683</fpage>&#x2013;<lpage>711</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tse.2018.2794977</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pandey</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Mishra</surname> <given-names>RB</given-names></string-name>, <string-name><surname>Tripathi</surname> <given-names>AK</given-names></string-name></person-group>. <article-title>BPDET: an effective software bug prediction model using deep representation and ensemble learning techniques</article-title>. <source>Expert Syst Appl</source>. <year>2020</year>;<volume>144</volume>(<issue>1</issue>):<fpage>113085</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2019.113085</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lyu</surname> <given-names>MR</given-names></string-name></person-group>. <article-title>Software defect prediction via convolutional neural network</article-title>. In: <conf-name>2017 IEEE International Conference on Software Quality, Reliability and Security (QRS)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>318</fpage>&#x2013;<lpage>28</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Manjula</surname> <given-names>C</given-names></string-name>, <string-name><surname>Florence</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Deep neural network based hybrid approach for software defect prediction using software metrics</article-title>. <source>Cluster Comput</source>. <year>2019</year>;<volume>22</volume>(<issue>4</issue>):<fpage>9847</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10586-018-1696-z</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ferenc</surname> <given-names>R</given-names></string-name>, <string-name><surname>B&#x00E1;n</surname> <given-names>D</given-names></string-name>, <string-name><surname>Gr&#x00F3;sz</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gyim&#x00F3;thy</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Deep learning in static, metric-based bug prediction</article-title>. <source>Array</source>. <year>2020</year>;<volume>6</volume>(<issue>2</issue>):<fpage>100021</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.array.2020.100021</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rhmann</surname> <given-names>W</given-names></string-name>, <string-name><surname>Pandey</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ansari</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pandey</surname> <given-names>DK</given-names></string-name></person-group>. <article-title>Software fault prediction based on change metrics using hybrid algorithms: an empirical study</article-title>. <source>J King Saud Univ Comput Inf Sci</source>. <year>2020</year>;<volume>32</volume>(<issue>4</issue>):<fpage>419</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jksuci.2019.03.006</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zeng</surname> <given-names>P</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Software vulnerability analysis and discovery using deep learning techniques: a survey</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>197158</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2020.3034766</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chakraborty</surname> <given-names>S</given-names></string-name>, <string-name><surname>Krishna</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ray</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Deep learning based vulnerability detection: are we there yet?</article-title> <source>IEEE Trans Softw Eng</source>. <year>2022</year>;<volume>48</volume>(<issue>9</issue>):<fpage>3280</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tse.2021.3087402</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Research on software vulnerability detection methods based on deep learning</article-title>. <source>J Comput Elec Inf Manag</source>. <year>2024</year>;<volume>14</volume>(<issue>3</issue>):<fpage>21</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.54097/q1rgkx18</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Steenhoek</surname> <given-names>B</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Jiles</surname> <given-names>R</given-names></string-name>, <string-name><surname>Le</surname> <given-names>W</given-names></string-name></person-group>. <article-title>An empirical study of deep learning models for vulnerability detection</article-title>. In: <conf-name>Proceedings of the 45th International Conference on Software Engineering (ICSE &#x2019;23)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>2237</fpage>&#x2013;<lpage>48</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tantithamthavorn</surname> <given-names>C</given-names></string-name></person-group>. <article-title>LineVul: a transformer-based line-level vulnerability prediction</article-title>. In: <conf-name>Proceedings of the 19th International Conference on Mining Software Repositories (MSR &#x2019;22)</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2022</year>. p. <fpage>608</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Le</surname> <given-names>THM</given-names></string-name>, <string-name><surname>Sabir</surname> <given-names>B</given-names></string-name>, <string-name><surname>Babar</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Automated software vulnerability assessment with concept drift</article-title>. In: <conf-name>Proceedings of the 16th International Conference on Mining Software Repositories (MSR &#x2019;19)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>371</fpage>&#x2013;<lpage>82</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Le</surname> <given-names>THM</given-names></string-name>, <string-name><surname>Du</surname> <given-names>X</given-names></string-name>, <string-name><surname>Babar</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Are latent vulnerabilities hidden gems for software vulnerability prediction? An empirical study</article-title>. <comment>arXiv:2401.11105. 2024</comment>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>T</given-names></string-name>, <string-name><surname>Williams</surname> <given-names>L</given-names></string-name>, <string-name><surname>Menzies</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Dazzle: using optimized generative adversarial networks to address security data class imbalance issue</article-title>. In: <conf-name>Proceedings of the 19th International Conference on Mining Software Repositories (MSR &#x2019;22)</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2022</year>. p. <fpage>144</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fehrer</surname> <given-names>T</given-names></string-name>, <string-name><surname>Lozoya</surname> <given-names>RC</given-names></string-name>, <string-name><surname>Sabetta</surname> <given-names>A</given-names></string-name>, <string-name><surname>di Nucci</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tamburri</surname> <given-names>DA</given-names></string-name></person-group>. <article-title>Detecting security fixes in open-source repositories using static code analyzers</article-title>. In: <conf-name>Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering (EASE &#x2019;24)</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2024</year>. p. <fpage>210</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wartschinski</surname> <given-names>L</given-names></string-name>, <string-name><surname>Noller</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Vogel</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kehrer</surname> <given-names>T</given-names></string-name>, <string-name><surname>Grunske</surname> <given-names>L</given-names></string-name></person-group>. <source>VUDENC: vulnerability detection with deep learning on a natural codebase for python</source>. <publisher-name>Inf Softw Technol</publisher-name>. <year>2022</year>;<issue>144</issue>:<fpage>106809</fpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Domain knowledge-based security bug reports prediction</article-title>. <source>Knowl Based Syst</source>. <year>2022</year>;<volume>241</volume>:<fpage>108293</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2022.108293</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Bo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name></person-group>. <article-title>A comprehensive study on security bug characteristics</article-title>. <source>J Softw Evol Process</source>. <year>2021</year>;<volume>33</volume>(<issue>10</issue>):<fpage>e2376</fpage>. doi:<pub-id pub-id-type="doi">10.1002/smr.2376</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mashhadi</surname> <given-names>E</given-names></string-name>, <string-name><surname>Chowdhury</surname> <given-names>S</given-names></string-name>, <string-name><surname>Modaberi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hemmati</surname> <given-names>H</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>G</given-names></string-name></person-group>. <article-title>An empirical study on bug severity estimation using source code metrics and static analysis</article-title>. <source>J Syst Softw</source>. <year>2024</year>;<volume>217</volume>(<issue>1</issue>):<fpage>112179</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jss.2024.112179</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yerramreddy</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mordahl</surname> <given-names>A</given-names></string-name>, <string-name><surname>Koc</surname> <given-names>U</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>S</given-names></string-name>, <string-name><surname>Foster</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Carpuat</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An empirical assessment of machine learning approaches for triaging reports of static analysis tools</article-title>. <source>Empir Softw Eng</source>. <year>2023</year>;<volume>28</volume>(<issue>2</issue>):<fpage>28</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10664-022-10253-z</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kalouptsoglou</surname> <given-names>I</given-names></string-name>, <string-name><surname>Siavvas</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tsoukalas</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kehagias</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chatzigeorgiou</surname> <given-names>A</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Cross-project vulnerability prediction based on software metrics and deep learning</article-title>. In: <conf-name>Computational Science and Its Applications&#x2013;ICCSA 2020. Lecture Notes in Computer Science</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2020</year>. p. <fpage>877</fpage>&#x2013;<lpage>93</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Basili</surname> <given-names>VR</given-names></string-name>, <string-name><surname>Weiss</surname> <given-names>DM</given-names></string-name></person-group>. <article-title>A methodology for collecting valid software engineering data</article-title>. <source>IEEE Trans Softw Eng</source>. <year>1984</year>;<volume>SE-10</volume>(<issue>6</issue>):<fpage>728</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tse.1984.5010301</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Blagus</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lusa</surname> <given-names>L</given-names></string-name></person-group>. <article-title>SMOTE for high-dimensional class-imbalanced data</article-title>. <source>BMC Bioinform</source>. <year>2013</year>;<volume>14</volume>(<issue>1</issue>):<fpage>106</fpage>. doi:<pub-id pub-id-type="doi">10.1186/1471-2105-14-106</pub-id>; <pub-id pub-id-type="pmid">23522326</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gupta</surname> <given-names>A</given-names></string-name>, <string-name><surname>Suri</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>V</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Extracting rules for vulnerabilities detection with static metrics using machine learning</article-title>. <source>Int J Syst Assur Eng Manag</source>. <year>2021</year>;<volume>12</volume>(<issue>1</issue>):<fpage>65</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s13198-020-01036-0</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>SciTools</collab></person-group>. <article-title>What Metrics Does Understand Have [Internet]? 2022 [cited 2026 Jan 2]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://support.scitools.com/support/solutions/articles/70000582223">https://support.scitools.com/support/solutions/articles/70000582223</ext-link>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ghaemi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Feizi-Derakhshi</surname> <given-names>MR</given-names></string-name></person-group>. <article-title>Feature selection using forest optimization algorithm</article-title>. <source>Pattern Recognit</source>. <year>2016</year>;<volume>60</volume>(<issue>1</issue>):<fpage>121</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2016.05.012</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Benesty</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cohen</surname> <given-names>I</given-names></string-name></person-group>. <chapter-title>Pearson correlation coefficient</chapter-title>. In: <source>Noise reduction in speech processing</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2009</year>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-642-00296-0_5</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Adler</surname> <given-names>J</given-names></string-name>, <string-name><surname>Parmryd</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Quantifying colocalization by correlation: the pearson correlation coefficient is superior to the mander&#x2019;s overlap coefficient</article-title>. <source>Cytometry Part A</source>. <year>2010</year>;<volume>77</volume>(<issue>8</issue>):<fpage>733</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1002/cyto.a.20896</pub-id>; <pub-id pub-id-type="pmid">20653013</pub-id></mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kursa</surname> <given-names>MB</given-names></string-name>, <string-name><surname>Rudnicki</surname> <given-names>WR</given-names></string-name></person-group>. <article-title>Feature selection with the boruta package</article-title>. <source>J Statistical Softw</source>. <year>2010</year>;<volume>36</volume>(<issue>11</issue>):<fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.18637/jss.v036.i11</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Hollander</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wolfe</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Chicken</surname> <given-names>E</given-names></string-name></person-group>. <source>Nonparametric statistical methods</source>. <edition>3rd ed</edition>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons</publisher-name>; <year>2013</year>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mel&#x00E9;ndez</surname> <given-names>R</given-names></string-name>, <string-name><surname>Giraldo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Leiva</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Wilcoxon and mann-whitney tests for functional data: an approach based on random projections</article-title>. <source>Mathematics</source>. <year>2020</year>;<volume>9</volume>(<issue>1</issue>):<fpage>44</fpage>. doi:<pub-id pub-id-type="doi">10.3390/math9010044</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Keung</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bennin</surname> <given-names>KE</given-names></string-name>, <string-name><surname>Mi</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Improving bug localization with word embedding and enhanced convolutional neural networks</article-title>. <source>Inf Softw Technol</source>. <year>2019</year>;<volume>105</volume>(<issue>11</issue>):<fpage>17</fpage>&#x2013;<lpage>29</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.infsof.2018.08.002</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Aleem</surname> <given-names>S</given-names></string-name>, <string-name><surname>Capretz</surname> <given-names>LF</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Comparative performance analysis of machine learning techniques for software bug detection</article-title>. In: <conf-name>Proceedings of the 4th International Conference on Software Engineering and Applications</conf-name>. <publisher-loc>Amsterdam, The Netherlands</publisher-loc>: <publisher-name>Elsevier</publisher-name>; <year>2015</year>. p. <fpage>71</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alenezi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zarour</surname> <given-names>M</given-names></string-name></person-group>. <article-title>On the relationship between software complexity and security</article-title>. <source>Int J Softw Eng Appl</source>. <year>2020</year>;<volume>14</volume>(<issue>1</issue>):<fpage>73</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.5121/ijsea.2020.11104</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bettaieb</surname> <given-names>S</given-names></string-name>, <string-name><surname>Casino</surname> <given-names>F</given-names></string-name></person-group>. <article-title>A comprehensive analysis on software vulnerability detection datasets: trends, challenges, and road ahead</article-title>. <source>Int J Inf Secur</source>. <year>2024</year>;<volume>23</volume>(<issue>5</issue>):<fpage>3311</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10207-024-00888-y</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Croft</surname> <given-names>R</given-names></string-name>, <string-name><surname>Babar</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Kholoosi</surname> <given-names>MM</given-names></string-name></person-group>. <article-title>Data quality for software vulnerability datasets</article-title>. In: <conf-name>Proceedings of the 45th International Conference on Software Engineering (ICSE &#x2019;23)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>121</fpage>&#x2013;<lpage>33</lpage>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Grichi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Abidi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gu&#x00E9;h&#x00E9;neuc</surname> <given-names>YG</given-names></string-name>, <string-name><surname>Khomh</surname> <given-names>F</given-names></string-name></person-group>. <article-title>State of practices of java native interface</article-title>. In: <conf-name>Proceedings of the 29th Annual International Conference on Computer Science and Software Engineering (CASCON &#x2019;19)</conf-name>. <publisher-loc>Riverton, NJ, USA</publisher-loc>: <publisher-name>IBM Corp.</publisher-name>; <year>2019</year>. p. <fpage>274</fpage>&#x2013;<lpage>83</lpage>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Grichi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Abidi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jaafar</surname> <given-names>F</given-names></string-name>, <string-name><surname>Eghan</surname> <given-names>EE</given-names></string-name>, <string-name><surname>Adams</surname> <given-names>B</given-names></string-name></person-group>. <article-title>On the impact of interlanguage dependencies in multilanguage systems: empirical case study on java native interface applications (JNI)</article-title>. <source>IEEE Trans Reliab</source>. <year>2021</year>;<volume>70</volume>(<issue>1</issue>):<fpage>428</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tr.2020.3024873</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>