<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">79034</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.079034</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Dendritic Cell Algorithm with Reinforcement Learning for Adaptive Signal Categorization</article-title>
<alt-title alt-title-type="left-running-head">Dendritic Cell Algorithm with Reinforcement Learning for Adaptive Signal Categorization</alt-title>
<alt-title alt-title-type="right-running-head">Dendritic Cell Algorithm with Reinforcement Learning for Adaptive Signal Categorization</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Abudaqqa</surname><given-names>Yousra</given-names></name><email>yousram.83@gmail.com</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Othman</surname><given-names>Zulaiha Ali</given-names></name></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Bakar</surname><given-names>Azuraliza Abu</given-names></name></contrib>
<aff id="aff-1"><institution>Research Center for Artificial Intelligent Technology, Faculty of Information Science and Technology, Universiti Kebangsaan Malaysia</institution>, <addr-line>Bangi, Selangor</addr-line>, <country>Malaysia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yousra Abudaqqa. Email: <email>yousram.83@gmail.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>5</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>2</issue>
<elocation-id>37</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>01</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>03</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_79034.pdf"></self-uri>
<abstract>
<p>Signal categorization is a critical component of the Dendritic Cell Algorithm (DCA), as it directly influences its anomaly detection capability. Conventional DCA implementations typically rely on heuristic or optimization-based approaches, such as Grouping Particle Swarm Optimization (GPSO), Grouping Genetic Algorithms (GGA), Principal Component Analysis (PCA), and Support Vector Machines (SVM), to determine mappings between input features and the three immunological signal categories: Pathogen-Associated Molecular Patterns (PAMP), Danger Signals (DS), and Safe Signals (SS). These approaches depend heavily on domain expertise and predefined rules, making the resulting signal mappings static and often dataset specific. Consequently, the traditional DCA lacks flexibility across diverse data domains and may fail to capture evolving patterns in complex datasets. To address this limitation, this study integrates Reinforcement Learning (RL) into the DCA framework to develop an adaptive signal categorization mechanism. The proposed RL-DCA model employs a Q-learning agent to dynamically assign features to the three signal categories based on reward feedback derived from classification performance. Through continuous interaction with the environment, the RL agent learns an optimal signal mapping policy that improves the quality of generated signals while reducing reliance on manually defined configurations. Experimental evaluations conducted on nine benchmark datasets from multiple domains demonstrate that the proposed RL-DCA framework consistently outperforms existing DCA variants in terms of anomaly detection accuracy and robustness. The results confirm that reinforcement learning provides an effective mechanism for enabling adaptive and data-driven signal categorization in immune-inspired anomaly detection systems.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Dendritic cell algorithm (DCA)</kwd>
<kwd>reinforcement learning (RL)</kwd>
<kwd>Q-learning</kwd>
<kwd>dynamic signal categorization</kwd>
<kwd>anomaly detection</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Fundamental Research Grant Scheme (FRGS) under the Ministry of Higher Education Malaysia</funding-source>
<award-id>FRGS/1/2023/ICT02/UKM/02/2</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The natural immune system exhibits remarkable properties such as adaptability, diversity, distributiveness, and robustness, enabling it to effectively recognize and eliminate harmful pathogens. Artificial Immune Systems (AIS) are a family of bio-inspired computational algorithms developed by modeling these biological immune mechanisms and applying them to a wide range of computational problems. From a computational perspective, AIS approaches offer several advantages that make them attractive for solving complex problems in computer science. By incorporating principles inspired by the natural immune system, AIS models often provide unique capabilities such as adaptability, fault tolerance, and robustness [<xref ref-type="bibr" rid="ref-1">1</xref>]. Many AIS algorithms are considered computationally lightweight in classification and anomaly detection tasks when compared with conventional machine learning methods. Furthermore, previous studies have reported that AIS-based approaches can achieve competitive or superior detection performance in several application domains while maintaining good generalization ability [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<p>Among the various AIS algorithms, the Dendritic Cell Algorithm (DCA) is one of the most widely studied models. DCA emulates the behavior of biological dendritic cells (DCs), which play a key role in the innate immune system by detecting potentially harmful signals and coordinating immune responses. Inspired by this biological mechanism, the DCA has been successfully applied to anomaly detection tasks in several domains, including intrusion detection, fault diagnosis, and spam filtering [<xref ref-type="bibr" rid="ref-3">3</xref>]. The algorithm is particularly suitable for binary classification problems, where each input instance is transformed into three biologically inspired signal types: Pathogen-Associated Molecular Patterns (PAMP), Danger Signals (DS), and Safe Signals (SS). These signals are processed together with data identifiers (antigens) through a context assessment mechanism that determines whether a given instance represents normal or anomalous behavior [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>In general, the DCA consists of four main processing phases: (i) preprocessing and initialization, (ii) detection, (iii) context assessment, and (iv) classification [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>]. Among these stages, the preprocessing phase plays a critical role because it transforms raw input features into the three signal categories (PAMP, DS, and SS) that drive the entire detection process. Consequently, the effectiveness of the DCA largely depends on the quality of feature selection and the correctness of feature-to-signal mapping. Accurate identification of informative features and appropriate assignment of those features to the corresponding signal categories are essential to ensure that the generated signals are meaningful and discriminative for anomaly detection.</p>
<p>Traditionally, the preprocessing stage of the DCA relies heavily on manual configuration or static feature transformation techniques. Methods such as Principal Component Analysis (PCA) [<xref ref-type="bibr" rid="ref-3">3</xref>], Correlation Coefficient (CC) [<xref ref-type="bibr" rid="ref-4">4</xref>], Information Gain (IG), Rough Set Theory (RST) [<xref ref-type="bibr" rid="ref-5">5</xref>], and Fuzzy Rough Set Theory (FRST) [<xref ref-type="bibr" rid="ref-6">6</xref>] have been widely used to construct reduced feature spaces with lower dimensionality. These techniques typically remove weak or irrelevant attributes based on statistical relevance measures. However, relevance according to these criteria does not necessarily guarantee that a feature belongs to the optimal DCA feature subset, nor does irrelevance imply that a feature is completely unsuitable. As a result, filter-based reduction methods may produce suboptimal signal mappings that negatively affect anomaly classification performance [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>To further improve preprocessing, several studies have incorporated machine learning and optimization techniques into the DCA framework. For example, Support Vector Machines (SVM) [<xref ref-type="bibr" rid="ref-8">8</xref>] and K-Nearest Neighbors (KNN) [<xref ref-type="bibr" rid="ref-9">9</xref>] have been used to assist feature selection or signal generation. In addition, optimization-based approaches such as Grouping Particle Swarm Optimization (GPSO) [<xref ref-type="bibr" rid="ref-10">10</xref>] and Grouping Genetic Algorithms (GGA) [<xref ref-type="bibr" rid="ref-11">11</xref>] formulate signal categorization as a combinatorial grouping problem. In these approaches, input features are divided into mutually disjoint subsets representing PAMP, DS, SS, and unused features. These techniques partially automate the signal categorization process and reduce manual intervention. However, the resulting mappings remain static, meaning that once the optimization process finishes, the feature-to-signal assignments remain fixed.</p>
<p>A major limitation of evolutionary optimization approaches such as PSO-DCA and GA-DCA is their limited generalization capability. Because these algorithms optimize feature groupings based on a specific dataset, the resulting mappings often reflect dataset-specific statistical patterns rather than generalizable relationships. Consequently, the learned mappings may perform poorly when applied to new datasets or different domains [<xref ref-type="bibr" rid="ref-12">12</xref>]. In addition, PSO is highly sensitive to parameter configuration and does not incorporate the semantic meaning of the DCA signal categories. Similarly, GA-based methods may suffer from instability due to stochastic operations such as mutation and crossover. As a result, these optimization methods typically require re-optimization for each new dataset, which limits their applicability in multi-domain anomaly detection tasks where adaptability and robustness are required.</p>
<p>Another limitation of the traditional DCA framework is the reliance on static and manually defined signal mappings. In many implementations, users must assign features to PAMP, DS, and SS based on intuition or domain knowledge [<xref ref-type="bibr" rid="ref-13">13</xref>]. Once defined, these mappings remain fixed and cannot adapt to evolving data distributions or changing anomaly patterns. This rigid structure limits the ability of the DCA to respond to new or previously unseen data characteristics, thereby reducing its effectiveness in dynamic environments.</p>
<p>To address these challenges, this study proposes a novel adaptive preprocessing framework called Reinforcement Learning-based DCA (RL-DCA). The proposed approach reformulates the feature-to-signal mapping process as a reinforcement learning (RL) task [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>]. Specifically, the RL-DCA model employs Q-learning [<xref ref-type="bibr" rid="ref-16">16</xref>] to iteratively optimize the assignment of features to the three signal categories (PAMP, DS, and SS) based on reward feedback derived from classification performance. Unlike static signal mappings, the RL agent continuously updates its signal categorization policy through interaction with the environment. This adaptive mechanism enables RL-DCA to automatically adjust signal mappings as data characteristics change, thereby reducing reliance on manual configuration.</p>
<p>Such adaptability is particularly important in domains such as cybersecurity, financial fraud detection, and behavioral analytics, where anomaly patterns evolve over time and static preprocessing methods may quickly become ineffective. By incorporating a self-learning mechanism into the preprocessing phase, RL-DCA can dynamically adapt to changing data distributions and maintain stable anomaly detection performance across heterogeneous datasets.</p>
<p>Motivated by these challenges, this study addresses the following research question:</p>
<p>How can the preprocessing phase of the DCA be made adaptive, learning-driven, and generalizable without relying on manual configuration or static optimization strategies?</p>
<p>To answer this question, the main contributions of this study are summarized as follows:<list list-type="order">
<list-item>
<p>Adaptive Signal Categorization: An RL-based signal categorization mechanism (RL-DCA) is introduced to automatically learn feature-to-signal mappings in the preprocessing stage of the DCA.</p></list-item>
<list-item>
<p>Comprehensive Evaluation: The proposed RL-DCA model is evaluated on nine benchmark datasets from multiple domains and compared with several existing DCA variants, including GGA-DCA, PSO-DCA, PCA-DCA, and SVM-DCA.</p></list-item>
<list-item>
<p>Computational Analysis: The time complexity and runtime behavior of the RL-DCA framework are analyzed and compared with existing DCA-based models.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> presents an overview of the Dendritic Cell Algorithm and its computational principles. <xref ref-type="sec" rid="s3">Section 3</xref> introduces the reinforcement learning framework based on Q-learning. <xref ref-type="sec" rid="s4">Section 4</xref> reviews related work on DCA preprocessing and signal categorization. <xref ref-type="sec" rid="s5">Section 5</xref> describes the proposed RL-DCA model. <xref ref-type="sec" rid="s6">Section 6</xref> presents the experimental setup and datasets used in the evaluation. <xref ref-type="sec" rid="s7">Section 7</xref> reports and analyzes the experimental results. <xref ref-type="sec" rid="s8">Section 8</xref> investigates the sensitivity of the MCAV threshold. <xref ref-type="sec" rid="s9">Section 9</xref> discusses the implications and limitations of the proposed approach. Finally, <xref ref-type="sec" rid="s10">Section 10</xref> concludes the paper and outlines directions for future research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Overviews of DCA</title>
<p>The DCA is an AIS technique inspired by the behaviour of biological DCs in the innate immune system. In immunology, DCs act as professional antigen-presenting cells with the ability to ingest, process, and present antigens to T-cells. Their behaviour is regulated by environmental signals: in the presence of pathogenic or danger cues, DCs transition to a mature state, triggering an immune response; in contrast, exposure primarily to safe signals leads to a semi-mature state that promotes immune tolerance. This dual-state mechanism is the biological foundation for distinguishing between normal and anomalous patterns [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>For clarity, the key variables and notations used throughout this study are summarized here. The DCA operates using three immunological signal categories: Pathogen-Associated Molecular Patterns (PAMP), Danger Signals (DS), and Safe Signals (SS). During the detection phase, each dendritic cell accumulates three output signals: the costimulatory signal (Csm), the mature output signal (MAT), and the semi-mature output signal (SEMI). The migration threshold (MT) determines when a dendritic cell stops sampling antigens and performs context assessment. The final classification decision is derived using the Mature Context Antigen Value (MCAV), which measures the proportion of mature contexts associated with each antigen. In the reinforcement learning component, &#x03B1; denotes the learning rate, &#x03B3; represents the discount factor, and &#x03B5; controls the exploration rate in the &#x03B5;-greedy policy. These parameters regulate how the RL agent updates its signal assignment strategy during training.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Basic Computational Definitions</title>
<p>In the DCA, each data instance is represented using two components: antigens and signals.</p>
<p><bold>Definition 1 An antigen is defined as</bold> <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mi mathvariant="bold-italic">g</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x27E8;</mml:mo><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mo fence="false" stretchy="false">&#x27E9;</mml:mo><mml:mo>,</mml:mo></mml:math></inline-formula> <bold>where</bold> <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi mathvariant="bold-italic">e</mml:mi></mml:math></inline-formula> <bold>is the unique identifier of a data instance, and</bold> <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi mathvariant="bold-italic">t</mml:mi></mml:math></inline-formula> <bold>is the timestamp:</bold> <italic>Antigens represent what is being classified</italic>.</p>
<p><bold>Definition 2 Signals are defined as follows:</bold> <italic>Each data instance is transformed into a three-dimensional real-valued signal vector</italic> <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>D</mml:mi><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></inline-formula> <italic>where: PAMP is a strong an indicator of anomaly, DS is a moderate indicator of anomaly, SS is an indicator of normality</italic>.</p>
<p><bold>Definition 3 DCs:</bold> <italic>A single artificial dendritic cell is defined as</italic> <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mi>g</mml:mi><mml:mi>s</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>T</mml:mi><mml:mo>.</mml:mo></mml:math></inline-formula> <italic>where:</italic> <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>A</mml:mi><mml:mi>g</mml:mi><mml:mi>s</mml:mi></mml:math></inline-formula> <italic>is a set of antigens sampled by the DC</italic>, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:math></inline-formula> <italic>is the cumulative signal profile of that DC, and</italic> <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>T</mml:mi></mml:math></inline-formula> <italic>is the migration threshold that determines when the DC should stop sampling and assess the environment. The DCA operates with a population of DCs, each contributing to the final anomaly decision</italic>.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Overall Structure of the DCA</title>
<p>Following Greensmith &#x0026; Aickelin&#x2019;s abstraction, the DCA can be written as a function: <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:math></inline-formula> where: <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow></mml:math></inline-formula> is the antigen set, <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow></mml:math></inline-formula> is the transformed signal set, and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow></mml:math></inline-formula> is the population of DCs. As formalized in Algorithm 1 and presented in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, The DCA algorithm consists of four main phases include: (1) Pre-processing and Initialization, (2) Detection Phase, (3) Context Assessment, and (4) Classification Phase. Each phase contributes essential mechanisms for transforming input data into anomaly labels.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The standard DCA model.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-1.tif"/>
</fig>
<p><list list-type="simple">
<list-item>
<label>1.</label>
<p><bold>Pre-processing and Initialization Phase:</bold> The first phase prepares the dataset for immune-inspired processing. Each data instance is transformed into two components: an antigen (identifier) and a signal vector consisting of PAMP, DS, and SS signals. These signals are derived through feature selection and domain mapping, enabling the algorithm to convert high-dimensional data into the three-signal structure required by DCA. A population of artificial DCs is then initialized. Each DC is assigned: a random migration threshold <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>T</mml:mi></mml:math></inline-formula>, empty antigen storage, and zero cumulative values for the output signals: Costimulatory molecule (Csm), Semi-mature (SEMI), and Mature (MAT). This initialization prepares the DCs to begin sampling and accumulating signals from the environment.</p></list-item>
<list-item>
<label>2.</label>
<p><bold>Detection Phase:</bold> The detection phase is the core operational stage of the DCA. Each dendritic cell continuously samples antigens and combines the input signals with a predefined weight matrix, shown in <xref ref-type="table" rid="table-1">Table 1</xref>, to produce three interim outputs: Costimulatory signal (Csm), Semi-mature (SEMI), and Mature (MAT).</p>
</list-item>
</list></p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Default weight matrix [<xref ref-type="bibr" rid="ref-5">5</xref>].</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Signal Type</th>
<th>Csm</th>
<th>Mature</th>
<th>Semi-Mature</th>
</tr>
</thead>
<tbody>
<tr>
<td>PAMP</td>
<td>2</td>
<td>1</td>
<td>2</td>
</tr>
<tr>
<td>DS</td>
<td>0</td>
<td>0</td>
<td>3</td>
</tr>
<tr>
<td>SS</td>
<td>2</td>
<td>1</td>
<td>&#x2013;1</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-11">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-11.tif"/>
</fig>
<p>The combination of signals is computed using the weighted linear transformation shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.</p>
<p>Let <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>S</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>D</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the sampled PAMP, safe, and danger signals, respectively, and let <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent their corresponding weights.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mi>S</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2217;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mi>D</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>S</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x2217;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>I</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Every DC continues accumulating the three output values until the Csm value exceeds the migration threshold <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>T</mml:mi></mml:math></inline-formula>. This mechanism prevents DCs from having an excessively long lifespan.
<list list-type="simple">
<list-item>
<label>1.</label>
<p><bold>Context Assessment Phase:</bold> Once a DC&#x2019;s accumulated Csm surpasses its migration threshold, it stops sampling new antigens and enters the context assessment phase. If the Mature output value is greater than the Semi-mature output value, the DC indicates that the sampled antigens are associated with an abnormal context. Conversely, if the Semi-mature output is higher, the DC marks its sampled antigens as normal. This biological analogy reflects the natural behaviour of dendritic cells, where mature DCs stimulate immune activation, while semi-mature DCs promote tolerance.</p></list-item>
<list-item>
<label>2.</label>
<p><bold>Classification Phase:</bold> The final step of the algorithm aggregates the decisions made by all migrated dendritic cells. Since each antigen may be sampled by multiple DCs, the algorithm computes the Mature Context Antigen Value (MCAV) for each antigen to measure the degree of abnormality. MCAV is defined in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>:</p></list-item>
</list></p>
<p><disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>M</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>matured-cells</mml:mtext></mml:mrow><mml:mrow><mml:mtext>presented-cells</mml:mtext></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>A value close to 1 indicates that most DCs handling that antigen matured before migration signifying a high likelihood of anomaly. A value close to 0 reflects normal behaviour. The MCAV is compared against an anomaly threshold, which may be determined automatically from the dataset or defined by the user. If <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>M</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>V</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, the antigen is classified as anomalous. Otherwise, it is classified as normal.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Reinforcement Learning Using Q-Learning</title>
<p>Reinforcement Learning (RL) [<xref ref-type="bibr" rid="ref-6">6</xref>] is a sequential decision-making paradigm in which an agent learns an optimal policy through direct interaction with an environment by maximizing cumulative reward. In the proposed RL-DCA framework, RL is not employed as a conventional classifier. Instead, it serves as a search and optimization mechanism to dynamically learn feature to signal mappings for the DCA. This design directly addresses a key limitation of classical DCA, namely its reliance on manually defined and static signal categorization rules that are often domain dependent and difficult to generalize.</p>
<p>Among various RL approaches, Q-learning is adopted due to its simplicity, model-free nature, and proven convergence properties. Q-learning learns an action&#x2013;value function <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which estimates the expected long-term reward of action <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>a</mml:mi></mml:math></inline-formula> in state <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>s</mml:mi></mml:math></inline-formula> and following the optimal policy thereafter. In the context of RL-DCA, the environment state <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>s</mml:mi></mml:math></inline-formula> is derived from a compact representation of the input features, while the action space consists of three biologically inspired signal categories: PAMP, DS, and SS. Each action corresponds to assigning a particular signal type to the current data instance.</p>
<p>At each interaction step, the RL agent selects an action using an <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula>-greedy strategy, which balances exploration of alternative signal assignments and exploitation of learned knowledge. The reward function is label guided and designed to reinforce biologically meaningful behavior: assigning PAMP to abnormal samples or SS to normal samples yields a positive reward, while incorrect assignments are penalized. This reward structure allows the agent to gradually learn which feature patterns correspond to each signal category without relying on expert-defined thresholds or fixed heuristics. The Q-values are updated iteratively using the Bellman optimality equation, shown as <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi></mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B3;</mml:mi></mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is the learning rate, <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is the discount factor, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>r</mml:mi></mml:math></inline-formula> denotes the immediate reward, and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the next observed state. Through repeated updates, the Q-table gradually converges toward an optimal signal mapping policy that reflects the relationship between feature patterns and immune signal categories. The learned signal vectors are subsequently passed to the DCA, which performs the final immune-inspired aggregation and decision-making process. In this way, Q-learning enhances the adaptability of DCA while preserving its biological interpretation and anomaly detection mechanism.</p>
<p>In the proposed RL-DCA framework, reinforcement learning is used as a label-guided optimization mechanism for signal categorization. Because the reward function uses ground-truth class labels during training, this stage introduces a supervised learning component. However, RL is applied only to optimize signal generation before the standard DCA processing stage, while the final anomaly classification is still performed by the original DCA context assessment and MCAV mechanism.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>Related Work</title>
<p>The preprocessing phase of DCA, particularly feature selection or feature reduction, and signal categorization, plays a crucial role in its anomaly detection performance. This phase determines how raw input features are transformed into the three core signal types: PAMP, DS, and SS [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>In classical DCA implementations, preprocessing is typically performed manually using rule-based or statistically driven mappings informed by domain expertise. Practitioners select relevant features and assign them to signal categories based on intuition, prior knowledge, or predefined thresholds [<xref ref-type="bibr" rid="ref-3">3</xref>]. Although this approach is simple and interpretable, it introduces subjectivity and rigidity, which limit scalability and adaptability, especially in high-dimensional, noisy, or dynamic environments.</p>
<p>To mitigate the limitations of manual preprocessing, several feature reduction techniques have been incorporated into the DCA framework. Principal Component Analysis (PCA) [<xref ref-type="bibr" rid="ref-8">8</xref>] is among the most widely used methods, where original features are transformed into a smaller set of uncorrelated components before being mapped to DCA signals. PCA improves computational efficiency and reduces redundancy; however, it often sacrifices semantic interpretability, making it difficult to trace how individual features influence signal behavior [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Other statistical feature selection methods, such as Information Gain, correlation coefficients, and mutual information, have also been employed to rank and select features prior to signal mapping [<xref ref-type="bibr" rid="ref-10">10</xref>]. While effective in certain scenarios, these approaches generally assume linear relationships and often fail to capture complex feature interactions, reducing their robustness in noisy or non-linear datasets [<xref ref-type="bibr" rid="ref-11">11</xref>].</p>
<p>More recent studies have explored automated feature reduction techniques to further enhance the DCA preprocessing phase. Kernel Principal Component Analysis (KPCA) has been proposed as a nonlinear extension of PCA, enabling projection into kernel-induced feature spaces that better preserve complex relationships and filter noise prior to DCA execution. Autoencoders (AE) have also been adopted to learn compact latent representations through unsupervised neural reconstruction, allowing nonlinear feature compression before mapping reduced features into DCA signals. Hybrid approaches, such as AEkPCA, combine AE-based representation learning with KPCA refinement to further decorrelate and enhance latent features. These methods have demonstrated improved anomaly detection performance across multiple domains. Despite their effectiveness, KPCA-, AE-, and AEkPCA-based pipelines introduce additional computational overhead, hyperparameter sensitivity, and static transformation stages that operate independently of the DCA decision process. Consequently, while these approaches reduce manual intervention, they remain external preprocessing modules rather than adaptive mechanisms integrated within the DCA framework itself [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>Several studies have focused specifically on improving signal categorization in DCA. Mean Squared Error (MSE) and signal sensitivity measures have been used to evaluate the quality of feature-to-signal assignments. Rough Set Theory (RST) has also been applied to identify minimal feature subsets (reducts) that preserve classification performance [<xref ref-type="bibr" rid="ref-13">13</xref>]. The QuickReduct algorithm, for example, derives efficient feature sets while maintaining DCA accuracy. To handle uncertainty in feature importance, Fuzzy Rough Set Theory (FRST) has been introduced, enabling soft boundaries during categorization [<xref ref-type="bibr" rid="ref-13">13</xref>]. Although these techniques improve robustness, their configurations remain static and highly dependent on dataset characteristics.</p>
<p>To further enhance input signal generation, machine learning classifiers have been integrated with the traditional DCA. In SVM-DCA, Support Vector Machines are used to compute sparse weight matrices, where attributes with larger weights are retained and mapped to DCA signals [<xref ref-type="bibr" rid="ref-14">14</xref>]. Similarly, K-Nearest Neighbors (KNN) has been employed to filter attributes based on pairwise instance differences, forming hybrid KNN-DCA models.</p>
<p>Metaheuristic optimization techniques have also been explored to automate signal mapping. Genetic Algorithms (GA) evolve feature-to-signal assignments using mutation and crossover operations, optimizing classification accuracy [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. Particle Swarm Optimization (PSO) and Genetic Grouping Algorithms (GGA) further frame signal categorization as a clustering task, grouping features into PAMP, DS, SS, and unassigned categories [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. These methods reduce manual configuration but typically converge to fixed mappings. Additional studies have incorporated Bayesian Optimization with resource-efficient strategies such as Hyperband to fine-tune signal fusion parameters [<xref ref-type="bibr" rid="ref-18">18</xref>]. Although these methods enhance discriminative power, they still produce static mappings that do not adapt after training.</p>
<p>A critical limitation across existing DCA-based methods is the lack of adaptability in signal categorization. Most approaches generate fixed feature-to-signal mappings that cannot respond to evolving data distributions, class imbalance, or noise. For example, PCA may perform well on one dataset but poorly on another, while GA- and PSO-based methods are often sensitive to parameter tuning and dataset complexity [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>Although the DCA is bio-inspired and widely used for anomaly detection [<xref ref-type="bibr" rid="ref-20">20</xref>], it also suffers from similar weaknesses [<xref ref-type="bibr" rid="ref-21">21</xref>]. In its standard form, the DCA relies on experience-based parameter settings and lacks a training phase, resulting in reduced categorization accuracy when handling large-scale or high-dimensional datasets. This rigid design limits robustness under real-world conditions.</p>
<p>Recent advances in artificial intelligence have significantly influenced the development of zero-day and anomaly detection systems. A comprehensive review by Yee et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] highlights that AI-based detection methods can be broadly categorized into machine learning based, deep learning based, hybrid, and anomaly based approaches. While many supervised and deep learning models achieve high detection accuracy, the review emphasizes recurring challenges such as limited adaptability, dependence on labeled data, sensitivity to data distribution shifts, and difficulty handling unseen attack patterns. These limitations indicate that static learning mechanisms may struggle to maintain performance in evolving threat environments.</p>
<p>Similarly, Dai et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] propose a hybrid intrusion detection framework that integrates an autoencoder with Random Forest and XGBoost classifiers to improve detection of zero-day attacks in unseen data. Their results demonstrate that incorporating anomaly detection improves generalization within the same dataset distribution. However, the model structure remains largely dependent on supervised classifiers after anomaly filtering, and adaptation does not occur within the internal feature transformation process itself. These findings collectively suggest that improving generalization requires not only anomaly aware modeling but also mechanisms that allow detection systems to adapt dynamically during operation.</p>
<p>Inspired by advances in adaptive optimization, recent studies have investigated Reinforcement Learning (RL) as a mechanism for enabling self-learning and adaptability in computational models. Liu et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] applied deep RL for feature selection in high-dimensional domains, outperforming static filter and wrapper methods. Other works have used RL to optimize decision thresholds in network anomaly detection [<xref ref-type="bibr" rid="ref-25">25</xref>], adapt detection parameters in artificial immune systems [<xref ref-type="bibr" rid="ref-26">26</xref>], and control parameter updates in swarm intelligence algorithms [<xref ref-type="bibr" rid="ref-27">27</xref>]. These studies consistently demonstrate that RL enables models to adjust their behavior based on environmental feedback rather than relying on fixed configurations.</p>
<p>Motivated by both the limitations identified in recent AI-based zero-day detection frameworks and the adaptive capability demonstrated by RL-driven models, this study introduces reinforcement learning into the Dendritic Cell Algorithm (DCA) to enhance its signal categorization mechanism. Unlike static or handcrafted mappings, RL frames signal categorization as a sequential decision-making process guided by feedback from classification performance.</p>
<p>Specifically, this work replaces the static signal categorization mechanism with a dynamic RL-based strategy using Q-learning [<xref ref-type="bibr" rid="ref-28">28</xref>]. Through continuous interaction with the environment, the RL agent evaluates outcomes and iteratively updates its signal assignment policy. This design enables the DCA to self-adjust its internal signal transformation process in real time, thereby improving adaptability, generalization capability, and robustness across diverse benchmark datasets. A comparative overview of existing DCA preprocessing methods, including their advantages, limitations, and adaptability, is presented in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Summary of DCA preprocessing methods in terms of advantages, limitations, and adaptability. The proposed RL-DCA is the only method that supports dynamic and feedback-driven signal categorization.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Application</th>
<th>Advantages</th>
<th>Limitations</th>
<th>Adaptability</th>
</tr>
</thead>
<tbody>
<tr>
<td>Manual Mapping [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>Signal Categorization</td>
<td>Simple and interpretable, easy to implement</td>
<td>Subjective, non-scalable, domain-specific</td>
<td>No</td>
</tr>
<tr>
<td>PCA [<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>Feature Reduction</td>
<td>Removes redundancy, improves computational efficiency</td>
<td>Loses interpretability, assumes linearity</td>
<td>No</td>
</tr>
<tr>
<td>Information Gain, Correlation, MI [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>Feature Selection</td>
<td>Fast feature selection, highlights relevance</td>
<td>Ignores feature interactions, not robust to noise</td>
<td>No</td>
</tr>
<tr>
<td>RST [<xref ref-type="bibr" rid="ref-31">31</xref>]</td>
<td>Feature Selection</td>
<td>Identifies minimal feature subsets (reducts)</td>
<td>Requires discretization, performance varies across datasets</td>
<td>No</td>
</tr>
<tr>
<td>FRST [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>Feature Selection</td>
<td>Models&#x2019; uncertainty, flexible signal boundaries</td>
<td>Computationally complex, static configuration</td>
<td>No</td>
</tr>
<tr>
<td>GA [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>Feature Selection</td>
<td>Explores large search spaces, adaptable to objective functions</td>
<td>Parameter sensitivity, performance inconsistency</td>
<td>Limited</td>
</tr>
<tr>
<td>GGA [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>Signal Categorization</td>
<td>Automates signal group assignment, retains interpretability</td>
<td>Produces static mappings, lacks online adaptability</td>
<td>Limited</td>
</tr>
<tr>
<td>GPSO [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>Signal Categorization</td>
<td>Fast convergence, fewer parameters than GA, global search capability</td>
<td>Susceptible to local optima, static output post-training</td>
<td>Limited</td>
</tr>
<tr>
<td>Bayesian Optimization &#x002B; Hyperband [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Signal Categorization</td>
<td>Efficient exploration of parameter space, improves signal fusion performance</td>
<td>Resource-intensive, mappings fixed post-training</td>
<td>No</td>
</tr>
<tr>
<td>Proposed RL-DCA</td>
<td>Signal Categorization</td>
<td>Learns from classification feedback, dynamically adapts signal mappings</td>
<td>Requires training and reward shaping</td>
<td>Yes (Fully Adaptive)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5">
<label>5</label>
<title>The Proposed Model: RL-DCA</title>
<sec id="s5_1">
<label>5.1</label>
<title>Model Overview</title>
<p>This study aims to transform the traditional feature selection and signal categorization procedures adopted by existing DCA-based approaches (e.g., PCA-DCA, GPSO-DCA, and GGA-DCA) into a dynamic and self-learning mechanism using RL with the Q-learning algorithm. Instead of relying on predefined or population-based feature groupings, the proposed RL-DCA framework enables the feature-to-signal assignment process to be learned adaptively through continuous interaction with the environment. In RL-DCA, an RL agent dynamically determines how input</p>
<p>Features should be mapped to the three primary immunological signal categories: PAMP, DS, and SS. Features that do not contribute meaningfully to these signal categories are implicitly excluded from further signal processing. Through this adaptive learning mechanism, RL replaces static signal categorization rules with a data driven optimization process.</p>
<p>A novel hybrid scheme, termed RL-DCA, is introduced by integrating RL-based dynamic signal transformation with the DCA, as illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. RL-DCA consists of three main components: the search space, the search engine, and the evaluation method. The search space represents all possible feature to signal assignments, the RL agent serves as the search engine that optimizes this mapping, and the DCA functions as the evaluation method by assessing classification performance.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The model of RL-DCA.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-2.tif"/>
</fig>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Search Space</title>
<p>In RL-DCA, the search space is defined as the set of all possible mappings between input features and immunological signal categories. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, The signal categorization process establishes a mapping relationship between features and signal types, resulting in a three-dimensional signal representation, as expressed in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mi>F</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>F</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>}</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>D</mml:mi><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Steps of feature selection and signal categorization.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-3.tif"/>
</fig>
<p>Each mapping in the search space represents a potential solution for generating input signals for the DCA. Through interaction with the environment and reward feedback, the RL agent explores this search space and progressively refines the feature to signal assignment strategy.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Search Method: Reinforcement Learning Using Q-Learning</title>
<p>This study employs reinforcement learning with the Q-learning algorithm as the search engine to identify optimal feature to signal mappings. Q-learning is selected due to its model free nature and its ability to learn optimal policies through iterative interaction with the environment.</p>
<p><bold>Step 1: Reinforcement Learning Based Signal Categorization</bold></p>
<p>In the proposed RL-DCA framework, reinforcement learning is employed to dynamically determine the assignment of input instances to immunological signal categories used by the DCA. Unlike conventional approaches that rely on manually predefined feature-to-signal mappings, the RL agent learns an adaptive signal assignment policy through interaction with the environment. Each state represents a compact representation of the input instance derived from the feature space, while each action corresponds to assigning the instance to one of the immunological signal categories defined in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>. The Q-learning algorithm iteratively updates the signal assignment policy based on reward feedback derived from classification consistency, allowing the agent to gradually learn which signal categories best represent the underlying data patterns.</p>
<p><bold>Step 2: RL State Representation via Random Projection and Discretization</bold></p>
<p>Let <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the normalized feature vector of the <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>i</mml:mi></mml:math></inline-formula>-th input instance after preprocessing, where <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>d</mml:mi></mml:math></inline-formula> represents the number of features. Directly using <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> as a state in tabular Q-learning is infeasible because the number of possible states grows exponentially with the dimensionality of the feature space. To obtain a compact yet informative state representation while still incorporating information from all input features, a Random Projection (RP) [<xref ref-type="bibr" rid="ref-32">32</xref>] transformation is employed. Specifically, each feature vector is projected into a two-dimensional embedding space using a randomly generated projection matrix as shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>R</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is a fixed random matrix, whose elements are sampled from a standard normal distribution <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. A fixed random seed is used to ensure reproducibility. The resulting projected components are therefore linear combinations of all original features as shown in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Thus, although the RL agent operates in a two-dimensional state space, the representation implicitly incorporates information from all original features through the projection transformation. This design follows the Johnson&#x2013;Linden Strauss principle, which states that random projections approximately preserve pairwise distances when reducing dimensionality.</p>
<p>To stabilize tabular learning and limit the number of possible states, the projected vector <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is discretized using a binning strategy. After discretization, the RL state is represented as a two-dimensional tuple <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denote the discretized bin indices corresponding to the two projected feature components. This compact state representation enables the tabular Q-learning agent to operate within a finite state space while still capturing information from the original high-dimensional feature set.</p>
<p>First, each component is clipped to a bounded interval: <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. The discretized state is then computed as: <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where shown in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>&#x230A;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>clip</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x230B;</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></disp-formula>and <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>B</mml:mi></mml:math></inline-formula> denotes the number of discretization bins. This discretization process ensures that the RL agent operates within a finite state space while still reflecting patterns from the full feature space through the random projection.</p>
<p><bold>Step 3: Signal Mapping via Q-Learning</bold></p>
<p>The signal categorization process is formulated as a Markov Decision Process (MDP) [<xref ref-type="bibr" rid="ref-33">33</xref>] defined by the state space <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>S</mml:mi></mml:math></inline-formula>, action space <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>A</mml:mi></mml:math></inline-formula>, reward function <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>R</mml:mi></mml:math></inline-formula>, and transition dynamics. The RL agent follows a Q-learning strategy using an <bold>&#x03B5;-</bold>greedy policy to balance exploration and exploitation. At each step, the agent observes the current state <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and selects.</p>
<p>The action space consists of three possible signal assignments defined as <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where action 0 assigns the PAMP signal, action 1 assigns the DS, and action 2 assigns the SS. As shown in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where:
<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mn>0</mml:mn></mml:math></inline-formula>: assign PAMP</p></list-item>
<list-item>
<p><inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mn>1</mml:mn></mml:math></inline-formula>: assign DS</p></list-item>
<list-item>
<p><inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mn>2</mml:mn></mml:math></inline-formula>: assign SS</p></list-item>
</list></p>
<p>The reward function evaluates the correctness of the signal assignment relative to the class label of the instance. A positive reward is given when anomalous instances are mapped to danger-related signals and normal instances are mapped to safe signals. Formally, the reward function is defined as shown in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:mi>a</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The Q-value function is iteratively updated according to the Bellman optimality equation, as shown in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mi>Q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>r</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mspace width="thinmathspace" /><mml:munder><mml:mo form="prefix">max</mml:mo><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:mi>A</mml:mi></mml:mrow></mml:munder><mml:mspace width="thinmathspace" /><mml:mi>Q</mml:mi><mml:mspace width="thinmathspace" /><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>Q</mml:mi><mml:mspace width="thinmathspace" /><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where
<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is the learning rate</p></list-item>
<list-item>
<p><inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is the discount factor</p></list-item>
<list-item>
<p><inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>r</mml:mi></mml:math></inline-formula> is the immediate reward</p></list-item>
<list-item>
<p><inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msup><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the next state.</p></list-item>
</list></p>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates the architecture of the Q-learning mechanism used for adaptive signal mapping in the proposed RL-DCA framework. The figure shows how the Q-learning update rule iteratively refines the signal mapping policy based on reward feedback, enabling the agent to dynamically learn appropriate signal assignments.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Architecture of the Q-learning mechanism for adaptive signal mapping in the RL-DCA framework.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-4.tif"/>
</fig>
<p>Through repeated interactions with the environment, the Q-table gradually converges toward an adaptive signal mapping policy.</p>
<p><bold>Step 4: Signal Vector Construction</bold></p>
<p>Based on the action selected by the RL agent, a three-dimensional signal vector is constructed for each instance as shown in <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>D</mml:mi><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula></p>
<p>Only the signal component corresponding to the selected action is populated, while the remaining components are set to zero. As shown in <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the signal magnitude extracted from the corresponding feature component of the normalized input vector. The resulting signal vectors form the input signal stream for the subsequent DCA processing stage, where immune-inspired signal aggregation and context evaluation are performed to produce the final anomaly classification.</p>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> illustrates the distribution of the immunological signals generated by the proposed RL-based signal categorization mechanism for the SP dataset. The figure presents the three signal types used in the DCA framework. The horizontal axis represents the sample index, while the vertical axis indicates the magnitude of the corresponding signal assigned by the reinforcement learning agent. As observed in the figure, the RL agent dynamically generates signal values according to the learned feature to signal mapping policy obtained through the Q-learning process. The PAMP signal exhibits several peaks across the dataset, indicating regions where abnormal behavioral patterns are detected. The DS appears sparsely with distinct spikes, highlighting instances that are strongly associated with potential anomalies. In contrast, the SS is more broadly distributed across the dataset and dominates in regions corresponding to normal instances. These RL-generated signals collectively form the three-dimensional signal vector defined in <xref ref-type="disp-formula" rid="eqn-11">Eqs. (11)</xref> and <xref ref-type="disp-formula" rid="eqn-12">(12)</xref> and serve as the input signal stream for the subsequent DCA processing phase. Through this adaptive signal generation process, the proposed RL-DCA framework replaces the static signal categorization used in traditional DCA variants with a data driven mechanism that dynamically learns informative signal representations from the dataset.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Distribution of signals generated by the proposed RL-DCA method for the SP dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-5.tif"/>
</fig>
<p><bold>Step 5: Algorithm of RL-DCA</bold></p>
<p>Algorithm 2 summarizes the complete RL-DCA search procedure. The algorithm integrates Q-learning as a search and optimization mechanism for adaptive signal categorization. For each input instance, a compact state representation defined in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> is constructed, and an <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula>-greedy policy is employed to select actions. The Q-values are updated using the Bellman equation in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>, and the resulting signal vectors defined in <xref ref-type="disp-formula" rid="eqn-5">Eqs. (5)</xref> and <xref ref-type="disp-formula" rid="eqn-6">(6)</xref> are provided as input to the DCA. In this way, reinforcement learning does not replace the dendritic cell algorithm; instead, it enhances its adaptability by enabling self-learning signal generation, thereby reducing reliance on static, expert-defined signal categorization rules and improving robustness across heterogeneous datasets.</p>
<fig id="fig-12">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-12.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Experimentations</title>
<p><bold><italic>A Data Sets</italic></bold></p>
<p>To evaluate the performance and generalizability of the proposed RL-DCA model, nine publicly available datasets were selected from reputable sources, including the UCI Machine Learning Repository [<xref ref-type="bibr" rid="ref-34">34</xref>], KEEL [<xref ref-type="bibr" rid="ref-35">35</xref>] and established cybersecurity research platforms [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. These datasets represent a diverse set of binary classification tasks, varying significantly in feature (ranging from 3 to 85 features), sample sizes, and class imbalance levels. The datasets are collected from a variety of domains, including healthcare, finance, social behaviour, spam detection, and network intrusion detection. which allows a comprehensive evaluation of RL-DCA across multiple real-world contexts. <xref ref-type="table" rid="table-3">Table 3</xref> provides the details of each dataset, including the number of features, the total number of samples, and the ratio of class imbalance. To ensure compatibility with the RL-based signal transformation process, all categorical attributes were numerically encoded using label encoding. Numerical features were standardized via Z-score normalization to align data scales and enhance the convergence behavior of the learning algorithm. Data cleaning procedures included the removal of features with over 40% missing values and attributes with zero variance (i.e., those containing a single unique value). Outliers with a Z-score greater than 2.5 were either filtered or imputed using the feature mean, based on the <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the value of the <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> instance for the <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>j</mml:mi><mml:mrow><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> feature, <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the mean, and <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the standard deviation of that feature. This preprocessing ensures that the selected datasets are standardized, clean, and suitable for evaluating the adaptability and robustness of the RL-DCA algorithm across both general classification tasks and real-world anomaly detection scenarios.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Description of datasets used for RL-DCA evaluation.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Features</th>
<th>Samples</th>
<th>Domain</th>
<th>IR (Imbalance Rate)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Cervical Cancer Behavioral Risk (CCBR)</td>
<td>19</td>
<td>72</td>
<td>Health (UCI)</td>
<td>&#x007E;1.6</td>
</tr>
<tr>
<td>German Credit Data (GCD)</td>
<td>20</td>
<td>1000</td>
<td>Finance (UCI)</td>
<td>&#x007E;1.2</td>
</tr>
<tr>
<td>Titanic Survival (TITANIC)</td>
<td>3</td>
<td>2201</td>
<td>General (Keel)</td>
<td>&#x007E;1.58</td>
</tr>
<tr>
<td>Divorce Predictors (DP)</td>
<td>54</td>
<td>170</td>
<td>Social (UCI)</td>
<td>&#x007E;1.43</td>
</tr>
<tr>
<td>Spambase (SP)</td>
<td>57</td>
<td>4601</td>
<td>Text/Spam</td>
<td>&#x007E;1.54</td>
</tr>
<tr>
<td>Cheese Spectroscopy (SPCHEES)</td>
<td>36</td>
<td>3196</td>
<td>Food science</td>
<td>&#x007E;1.2</td>
</tr>
<tr>
<td>NSL-KDD Intrusion Detection</td>
<td>41</td>
<td>125,973</td>
<td>Cyber Security</td>
<td>&#x007E;5.3</td>
</tr>
<tr>
<td>UNSW15 Intrusion Detection</td>
<td>49</td>
<td>175,341</td>
<td>Cyber Security</td>
<td>&#x007E;9.4</td>
</tr>
<tr>
<td>Insurance Company Benchmark (COIL2000)</td>
<td>85</td>
<td>9822</td>
<td>Insurance/Finance</td>
<td>&#x007E;5.9</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold><italic>B Experiment Setup</italic></bold></p>
<p>In this study, a set of experiments was conducted to evaluate the effectiveness and robustness of the proposed RL-DCA, which replaces the traditional signal transformation process in the DCA. The goal is to assess the ability of RL-DCA to dynamically learn and assign signals (PAMP, DS, SS) to features, thereby improving anomaly detection performance across various types of datasets. The experiments were carried out on nine publicly available datasets, which vary in feature dimensionality, data size, and domain characteristics. Each dataset was pre-processed by removing missing values, encoding categorical features using Label Encoder, and scaling numerical features using</p>
<p>Standard Scaler, following commonly adopted preprocessing practices in anomaly detection studies [<xref ref-type="bibr" rid="ref-38">38</xref>]. To ensure fair comparison and consistency across datasets, the same preprocessing pipeline was applied to all datasets. This approach is widely used in comparative benchmarking studies to avoid dataset dependent bias and to ensure that performance differences are attributed to the learning model rather than preprocessing variations [<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
<p>The experimental protocol involves two main evaluations:</p>
<p><bold><italic>C Reinforcement Learning&#x2013;Based Signal Transformation</italic></bold></p>
<p>The proposed RL-DCA framework employs a Q-learning agent to dynamically map input features to the DCA signal categories, namely PAMP, DS, and SS. Through interaction with the environment, the agent learns a signal transformation policy that maximizes the expected reward associated with correct anomaly classification. The behaviour of the RL agent is controlled by three main hyperparameters: the learning rate (&#x03B1;), the discount factor (&#x03B3;), and the exploration rate (&#x03B5;). These parameters regulate the learning dynamics of the agent, including how new information updates the Q-values, how future rewards are considered, and how the balance between exploration and exploitation is maintained during training.</p>
<p>The learning rate &#x03B1; determines how strongly newly observed rewards influence the update of the Q-values. In this study, &#x03B1; &#x003D; 0.3 was selected to allow gradual updates to the Q-table while avoiding unstable fluctuations during learning. A moderate learning rate helps maintain stable convergence while still allowing the agent to adapt its policy. The discount factor &#x03B3; controls the contribution of future rewards in the Q-value update. A value of &#x03B3; &#x003D; 0.9 enables the agent to consider long-term rewards while still emphasizing immediate feedback from the environment. This encourages the agent to learn signal categorization strategies that improve overall detection performance over time.</p>
<p>The exploration rate &#x03B5; defines the probability of selecting a random action instead of the currently optimal action. In this work, &#x03B5; &#x003D; 0.1 was used to ensure sufficient exploration of the action space while allowing the agent to increasingly exploit the best-performing signal mappings identified during training.</p>
<p>The selected parameter configuration <bold>(</bold>&#x03B1; &#x003D; 0.3, &#x03B3; &#x003D; 0.9, &#x03B5; &#x003D; 0.1<bold>)</bold> follows commonly adopted settings in Q-learning-based optimisation methods and was further validated through empirical experimentation. The convergence behaviour of the RL agent is illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, which shows the evolution of the average reward during training on the SP dataset. As training progresses, the learning curve gradually stabilizes, indicating that the RL agent converges toward a consistent signal transformation policy. This behaviour demonstrates that the selected hyperparameter configuration supports stable policy learning within the RL-DCA framework.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Convergence behaviour of the RL-DCA agent during training.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-6.tif"/>
</fig>
<p><bold><italic>D Enhanced DCA Classification</italic></bold></p>
<p>The transformed signals are processed using the standard DCA mechanism, where classification decisions are made based on the Mature Context Antigen Value (MCAV). A decision threshold of 0.85 is applied to distinguish between normal and anomalous antigens. This threshold value has been widely adopted in DCA related studies [<xref ref-type="bibr" rid="ref-40">40</xref>] and was further vali dated experimentally in this work to achieve a suitable balance between detection accuracy and false alarm rate. To ensure statistical reliability, each experiment was repeated for 10 independent runs. Datasets were randomly split into 80% training and 20% testing sets, and 10-fold cross-validation was employed. Model performance was evaluated using multiple classification metrics [<xref ref-type="bibr" rid="ref-41">41</xref>], including accuracy <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>, precision <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>, sensitivity (recall) <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref>, specificity <xref ref-type="disp-formula" rid="eqn-17">Eq. (17)</xref>, and F1-score <xref ref-type="disp-formula" rid="eqn-18">Eq. (18)</xref>, which are standard evaluation measures in anomaly detection and classification research. These metrics are defined as follows:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>To reduce the effect of randomness due to different data splits and initialization conditions, the reported results represent the mean and standard deviation across the repeated runs. The relatively small standard deviation values observed in the experiments indicate that the proposed RL-DCA model produces stable and consistent performance. Where TP, TN, FP, and FN denote true positives, true negatives, false positives, and false negatives, respectively. These metrics collectively provide a comprehensive assessment of classification performance by capturing overall correctness, anomaly detection reliability, detection completeness, normal-class recognition capability, and the balance between precision and recall. All simulations were executed using Python 3.10 in PyCharm on a Windows 10 machine with an Intel Core i7-10510U CPU (1.8 GHz) and 16 GB RAM. Visualization of results was performed using the matplotlib library.</p>
</sec>
<sec id="s7">
<label>7</label>
<title>Complexity Analysis</title>
<p>The RL-DCA introduces an RL mechanism for signal categorization prior to executing the standard DCA pipeline. The runtime of RL-DCA therefore depends on the number of iterations over the dataset, the Q-table update operations, and the signal construction steps defined in Algorithm 1. According to Gu et al. [<xref ref-type="bibr" rid="ref-42">42</xref>], the runtime complexity of the standard DCA is bounded by O(n&#x00B2;), where <italic>n</italic> is the data size. Thus, this study focuses on the runtime of the complete RL-DCA. <xref ref-type="table" rid="table-4">Table 4</xref> shows the details of the primitive operations Algorithm 1, together with the number of times each operation is executed. For each input sample, the algorithm performs state construction, Q-table initialization (if required), &#x03B5;-greedy action selection, reward calculation, Q-value update, and signal construction. All operations are bound by constant time except the Q-update, which requires computing a maximum over the action space. Since the action space is fixed to three categories PAMP, DS, SS), this operation also executes in constant time. Therefore, the runtime of the Q-learning stage scales linearly with the dataset size n. The runtime of the Q-learning stage can be expressed as <xref ref-type="disp-formula" rid="eqn-19">Eq. (19)</xref>:<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>n</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mtext>n</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the initialization cost, and <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denote the constant costs of state processing, action selection, reward computation, Q-table update, and signal construction. Simplifying, we obtain <xref ref-type="disp-formula" rid="eqn-20">Eq. (20)</xref>:<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mi>T</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>o</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Details of primitive operations of Algorithm 1, where <italic>n</italic> is the dataset size.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Line No.</th>
<th>Description</th>
<th>Times</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>Initialize Q-table</td>
<td>1</td>
</tr>
<tr>
<td>2</td>
<td>For-loop over dataset</td>
<td><inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>3&#x2013;5</td>
<td>State formation and Q-table initialization</td>
<td><inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>6&#x2013;9</td>
<td>&#x03B5;-greedy action selection</td>
<td><inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>10</td>
<td>Reward computation</td>
<td><inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>11&#x2013;13</td>
<td>Next-state formation and Q-table initialization</td>
<td><inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>14</td>
<td>Q-value update</td>
<td><inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>16&#x2013;21</td>
<td>Construct signal vector (PAMP, Danger, Safe)</td>
<td><inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>23</td>
<td>Append signal to output set</td>
<td><inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>n</mml:mi></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Since the Q-learning stage is followed by the standard DCA classification phase with complexity <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>o</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, the overall runtime of RL-DCA using is:<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>R</mml:mi><mml:mi>L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>o</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>o</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>o</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>As shown in <xref ref-type="disp-formula" rid="eqn-21">Eq. (21)</xref>, the worst-case runtime complexity of RL-DCA is <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, dominated by the standard DCA phase. The Q-learning signal mapping contributes only to the linear overhead. To further verify this analysis, the empirical runtime was calculated across multiple benchmark datasets. <xref ref-type="fig" rid="fig-7">Fig. 7</xref> illustrates the average execution time of RL-DCA in seconds. The results show that RL-DCA maintains competitive runtime performance, scaling efficiently even on large datasets such as NSL-KDD and UNSW15. Considering the considerable accuracy improvements achieved, the additional computational overhead introduced by the RL mechanism is considered acceptable.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Mean execution time (in seconds) acquired by DCA versions (RL-DCA, GGA-DCA, GPSO-DCA, PCA-DCA, and SVM-DCA) when performing classification across the benchmark datasets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-7.tif"/>
</fig>
</sec>
<sec id="s8">
<label>8</label>
<title>Experimental Results and Discussion</title>
<p>The proposed RL-DCA model was evaluated against six baseline methods, namely GGA-DCA [<xref ref-type="bibr" rid="ref-14">14</xref>], GPSO-DCA [<xref ref-type="bibr" rid="ref-17">17</xref>], PCA-DCA, SVM-DCA, KNN, DT [<xref ref-type="bibr" rid="ref-14">14</xref>] and SVM [<xref ref-type="bibr" rid="ref-43">43</xref>] using nine benchmark datasets spanning biomedical, financial, behavioral, and cybersecurity domains. The evaluation employed multiple performance metrics, including accuracy (<xref ref-type="table" rid="table-5">Table 5</xref>), precision (<xref ref-type="table" rid="table-6">Table 6</xref>), specificity (<xref ref-type="table" rid="table-7">Table 7</xref>), F1-score, and area under the ROC curve (AUC). For all metrics, the reported results represent mean performance values with corresponding standard deviations computed over multiple experimental runs. <xref ref-type="table" rid="table-5">Tables 5</xref>&#x2013;<xref ref-type="table" rid="table-7">7</xref> present detailed numerical comparisons of accuracy, precision, and specificity across all datasets. To improve readability and facilitate interpretation of the remaining performance metrics, the F1-score and AUC results are summarized using graphical representations rather than extensive numerical tables.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Mean accuracy (%) &#x00B1; standard deviation for all algorithms across benchmark datasets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th> Dataset</th>
<th>RL-DCA</th>
<th>GGA-DCA</th>
<th>GPSO-DCA</th>
<th>PCA-DCA</th>
<th>SVM-DCA</th>
<th>KNN</th>
<th>DT</th>
<th>SVM</th>
</tr>
</thead>
<tbody>
<tr>
<td>CCBR</td>
<td>86.9 &#x00B1; 7.3</td>
<td>84.3 &#x00B1; 1.5</td>
<td>83.2 &#x00B1; 2.1</td>
<td>81.5 &#x00B1; 2.4</td>
<td>68.7 &#x00B1; 4.0</td>
<td>86.2 &#x00B1; 10.7</td>
<td>83.5 &#x00B1; 15.0</td>
<td>83.33 &#x00B1; 6.15</td>
</tr>
<tr>
<td>GCD</td>
<td>83.8 &#x00B1; 2.2</td>
<td>78.6 &#x00B1; 0.9</td>
<td>79.5 &#x00B1; 1.7</td>
<td>76.8 &#x00B1; 1.9</td>
<td>50.8 &#x00B1; 3.2</td>
<td>58.5 &#x00B1; 7.1</td>
<td>71.0 &#x00B1; 2.8</td>
<td>72.55 &#x00B1; 1.25</td>
</tr>
<tr>
<td>Titanic</td>
<td>95.2 &#x00B1; 1.0</td>
<td>79.1 &#x00B1; 0.3</td>
<td>81.6 &#x00B1; 0.7</td>
<td>78.5 &#x00B1; 0.8</td>
<td>74.9 &#x00B1; 0.4</td>
<td>74.1 &#x00B1; 3.9</td>
<td>79.0 &#x00B1; 1.2</td>
<td>78.73 &#x00B1; 1.45</td>
</tr>
<tr>
<td>DP</td>
<td>97.3 &#x00B1; 0.7</td>
<td>96.0 &#x00B1; 0.4</td>
<td>96.5 &#x00B1; 0.5</td>
<td>95.8 &#x00B1; 0.6</td>
<td>85.5 &#x00B1; 1.7</td>
<td>97.6 &#x00B1; 3.9</td>
<td>96.4 &#x00B1; 3.9</td>
<td>97.65 &#x00B1; 1.18</td>
</tr>
<tr>
<td>SP</td>
<td>97.3 &#x00B1; 0.7</td>
<td>94.7 &#x00B1; 1.5</td>
<td>95.4 &#x00B1; 0.9</td>
<td>94.1 &#x00B1; 1.1</td>
<td>88.2 &#x00B1; 0.7</td>
<td>80.6 &#x00B1; 3.0</td>
<td>91.6 &#x00B1; 11.0</td>
<td>85.82 &#x00B1; 1.12</td>
</tr>
<tr>
<td>SPCHEES</td>
<td>96.7 &#x00B1; 0.8</td>
<td>93.9 &#x00B1; 1.4</td>
<td>94.5 &#x00B1; 0.9</td>
<td>93.2 &#x00B1; 1.0</td>
<td>90.1 &#x00B1; 1.9</td>
<td>89.5 &#x00B1; 2.0</td>
<td>88.7 &#x00B1; 2.1</td>
<td>88.22 &#x00B1; 1.25</td>
</tr>
<tr>
<td>NSL-KDD</td>
<td>99.1 &#x00B1; 0.5</td>
<td>96.8 &#x00B1; 0.6</td>
<td>97.1 &#x00B1; 0.5</td>
<td>95.9 &#x00B1; 0.7</td>
<td>84.8 &#x00B1; 2.5</td>
<td>83.6 &#x00B1; 2.4</td>
<td>82.1 &#x00B1; 2.6</td>
<td>95.68 &#x00B1; 0.55</td>
</tr>
<tr>
<td>UNSW15</td>
<td>93.9 &#x00B1; 1.3</td>
<td>92.4 &#x00B1; 1.0</td>
<td>92.8 &#x00B1; 0.9</td>
<td>91.5 &#x00B1; 1.1</td>
<td>86.2 &#x00B1; 2.2</td>
<td>85.1 &#x00B1; 2.3</td>
<td>83.5 &#x00B1; 2.4</td>
<td>95.75 &#x00B1; 1.08</td>
</tr>
<tr>
<td>COIL2000</td>
<td>97.6 &#x00B1; 0.3</td>
<td>96.2 &#x00B1; 0.4</td>
<td>96.5 &#x00B1; 0.3</td>
<td>95.8 &#x00B1; 0.4</td>
<td>91.7 &#x00B1; 1.7</td>
<td>90.6 &#x00B1; 1.6</td>
<td>89.8 &#x00B1; 1.9</td>
<td>94.05 &#x00B1; 0.00</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Mean precision (%) &#x00B1; standard deviation for all algorithms across benchmark datasets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>RL-DCA</th>
<th>GGA-DCA</th>
<th>GPSO-DCA</th>
<th>PCA-DCA</th>
<th>SVM-DCA</th>
<th>KNN</th>
<th>DT</th>
<th>SVM</th>
</tr>
</thead>
<tbody>
<tr>
<td>CCBR</td>
<td>73.7 &#x00B1; 14.7</td>
<td>72.9 &#x00B1; 4.0</td>
<td>83.9 &#x00B1; 2.0</td>
<td>82.1 &#x00B1; 2.3</td>
<td>70.2 &#x00B1; 3.8</td>
<td>86.7 &#x00B1; 10.1</td>
<td>84.1 &#x00B1; 14.7</td>
<td>76.50 &#x00B1; 17.13</td>
</tr>
<tr>
<td>GCD</td>
<td>84.8 &#x00B1; 1.7</td>
<td>76.6 &#x00B1; 0.7</td>
<td>80.2 &#x00B1; 1.6</td>
<td>77.5 &#x00B1; 1.8</td>
<td>52.1 &#x00B1; 3.0</td>
<td>59.3 &#x00B1; 6.8</td>
<td>71.8 &#x00B1; 2.7</td>
<td>73.64 &#x00B1; 0.85</td>
</tr>
<tr>
<td>Titanic</td>
<td>95.8 &#x00B1; 0.3</td>
<td>60.1 &#x00B1; 0.3</td>
<td>82.0 &#x00B1; 0.6</td>
<td>79.1 &#x00B1; 0.7</td>
<td>75.3 &#x00B1; 0.4</td>
<td>74.5 &#x00B1; 3.7</td>
<td>79.5 &#x00B1; 1.1</td>
<td>77.29 &#x00B1; 1.57</td>
</tr>
<tr>
<td>DP</td>
<td>99.6 &#x00B1; 0.5</td>
<td>94.8 &#x00B1; 0.5</td>
<td>96.7 &#x00B1; 0.5</td>
<td>96.0 &#x00B1; 0.6</td>
<td>86.1 &#x00B1; 1.6</td>
<td>97.9 &#x00B1; 3.7</td>
<td>96.6 &#x00B1; 3.7</td>
<td>100.00 &#x00B1; 0.00</td>
</tr>
<tr>
<td>SP</td>
<td>99.6 &#x00B1; 0.5</td>
<td>89.4 &#x00B1; 1.3</td>
<td>95.7 &#x00B1; 0.8</td>
<td>94.5 &#x00B1; 1.0</td>
<td>88.9 &#x00B1; 0.6</td>
<td>81.3 &#x00B1; 2.8</td>
<td>92.0 &#x00B1; 10.8</td>
<td>87.14 &#x00B1; 1.44</td>
</tr>
<tr>
<td>SPCHEES</td>
<td>97.6 &#x00B1; 0.6</td>
<td>97.5 &#x00B1; 0.6</td>
<td>94.7 &#x00B1; 0.8</td>
<td>93.5 &#x00B1; 1.0</td>
<td>90.7 &#x00B1; 1.8</td>
<td>89.9 &#x00B1; 1.9</td>
<td>89.1 &#x00B1; 2.0</td>
<td>86.55 &#x00B1; 1.55</td>
</tr>
<tr>
<td>NSL-KDD</td>
<td>99.8 &#x00B1; 0.2</td>
<td>99.0 &#x00B1; 0.2</td>
<td>97.3 &#x00B1; 0.4</td>
<td>96.1 &#x00B1; 0.6</td>
<td>85.2 &#x00B1; 2.4</td>
<td>83.9 &#x00B1; 2.3</td>
<td>82.5 &#x00B1; 2.5</td>
<td>95.82 &#x00B1; 0.95</td>
</tr>
<tr>
<td>UNSW15</td>
<td>94.6 &#x00B1; 0.9</td>
<td>94.5 &#x00B1; 0.9</td>
<td>93.0 &#x00B1; 0.9</td>
<td>91.8 &#x00B1; 1.0</td>
<td>86.5 &#x00B1; 2.1</td>
<td>85.4 &#x00B1; 2.2</td>
<td>83.8 &#x00B1; 2.3</td>
<td>95.52 &#x00B1; 1.23</td>
</tr>
<tr>
<td>COIL2000</td>
<td>73.1 &#x00B1; 2.6</td>
<td>73.5 &#x00B1; 2.6</td>
<td>96.7 &#x00B1;0.3</td>
<td>96.0 &#x00B1; 0.4</td>
<td>91.9 &#x00B1; 1.6</td>
<td>90.9 &#x00B1; 1.5</td>
<td>90.1 &#x00B1; 1.8</td>
<td>0.00 &#x00B1; 0.00</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Mean specificity (%) &#x00B1; standard deviation for all algorithms across benchmark datasets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>RL-DCA</th>
<th>GGA-DCA</th>
<th>GPSO-DCA</th>
<th>PCA-DCA</th>
<th>SVM-DCA</th>
<th>KNN</th>
<th>DT</th>
<th>SVM</th>
</tr>
</thead>
<tbody>
<tr>
<td>CCBR</td>
<td>83.1 &#x00B1; 11.9</td>
<td>82.6&#x00B1; 1.8</td>
<td>84.3 &#x00B1; 2.0</td>
<td>82.5 &#x00B1; 2.2</td>
<td>69.8 &#x00B1; 3.9</td>
<td>87.0 &#x00B1; 9.9</td>
<td>84.4 &#x00B1; 14.5</td>
<td>92.73 &#x00B1; 5.45</td>
</tr>
<tr>
<td>GCD</td>
<td>60.6 &#x00B1; 5.0</td>
<td>79.3 &#x00B1; 1.0</td>
<td>80.5 &#x00B1; 1.5</td>
<td>77.9 &#x00B1; 1.8</td>
<td>51.9 &#x00B1; 3.1</td>
<td>59.5 &#x00B1; 6.7</td>
<td>72.0 &#x00B1; 2.6</td>
<td>20.83 &#x00B1; 3.82</td>
</tr>
<tr>
<td>Titanic</td>
<td>91.3 &#x00B1; 0.7</td>
<td>69.1 &#x00B1; 0.4</td>
<td>82.3 &#x00B1; 0.6</td>
<td>79.3 &#x00B1; 0.7</td>
<td>75.0 &#x00B1; 0.4</td>
<td>74.7 &#x00B1; 3.6</td>
<td>79.7 &#x00B1; 1.1</td>
<td>39.86 &#x00B1; 6.29</td>
</tr>
<tr>
<td>DP</td>
<td>99.8 &#x00B1; 0.3</td>
<td>95.6 &#x00B1; 0.4</td>
<td>96.9 &#x00B1; 0.4</td>
<td>96.2 &#x00B1; 0.5</td>
<td>85.9 &#x00B1; 1.6</td>
<td>98.0 &#x00B1; 3.6</td>
<td>96.8 &#x00B1; 3.6</td>
<td>100.00 &#x00B1; 0.00</td>
</tr>
<tr>
<td>SP</td>
<td>99.8 &#x00B1; 0.3</td>
<td>92.4 &#x00B1; 1.3</td>
<td>95.8 &#x00B1; 0.8</td>
<td>94.8 &#x00B1; 1.0</td>
<td>88.6 &#x00B1; 0.7</td>
<td>81.5 &#x00B1; 2.7</td>
<td>92.3 &#x00B1; 10.6</td>
<td>92.81 &#x00B1; 0.86</td>
</tr>
<tr>
<td>SPCHEES</td>
<td>97.8 &#x00B1; 0.6</td>
<td>94.4 &#x00B1; 1.3</td>
<td>94.9 &#x00B1; 0.8</td>
<td>93.7 &#x00B1; 0.9</td>
<td>90.5 &#x00B1; 1.8</td>
<td>90.1 &#x00B1; 1.9</td>
<td>89.3 &#x00B1; 1.9</td>
<td>87.31 &#x00B1; 1.75</td>
</tr>
<tr>
<td>NSL-KDD</td>
<td>99.9 &#x00B1; 0.2</td>
<td>97.2 &#x00B1; 0.5</td>
<td>97.4 &#x00B1; 0.4</td>
<td>96.2 &#x00B1; 0.5</td>
<td>85.0 &#x00B1; 2.4</td>
<td>84.1 &#x00B1; 2.3</td>
<td>82.8 &#x00B1; 2.4</td>
<td>96.73 &#x00B1; 0.79</td>
</tr>
<tr>
<td>UNSW15</td>
<td>58.8 &#x00B1; 7.1</td>
<td>92.9 &#x00B1; 0.9</td>
<td>93.2 &#x00B1; 0.9</td>
<td>92.0 &#x00B1; 1.0</td>
<td>86.3 &#x00B1; 2.1</td>
<td>85.6 &#x00B1; 2.2</td>
<td>84.0 &#x00B1; 2.2</td>
<td>66.12 &#x00B1; 9.71</td>
</tr>
<tr>
<td>COIL2000</td>
<td>97.8 &#x00B1; 0.3</td>
<td>96.6 &#x00B1; 0.3</td>
<td>96.8 &#x00B1; 0.3</td>
<td>96.1 &#x00B1; 0.4</td>
<td>91.8 &#x00B1; 1.6</td>
<td>91.0 &#x00B1; 1.5</td>
<td>90.3 &#x00B1; 1.8</td>
<td>100.00 &#x00B1; 0.00</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-5">Tables 5</xref>&#x2013;<xref ref-type="table" rid="table-7">7</xref>, RL-DCA demonstrates consistently strong performance across mostbenchmark datasets. In terms of classification accuracy, RL-DCA achieves high detection rates on complex datasets such as NSL-KDD, DP, SP, and SPCHEES, outperforming or matching classical DCA variants and conventional machine learning classifiers. Similar trends are observed for precision and specificity, indicating that RL-DCA effectively balances correct positive detection while minimizing false alarms across diverse data characteristics.</p>

<p>Since the reward function in RL-DCA is label-guided during the learning phase, the reinforcement learning component introduces supervised information into the signal categorization process. To provide a fair and comprehensive evaluation, the experimental comparison includes not only traditional DCA variants but also conventional supervised machine learning classifiers such as SVM. This comparison allows the proposed method to be evaluated against both immune-inspired approaches and standard supervised learning baselines.</p>
<p>The comparative F1-score performance of all algorithms across the benchmark datasets is illustrated in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. This figure highlights the overall balance between precision and recall achieved by each method. RL-DCA attains superior or competitive F1-scores on most datasets, particularly in cybersecurity and structured domains such as NSL-KDD, SP, DP, and UNSW15. These results demonstrate that the adaptive signal categorization enabled by reinforcement learning enhances the robustness of the DCA framework, especially in complex or high-dimensional environments where static signal mappings and traditional classifiers exhibit reduced effectiveness. RL-DCA demonstrates consistently competitive or superior performance across most datasets, highlighting its ability to balance precision and recall across diverse domains.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>F1-score comparison of all algorithms across benchmark datasets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-8.tif"/>
</fig>
<p>The discriminative capability of the evaluated algorithms is further examined using the AUC metric, as shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. The graphical comparison reveals that RL-DCA maintains strong discrimination performance across most datasets, achieving high AUC values on NSL-KDD, DP, SP, and SPCHEES. Although RL-DCA exhibits relatively lower AUC performance on certain imbalanced financial datasets such as GCD and COIL2000, it remains competitive with existing DCA variants and conventional classifiers, confirming its ability to generalize across heterogeneous data distributions.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>The AUC performance comparison of all algorithms across benchmark datasets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-9.tif"/>
</fig>
<p>To further analyze the experimental results, a paired <italic>t</italic>-test was conducted to determine whether statistically significant differences exist between the proposed RL-DCA model and other signal categorization algorithms of the DCA framework (referred to as &#x201C;Comparisons&#x201D;). The statistical test was performed under a significance level of &#x03B1; &#x003D; 0.05.</p>
<p>The following hypotheses were considered:<disp-formula id="ueqn-21"><mml:math id="mml-ueqn-21" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003A;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>R</mml:mi><mml:mi>L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="ueqn-22"><mml:math id="mml-ueqn-22" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003A;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>R</mml:mi><mml:mi>L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2260;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>R</mml:mi><mml:mi>L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the mean classification accuracy obtained by the proposed RL-DCA model, and <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the mean accuracy obtained by the baseline algorithms. <xref ref-type="table" rid="table-8">Table 8</xref> presents the <italic>t</italic>-test results based on classification accuracy between RL-DCA and other DCA-based signal categorization algorithms. With ten experimental runs, the degree of freedom is df &#x003D; 9. At a significance level of &#x03B1; &#x003D; 0.05, the corresponding critical t-value from the student&#x2019;s t-distribution table is 2.262. If the calculated t-value is less than 2.262, the null hypothesis <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is accepted, indicating that no statistically significant difference exists between the compared methods. Conversely, if the calculated t-value exceeds 2.262, the null hypothesis is rejected, indicating a statistically significant performance difference. As shown in <xref ref-type="table" rid="table-8">Table 8</xref>, all calculated t-values are greater than 2.262, demonstrating that the proposed RL-DCA model achieves statistically significant improvements compared with the other signal categorization algorithms of the DCA framework across the evaluated benchmark datasets. These results confirm that the reinforcement learning based signal categorization mechanism effectively enhances the classification capability of the traditional DCA model.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Details of primitive operations of Algorithm 1, where <italic>n</italic> is the dataset size.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th rowspan="2">Category</th>
<th rowspan="2">Dataset</th>
<th colspan="4">RL-DCA</th>
</tr>
<tr>
<th>GGA-DCA</th>
<th>GPSO-DCA</th>
<th>PCA-DCA</th>
<th>SVM-DCA</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>CCBR</td>
<td>6.51</td>
<td>7.00</td>
<td>7.07</td>
<td>8.74</td>
</tr>
<tr>
<td>2</td>
<td>GCD</td>
<td>30.25</td>
<td>34.91</td>
<td>38.24</td>
<td>25.78</td>
</tr>
<tr>
<td>3</td>
<td>Titanic</td>
<td>26.26</td>
<td>48.16</td>
<td>40.11</td>
<td>27.76</td>
</tr>
<tr>
<td>4</td>
<td>DP</td>
<td>11.23</td>
<td>8.85</td>
<td>8.97</td>
<td>15.84</td>
</tr>
<tr>
<td>5</td>
<td>SP</td>
<td>44.07</td>
<td>107.00</td>
<td>48.17</td>
<td>27.96</td>
</tr>
<tr>
<td>6</td>
<td>SPCHEES</td>
<td>29.43</td>
<td>53.18</td>
<td>29.56</td>
<td>38.63</td>
</tr>
<tr>
<td>7</td>
<td>NSL-KDD</td>
<td>204.52</td>
<td>234.75</td>
<td>116.55</td>
<td>182.22</td>
</tr>
<tr>
<td>8</td>
<td>UNSW15</td>
<td>36.87</td>
<td>158.68</td>
<td>127.69</td>
<td>134.57</td>
</tr>
<tr>
<td>9</td>
<td>COIL2000</td>
<td>25.24</td>
<td>18.24</td>
<td>21.62</td>
<td>27.88</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s9">
<label>9</label>
<title>Threshold Sensitivity Analysis</title>
<p>In the traditional DCA, classification is performed using the MCAV, where samples are labeled as anomalous when their MCAV exceeds a predefined threshold <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>. Previous studies commonly adopted a threshold value of <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>&#x03B8;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.85</mml:mn><mml:mo>,</mml:mo></mml:math></inline-formula> however, the influence of this parameter on detection performance requires further investigation. Therefore, a threshold sensitivity analysis was conducted by varying the MCAV threshold within the range:<disp-formula id="ueqn-23"><mml:math id="mml-ueqn-23" display="block"><mml:mi>&#x03B8;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>0.50</mml:mn><mml:mo>,</mml:mo><mml:mn>0.65</mml:mn><mml:mo>,</mml:mo><mml:mn>0.75</mml:mn><mml:mo>,</mml:mo><mml:mn>0.85</mml:mn><mml:mo>,</mml:mo><mml:mn>0.95</mml:mn><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The analysis was performed on two representative datasets with different class distributions. The NSL-KDD dataset, which contains relatively balanced attack and normal samples, was selected to represent balanced scenarios. In contrast, the COIL2000 dataset, which is highly imbalanced, was used to evaluate the behavior of the algorithm under skewed class distributions.</p>
<p><xref ref-type="fig" rid="fig-10">Fig. 10</xref> illustrates the F1-score obtained under different threshold values for both datasets. The results indicate that the proposed RL-DCA model maintains relatively stable performance across a wide range of threshold settings. For the NSL-KDD dataset, the performance remains consistently high across all evaluated thresholds, indicating strong separability between normal and attack traffic. For the COIL2000 dataset, the F1-score improves as the threshold increases from 0.50 to approximately 0.65 and remains relatively stable between 0.65 and 0.95.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>The F1-score obtained under different threshold values for both datasets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79034-fig-10.tif"/>
</fig>
<p>Overall, the results demonstrate that RL-DCA is not overly sensitive to moderate variations in the MCAV threshold. Among the tested values, the threshold 0.85 achieves the highest or near-highest F1-score across the evaluated datasets, providing a balanced trade-off between anomaly detection capability and false positive control. Therefore, this threshold was selected for the subsequent experiments.</p>
<p>Furthermore, the threshold sensitivity analysis highlights a limitation of the traditional DCA classification mechanism: the final decision depends on a manually selected threshold parameter, which may not adapt optimally to different data distributions.</p>
</sec>
<sec id="s10">
<label>10</label>
<title>Discussion</title>
<p>This section discusses the experimental results of the proposed RL-DCA model and explains their implications with respect to the research question. Instead of repeating the numerical values reported in <xref ref-type="table" rid="table-5">Tables 5</xref>&#x2013;<xref ref-type="table" rid="table-7">7</xref>, the discussion focuses on the main performance patterns, the observed trade-offs, and the limitations identified across the evaluated datasets.</p>

<p>Overall, the results show that RL-DCA achieves the highest or near-highest performance on most datasets and evaluation metrics. The improvements are particularly clear in datasets with high dimensionality or complex feature relationships, such as NSL-KDD and UNSW15. In these datasets, traditional DCA variants that rely on static signal mappings or heuristic optimization methods tend to perform less effectively. The results suggest that learning the signal categorization directly from classification feedback helps the model generate more informative signals. As a result, RL-DCA improves the ability of the DCA framework to distinguish between normal and anomalous patterns across different application domains.</p>
<p>On structured datasets such as DP, RL-DCA achieves performance comparable to the best classical machine learning classifiers while showing lower variation across repeated runs. This indicates that the reinforcement learning component converges to a stable signal mapping strategy and reduces sensitivity to initialization. In contrast, conventional classifiers tend to show larger performance degradation on cybersecurity datasets. This behaviour is consistent with known challenges in these datasets, such as class imbalance, sparse features, and complex attack patterns.</p>
<p>Although RL-DCA improves performance in most cases, the results also reveal several dataset-dependent limitations. For example, the GCD and COIL2000 datasets show relatively high accuracy but lower F1-score and precision. This behaviour is mainly caused by class imbalance, where the large number of majority-class instances makes it more difficult to correctly identify rare events. A similar trade-off appears in the UNSW15 dataset. In this case, RL-DCA achieves strong accuracy, but slightly lower specificity and AUC compared with GGA-DCA and GPSO-DCA. This suggests that the model tends to prioritize detecting anomalous instances, which increases sensitivity but may also produce more false positives. Such behaviour can be useful in security applications where missing an attack is more critical than raising additional alarms. However, it also highlights the need for better balance between sensitivity and specificity when dealing with highly imbalanced datasets.</p>
<p>The results provide a clear answer to the research question of this study. The findings show that replacing fixed signal mappings with a learning-based signal categorization mechanism allows the DCA preprocessing stage to become more adaptive and less dependent on manual configuration. By learning signal assignments directly from classification outcomes, RL-DCA can adjust to different dataset characteristics and maintain competitive performance across multiple domains.</p>
<p>It should be noted that the reinforcement learning component uses class labels within the reward function to guide the learning of signal assignments. Therefore, the signal categorization stage becomes label guided during training. However, reinforcement learning is used only to optimize the preprocessing stage of the DCA, while the final anomaly decision is still performed by the original DCA context assessment and MCAV classification mechanism.</p>
<p>The findings also support the main contributions of this work. First, the proposed adaptive signal categorization mechanism improves anomaly detection performance in most evaluated datasets. Second, the experiments conducted on nine benchmark datasets demonstrate that the approach can generalize across different application areas. Third, the runtime analysis shows that the additional computational cost introduced by the learning mechanism remains moderate and acceptable compared with the benefits of removing manual signal design.</p>
<p>Despite these advantages, several limitations remain. The results show that the model is still affected by class imbalance in some datasets, particularly GCD, COIL2000, and UNSW15, where minority-class detection performance is reflected by lower F1-score, specificity, or AUC values. In addition, the current classification stage still relies on a fixed MCAV threshold. A single threshold value may not always be optimal for datasets with different class distributions. Future work will therefore focus on improving the classification stage of the DCA and exploring adaptive threshold strategies. A data-driven threshold adjustment mechanism could further improve the balance between detection sensitivity and false-positive rates.</p>
</sec>
<sec id="s11">
<label>11</label>
<title>Conclusion</title>
<p>This study proposes a new adaptive framework, called RL-DCA, which integrates reinforcement learning with the DCA to achieve dynamic signal categorization. Unlike traditional DCA approaches that rely on fixed rules, static feature transformations, or manually designed preprocessing, the proposed method learns the mapping between features and immunological signals directly from data. By using Q-learning, the RL agent iteratively updates signal assignments based on classification feedback, allowing the model to automatically discover more informative signal representations for anomaly detection.</p>
<p>The proposed framework preserves the immune-inspired structure of the original DCA while improving its preprocessing stage through an adaptive learning mechanism. In RL-DCA, the agent explores different signal assignments for PAMP, DS, and SS using a Q-table and an &#x03B5;-greedy learning strategy. This dynamic process allows the model to adjust signal categorization according to the characteristics of the dataset, reducing dependence on manual configuration and improving generalization.</p>
<p>Experimental results on nine benchmark datasets from healthcare, finance, social behavior, spam detection, and cybersecurity domains demonstrate the effectiveness of the proposed approach. RL-DCA consistently achieves higher or competitive performance compared with existing DCA variants, including PCA-DCA, GGA-DCA, and GPSO-DCA, as well as classical machine learning classifiers such as SVM, KNN, and Decision Tree. These results show that adaptive signal categorization improves anomaly detection performance and enhances the robustness of the DCA framework across different application domains.</p>
<p>Future work will focus on extending the RL-DCA framework to support multi-class classification and online anomaly detection in streaming environments. Such extensions will enable the model to operate in real-time scenarios and further improve its applicability to practical problems such as network intrusion detection, fraud detection, and adaptive cybersecurity monitoring.</p>
<p>Overall, the results demonstrate that incorporating reinforcement learning into the DCA provides an effective way to transform static signal categorization into a data-driven and adaptive process, improving the flexibility and generalization capability of immune-inspired anomaly detection systems.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to express their appreciation to the Center for Artificial Intelligence Technology (CAIT), Universiti Kebangsaan Malaysia (UKM) for providing the research environment and facilities that supported this study.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was supported by the Fundamental Research Grant Scheme (FRGS) under the Ministry of Higher Education Malaysia, grant number FRGS/1/2023/ICT02/UKM/02/2.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design were carried out by Yousra Abudaqqa, Zulaiha Ali Othman, and Azuraliza Abu Bakar. Yousra Abudaqqa conducted the data collection, implementation, and experimental analysis. Interpretation and validation of the results were performed by Yousra Abudaqqa, Zulaiha Ali Othman, and Azuraliza Abu Bakar. Yousra Abudaqqa prepared the initial draft of the manuscript. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The datasets used in this study are publicly available benchmark datasets widely used in machine learning and anomaly detection. These datasets are available from publicly accessible repositories through the UCI Machine Learning Repository and Kaggle. The datasets are available at: UCI Machine Learning Repository: <ext-link ext-link-type="uri" xlink:href="https://archive.ics.uci.edu/">https://archive.ics.uci.edu/</ext-link>, Kaggle Datasets: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets">https://www.kaggle.com/datasets</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Myakala</surname> <given-names>PK</given-names></string-name>, <string-name><surname>Bura</surname> <given-names>C</given-names></string-name>, <string-name><surname>Jonnalagadda</surname> <given-names>AK</given-names></string-name></person-group>. <article-title>Artificial immune systems: a bio-inspired paradigm for computational intelligence</article-title>. <source>J Artif Intell Big Data</source>. <year>2025</year>;<volume>5</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.31586/jaibd.2025.1233</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Widuli&#x0144;ski</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Artificial immune systems in local and network cybersecurity: an overview of intrusion detection strategies</article-title>. <source>Appl Cybersecur Internet Gov</source>. <year>2023</year>;<volume>2</volume>(<issue>1</issue>):<fpage>184306</fpage>. doi:<pub-id pub-id-type="doi">10.60097/acig/162896</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chelly</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Elouedi</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A survey of the dendritic cell algorithm</article-title>. <source>Knowl Inf Syst</source>. <year>2016</year>;<volume>48</volume>(<issue>3</issue>):<fpage>505</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10115-015-0891-y</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Greensmith</surname> <given-names>J</given-names></string-name>, <string-name><surname>Aickelin</surname> <given-names>U</given-names></string-name>, <string-name><surname>Cayzer</surname> <given-names>S</given-names></string-name></person-group>. <chapter-title>Introducing dendritic cells as a novel immune-inspired algorithm for anomaly detection</chapter-title>. In: <source>Artificial immune systems</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2005</year>. p. <fpage>153</fpage>&#x2013;<lpage>67</lpage>. doi:<pub-id pub-id-type="doi">10.1007/11536444_12</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Belhadj</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cherif</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cheriet</surname> <given-names>M</given-names></string-name></person-group>. <article-title>NMF-DCA: an efficient dendritic cell algorithm based on non-negative matrix factorization</article-title>. <source>Int J Comput Digit Syst</source>. <year>2021</year>;<volume>10</volume>:<fpage>575</fpage>&#x2013;<lpage>83</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kaelbling</surname> <given-names>LP</given-names></string-name>, <string-name><surname>Littman</surname> <given-names>ML</given-names></string-name>, <string-name><surname>Moore</surname> <given-names>AW</given-names></string-name></person-group>. <article-title>Reinforcement learning: a survey</article-title>. <source>J Artif Intell Res</source>. <year>1996</year>;<volume>4</volume>:<fpage>237</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1613/jair.301</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Clifton</surname> <given-names>J</given-names></string-name>, <string-name><surname>Laber</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Q-learning: theory and applications</article-title>. <source>Annu Rev Stat Appl</source>. <year>2020</year>;<volume>7</volume>(<issue>1</issue>):<fpage>279</fpage>&#x2013;<lpage>301</lpage>. doi:<pub-id pub-id-type="doi">10.1146/annurev-statistics-031219-041220</pub-id>; <pub-id pub-id-type="pmid">41139587</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mohsin</surname> <given-names>MFM</given-names></string-name>, <string-name><surname>Hamdan</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Abu Bakar</surname> <given-names>A</given-names></string-name></person-group>. <article-title>An upper and lower CUSUM for signal normalization in the dendritic cell algorithm</article-title>. <source>Evol Intell</source>. <year>2016</year>;<volume>9</volume>(<issue>1</issue>):<fpage>37</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s12065-016-0136-3</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Farzadnia</surname> <given-names>E</given-names></string-name>, <string-name><surname>Shirazi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Nowroozi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A new intrusion detection system using the improved dendritic cell algorithm</article-title>. <source>Comput J</source>. <year>2021</year>;<volume>64</volume>(<issue>8</issue>):<fpage>1193</fpage>&#x2013;<lpage>214</lpage>. doi:<pub-id pub-id-type="doi">10.1093/comjnl/bxaa140</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Greensmith</surname> <given-names>J</given-names></string-name>, <string-name><surname>Aickelin</surname> <given-names>U</given-names></string-name>, <string-name><surname>Tedesco</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Information fusion for anomaly detection with the dendritic cell algorithm</article-title>. <source>Inf Fusion</source>. <year>2010</year>;<volume>11</volume>(<issue>1</issue>):<fpage>21</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2009.04.006</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Al-Hasan</surname> <given-names>AA</given-names></string-name>, <string-name><surname>El-Alfy</surname> <given-names>EM</given-names></string-name></person-group>. <article-title>Dendritic cell algorithm for mobile phone spam filtering</article-title>. <source>Procedia Comput Sci</source>. <year>2015</year>;<volume>52</volume>(<issue>3</issue>):<fpage>244</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.procs.2015.05.067</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pereira</surname> <given-names>VE</given-names></string-name></person-group>. <article-title>Automatic signal characterization for the Dendritic Cell Algorithm [dissertation]. Porto, Portugal: Universidade do Porto</article-title>; <year>2024</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chelly</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Elouedi</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Improving the dendritic cell algorithm performance using fuzzy-rough set theory as a pattern discovery technique</article-title>. In: <conf-name>Proceedings of the Fifth International Conference on Innovations in Bio-Inspired Computing and Applications IBICA 2014</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2014</year>. p. <fpage>23</fpage>&#x2013;<lpage>32</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-08156-4_3</pub-id>; <pub-id pub-id-type="pmid">11826839</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Dendritic cell algorithm with grouping genetic algorithm for input signal generation</article-title>. <source>Comput Model Eng Sci</source>. <year>2023</year>;<volume>135</volume>(<issue>3</issue>):<fpage>2025</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2023.022864</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Elisa</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Naik</surname> <given-names>N</given-names></string-name></person-group>. <article-title>A comparative study of genetic algorithm and particle swarm optimisation for dendritic cell algorithm</article-title>. In: <conf-name>2020 IEEE Congress on Evolutionary Computation (CEC); 2020 Jul 19&#x2013;24</conf-name>; <publisher-loc>Glasgow, UK</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cec48606.2020.9185497</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Elisa</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Naik</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Dendritic cell algorithm with optimised parameters using genetic algorithm</article-title>. In: <conf-name>2018 IEEE Congress on Evolutionary Computation (CEC); 2018 Jul 8&#x2013;13</conf-name>; <publisher-loc>Rio de Janeiro, Brazil</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CEC.2018.8477932</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name></person-group>. <chapter-title>Dendritic cell algorithm with group particle swarm optimization for input signal generation</chapter-title>. In: <source>PRICAI 2021: trends in artificial intelligence</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2021</year>. p. <fpage>527</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-89188-6_39</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Dendritic cell algorithm with Bayesian optimization hyperband for signal fusion</article-title>. <source>Comput Mater Contin</source>. <year>2023</year>;<volume>76</volume>(<issue>2</issue>):<fpage>2317</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2023.038026</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>An immune optimization based deterministic dendritic cell algorithm</article-title>. <source>Appl Intell</source>. <year>2022</year>;<volume>52</volume>(<issue>2</issue>):<fpage>1461</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10489-020-02098-0</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>G</given-names></string-name>, <string-name><surname>Shuo</surname> <given-names>P</given-names></string-name>, <string-name><surname>Rong</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chao</surname> <given-names>L</given-names></string-name></person-group>. <article-title>An anomaly detection system based on dendritic cell algorithm</article-title>. In: <conf-name>2009 Third International Conference on Genetic and Evolutionary Computing; 2019 Oct 14&#x2013;17</conf-name>; <publisher-loc>Guilin, China</publisher-loc>. p. <fpage>192</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/WGEC.2009.129</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abudaqqa</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ali Othman</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Abu Bakar</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Enhancing dendritic cell algorithm by integration with multi-layer perceptron for anomaly detection</article-title>. <source>Int J Adv Comput Sci Appl</source>. <year>2025</year>;<volume>16</volume>(<issue>7</issue>):<fpage>1035</fpage>. doi:<pub-id pub-id-type="doi">10.14569/ijacsa.2025.0160798</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yee Por</surname> <given-names>L</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Juan Leem</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Binbeshr</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A systematic literature review on AI-based methods and challenges in detecting zero-day attacks</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>144150</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3455410</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Por</surname> <given-names>LY</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>YL</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ku</surname> <given-names>CS</given-names></string-name>, <string-name><surname>Alizadehsani</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An intrusion detection model to detect zero-day attacks in unseen data using machine learning</article-title>. <source>PLoS One</source>. <year>2024</year>;<volume>19</volume>(<issue>9</issue>):<fpage>e0308469</fpage>. doi:<pub-id pub-id-type="doi">10.1371/journal.pone.0308469</pub-id>; <pub-id pub-id-type="pmid">39259729</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>DR</given-names></string-name>, <string-name><surname>Li</surname> <given-names>HL</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Feature selection and feature learning for high-dimensional batch reinforcement learning: a survey</article-title>. <source>Int J Autom Comput</source>. <year>2015</year>;<volume>12</volume>(<issue>3</issue>):<fpage>229</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11633-015-0893-y</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>ID-RDRL: a deep reinforcement learning-based feature selection intrusion detection model</article-title>. <source>Sci Rep</source>. <year>2022</year>;<volume>12</volume>(<issue>1</issue>):<fpage>15370</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-022-19366-3</pub-id>; <pub-id pub-id-type="pmid">36100644</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Howley</surname> <given-names>E</given-names></string-name>, <string-name><surname>Schukat</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Agent-based dynamic thresholding for adaptive anomaly detection using reinforcement learning</article-title>. <source>Neural Comput Appl</source>. <year>2025</year>;<volume>37</volume>(<issue>23</issue>):<fpage>18775</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-024-10536-0</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Basak</surname> <given-names>H</given-names></string-name>, <string-name><surname>Das</surname> <given-names>M</given-names></string-name>, <string-name><surname>Modak</surname> <given-names>S</given-names></string-name></person-group>. <article-title>RSO: a novel reinforced swarm optimization algorithm for feature selection</article-title>. In: <conf-name>IEEE EUROCON 2021 19th International Conference on Smart Technologies; 2021 Jul 6&#x2013;8</conf-name>; <publisher-loc>Lviv, Ukraine</publisher-loc>. p. <fpage>203</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/eurocon52738.2021.9535639</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Roslan</surname> <given-names>SNRB</given-names></string-name>, <string-name><surname>Huddin</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Bahari</surname> <given-names>MRK</given-names></string-name>, <string-name><surname>Zaki</surname> <given-names>WMDW</given-names></string-name>, <string-name><surname>Mokri</surname> <given-names>SS</given-names></string-name></person-group>. <article-title>Reinforcement learning for automated anterior commissure localization in brain CT scans for tumor detection</article-title>. In: <conf-name>2025 21st IEEE International Colloquium on Signal Processing &#x0026; Its Applications (CSPA); 2025 Feb 7&#x2013;8</conf-name>; <publisher-loc>Pulau Pinang, Malaysia</publisher-loc>. p. <fpage>123</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CSPA64953.2025.10933268</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ng</surname> <given-names>DV</given-names></string-name>, <string-name><surname>Hwang</surname> <given-names>JG</given-names></string-name></person-group>. <article-title>Android malware detection using the dendritic cell algorithm</article-title>. In: <conf-name>2014 International Conference on Machine Learning and Cybernetics; 2014 Jul 13&#x2013;16</conf-name>; <publisher-loc>Lanzhou, China</publisher-loc>. p. <fpage>257</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICMLC.2014.7009126</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Theoretical and empirical extensions of the dendritic cell algorithm [dissertation]. Nottingham, UK: University of Nottingham</article-title>; <year>2011</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Chelly</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Elouedi</surname> <given-names>Z</given-names></string-name></person-group>. <chapter-title>RC-DCA: a new feature selection and signal categorization technique for the dendritic cell algorithm based on rough set theory</chapter-title>. In: <source>Artificial immune systems</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2012</year>. p. <fpage>152</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-642-33757-4_12</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Comparison among dimensionality reduction techniques based on Random Projection for cancer classification</article-title>. <source>Comput Biol Chem</source>. <year>2016</year>;<volume>65</volume>(<issue>18</issue>):<fpage>165</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiolchem.2016.09.010</pub-id>; <pub-id pub-id-type="pmid">27687329</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Boucherie</surname> <given-names>RJ</given-names></string-name>, <string-name><surname>van Dijk</surname> <given-names>NM</given-names></string-name></person-group>. <source>Markov decision processes in practice</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2017</year>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-47766-4</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Frank</surname> <given-names>A</given-names></string-name>, <string-name><surname>Asuncion</surname> <given-names>A</given-names></string-name></person-group>. <article-title>UCI machine learning repository [Internet]</article-title>. <comment>[cited 2026 Jan 1]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://archive.ics.uci.edu/">https://archive.ics.uci.edu/</ext-link>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Triguero</surname> <given-names>I</given-names></string-name>, <string-name><surname>Gonz&#x00E1;lez</surname> <given-names>S</given-names></string-name>, <string-name><surname>Moyano</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Garc&#x00ED;a</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alcal&#x00E1;-Fdez</surname> <given-names>J</given-names></string-name>, <string-name><surname>Luengo</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>KEEL 3.0: an open source software for multi-stage analysis in data mining</article-title>. <source>Int J Comput Intell Syst</source>. <year>2017</year>;<volume>10</volume>(<issue>1</issue>):<fpage>1238</fpage>. doi:<pub-id pub-id-type="doi">10.2991/ijcis.10.1.82</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Moustafa</surname> <given-names>N</given-names></string-name>, <string-name><surname>Slay</surname> <given-names>J</given-names></string-name></person-group>. <article-title>UNSW-NB15: a comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set)</article-title>. In: <conf-name>2015 Military Communications and Information Systems Conference (MilCIS); 2015 Nov 10&#x2013;12</conf-name>; <publisher-loc>Canberra, Australia</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/MilCIS.2015.7348942</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tavallaee</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bagheri</surname> <given-names>E</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ghorbani</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>A detailed analysis of the KDD CUP 99 data set</article-title>. In: <conf-name>2009 IEEE Symposium on Computational Intelligence for Security and Defense Applications; 2009 Jul 8&#x2013;10</conf-name>; <publisher-loc>Ottawa, ON, Canada</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CISDA.2009.5356528</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Davis</surname> <given-names>JJ</given-names></string-name>, <string-name><surname>Clark</surname> <given-names>AJ</given-names></string-name></person-group>. <article-title>Data preprocessing for anomaly based network intrusion detection: a review</article-title>. <source>Comput Secur</source>. <year>2011</year>;<volume>30</volume>(<issue>6&#x2013;7</issue>):<fpage>353</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cose.2011.05.008</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>J</given-names></string-name></person-group>. <source>The elements of statistical learning: data mining, inference, and prediction</source>. <publisher-loc>Oxford, UK</publisher-loc>: <publisher-name>Oxford University Press (OUP)</publisher-name>; <year>2010</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Elisa</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Naik</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Dendritic cell algorithm enhancement using fuzzy inference system for network intrusion detection</article-title>. In: <conf-name>2019 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE); 2019 Jun 23&#x2013;26</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/fuzz-ieee.2019.8859006</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chicco</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jurman</surname> <given-names>G</given-names></string-name></person-group>. <article-title>The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation</article-title>. <source>BMC Genomics</source>. <year>2020</year>;<volume>21</volume>(<issue>1</issue>):<fpage>6</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s12864-019-6413-7</pub-id>; <pub-id pub-id-type="pmid">31898477</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Greensmith</surname> <given-names>J</given-names></string-name>, <string-name><surname>Aickelin</surname> <given-names>U</given-names></string-name></person-group>. <article-title>Theoretical formulation and analysis of the deterministic dendritic cell algorithm</article-title>. <source>Biosystems</source>. <year>2013</year>;<volume>111</volume>(<issue>2</issue>):<fpage>127</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.biosystems.2013.01.001</pub-id>; <pub-id pub-id-type="pmid">23337179</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chandra</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Bedi</surname> <given-names>SS</given-names></string-name></person-group>. <article-title>Survey on SVM and their application in imageclassification</article-title>. <source>Int J Inf Technol</source>. <year>2021</year>;<volume>13</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s41870-017-0080-1</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>