<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">64896</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.064896</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>SAMI-FGSM: Towards Transferable Attacks with Stochastic Gradient Accumulation</article-title>
<alt-title alt-title-type="left-running-head">SAMI-FGSM: Towards Transferable Attacks with Stochastic Gradient Accumulation</alt-title>
<alt-title alt-title-type="right-running-head">SAMI-FGSM: Towards Transferable Attacks with Stochastic Gradient Accumulation</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Feng</surname><given-names>Haolang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Chen</surname><given-names>Yuling</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><email>ylchen3@gzu.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Huang</surname><given-names>Yang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Xuewei</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Sang</surname><given-names>Haiwei</given-names></name><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<aff id="aff-1"><label>1</label><institution>State Key Laboratory of Public Big Data, Guizhou University</institution>, <addr-line>Guiyang, 550025</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>College of Computer Science and Technology, Guizhou University</institution>, <addr-line>Guiyang, 550025</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>Computer College, Weifang University of Science and Technology</institution>, <addr-line>Weifang, 262700</addr-line>, <country>China</country></aff>
<aff id="aff-4"><label>4</label><institution>School of Mathematics and Big Data, Guizhou Education University</institution>, <addr-line>Guiyang, 550018</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yuling Chen. Email: <email>ylchen3@gzu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>30</day><month>07</month><year>2025</year>
</pub-date>
<volume>84</volume>
<issue>3</issue>
<fpage>4469</fpage>
<lpage>4490</lpage>
<history>
<date date-type="received">
<day>26</day>
<month>2</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>5</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_64896.pdf"></self-uri>
<abstract>
<p>Deep neural networks remain susceptible to adversarial examples, where the goal of an adversarial attack is to introduce small perturbations to the original examples in order to confuse the model without being easily detected. Although many adversarial attack methods produce adversarial examples that have achieved great results in the white-box setting, they exhibit low transferability in the black-box setting. In order to improve the transferability along the baseline of the gradient-based attack technique, we present a novel Stochastic Gradient Accumulation Momentum Iterative Attack (SAMI-FGSM) in this study. In particular, during each iteration, the gradient information is calculated using a normal sampling approach that randomly samples around the sample points, with the highest probability of capturing adversarial features. Meanwhile, the accumulated information of the sampled gradient from the previous iteration is further considered to modify the current updated gradient, and the original gradient attack direction is changed to ensure that the updated gradient direction is more stable. Comprehensive experiments conducted on the ImageNet dataset show that our method outperforms existing state-of-the-art gradient-based attack techniques, achieving an average improvement of 10.2% in transferability.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Adversarial examples</kwd>
<kwd>normal sampling</kwd>
<kwd>gradient accumulation</kwd>
<kwd>adversarial transferability</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation</funding-source>
<award-id>62202118</award-id>
<award-id>U24A20241</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Major Scientific and Technological Special Project of Guizhou Province</funding-source>
<award-id>[2024]014</award-id>
<award-id>[2024]003</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Technological Research Projects from Guizhou Education Department</funding-source>
<award-id>2023]003</award-id>
</award-group>
<award-group id="awg4">
<funding-source>Guizhou Science and Technology Department Hundred Level Innovative Talents Project</funding-source>
<award-id>GCC[2023]018</award-id>
</award-group></funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In recent advancements, neural networks have proven to be highly effective for various complex tasks. Notably, their ability to classify images into multiple categories has been one of the most prominent uses [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. Despite their effectiveness, image classification models are vulnerable to adversarial attacks, where imperceptible perturbations are applied to the input data, leading the models to make erroneous predictions. The process of crafting adversarial examples has garnered increasing attention, as studying these examples helps uncover the weaknesses of models, thereby contributing to improving their robustness. However, adversarial examples have also posed significant security threats, particularly in critical applications such as facial recognition [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>], autonomous driving [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>], and 3D target detection [<xref ref-type="bibr" rid="ref-7">7</xref>]. Although these difficulties exist, adversarial examples play a vital role in revealing vulnerabilities within neural networks and are key to enhancing the models&#x2019; robustness.</p>
<p>Adversarial attack methods are generally classified into two primary types: white-box and black-box attacks. In the case of white-box attacks, the attacker is granted total control over the internal structure and parameters of the target model, allowing for targeted modifications [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. Powerful attacks can be developed by directly creating adversarial examples that leverage the gradient data from a target model. In contrast, black-box attacks often involve studying multiple models, where adversarial examples generated on a surrogate model are transferred to other target models [<xref ref-type="bibr" rid="ref-10">10</xref>]. The ability of adversarial examples to be more effective in black-box attacks is often linked to boosting their transferability across different models. Given that it is often challenging to obtain specific parameters and structural details of target models in real-world scenarios, research on black-box attacks has become increasingly crucial.</p>
<p>Adversarial examples generated under white-box settings have demonstrated outstanding attack performance against models in such settings. Nevertheless, the transferability of these adversarial examples can usually be poor, particularly when applied to models that employ adversarial training or advanced defenses. Several methods for adversarial attacks have been introduced with the aim of boosting the transferability of adversarial examples in black-box settings. For instance, Wang et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] leveraged gradient variance from previous iterations to stabilize the direction of gradient updates, thereby improving the effectiveness of adversarial attacks. Similarly, Wang et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] combined momentum accumulation methods in both spatial and temporal domains by incorporating contextual gradient information from different regions, significantly boosting the attack success rate on most models. Although these methods have proven effective against normally trained models, there remains significant room for improvement in attacking adversarially trained and defended models. Therefore, this study focuses primarily on attacking models with defense capabilities.</p>
<p>In this paper, we introduce a novel attack method named Stochastic Gradient Accumulation Momentum Iterative Attack (SAMI-FGSM). This technique mitigates the overfitting of adversarial examples and enhances their effectiveness against models that have undergone adversarial training and defense mechanisms, by accumulating the stochastic gradient data from each iteration. Specifically, in [<xref ref-type="bibr" rid="ref-11">11</xref>], uniform sampling over a uniform distribution is used to obtain gradient variance information. However, this uniform sampling method can easily suppress the gradient&#x2019;s update towards the optimal direction, especially for points farther from the sample, which exhibit greater feature differences. To address this, we employ normal distribution sampling because points sampled from a normal distribution are concentrated around the sample point, having feature information similar to that of the sample. This approach further enhances the attack effectiveness compared to the original method. Although traditional momentum or variance tuning methods such as variance tuning momentum iterative fast gradient sign method (VMI-FGSM) stabilize the update direction by introducing gradient variance, they still use uniform distribution sampling and are prone to fall into local optima near highly nonlinear decision boundaries. The normal distribution sampling proposed in this paper captures key feature changes with a higher probability by sampling neighborhood points closer to the original sample, which effectively reduces noise interference and is easier to jump out of local optimum. As depicted in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the sampling method introduced in this paper achieves better attack success rates than uniform sampling on adversarially trained models. Furthermore, to alter the attack direction of the original sample, we aggregate gradient information from the selected points. Most gradient attack and defense work focuses on input transformation or integration of multiple models to improve the diversity and transferability of adversarial samples, but there is still limited performance in adversarial training or robustness enhanced defense models. By combining the historical accumulated gradient with the current normal sampling gradient, SAMI&#x2013;FGSM takes into account the sensitive area of the model in the update direction, which significantly improves the success rate of attacking multiple defense models. As illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, since decision boundary of model is highly nonlinear, original attack direction targets only one model. By accumulating the gradient from all sampled points, the attack direction can be modified to target both Model 1 and Model 2, thereby enhancing the transferability of adversarial examples.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The effectiveness of SAMI-FGSM-based adversarial examples in terms of attack success on the Inc-v3 model. The blue represents uniform sampling, and the red represents normal sampling used in our method. The results clearly demonstrate that normal sampling significantly outperforms uniform sampling</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-1.tif"/>
</fig><fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Schematic diagram illustrating our proposed method SAMI-FGSM. The black arrow represents the original attack direction. Our approach optimizes this direction by accumulating gradient information from sampling points near the x samples, enabling simultaneous attacks on both Model 1 and Model 2</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-2.tif"/>
</fig>
<p>The main contributions of this paper are as follows:
<list list-type="bullet">
<list-item>
<p>The traditional uniform sampling approach in variance adjustment often samples features that hinder gradient updates. In this work, we employ normal sampling around the sample points to capture features that enhance attack performance, thereby effectively improving the black-box attack performance of adversarial examples.</p></list-item>
<list-item>
<p>Additionally, by accumulating the gradient information obtained from stochastic sampling, our method can alter the original attack direction of adversarial examples, steering it closer to the optimal direction. This approach enhances the attack effectiveness of adversarial examples against adversarially trained and defended models.</p></list-item>
<list-item>
<p>Comprehensive experiments conducted on the ImageNet dataset demonstrate the applicability of our method. The proposed adversarial attack technique outperforms existing methods, as evidenced by the result on various models.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Adversarial Attacks</title>
<p>This section classifies gradient-based iterative attacks into two categories: traditional gradient attack methods and those that evolve from gradient-based strategies.</p>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Baseline Based on Gradient Attack</title>
<p>Threats to deep neural networks are typically classified into two types: black-box and white-box attacks. Research into the transferability of adversarial examples is classified as part of black-box attack techniques, where the attacker is unable to access details such as the parameters or structure of the victim model. Additionally, black-box attacks can target multiple other models simultaneously, making black-box transferability methods highly sought after. Fast Gradient Sign Method (FGSM) [<xref ref-type="bibr" rid="ref-13">13</xref>], originally developed for white-box attacks, inspired the creation of iterative gradient-based methods tailored for black-box research. This, in turn, spurred the quick advancement of more effective techniques to improve transferability in black-box settings. Dong et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] integrated momentum in gradient-based iterative attacks, while Liu et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] combined the accelerated gradient of Nesterov with gradient attack methods using a momentum-based approach. Wang et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] addressed the issue of local optima by utilizing the gradient variance from earlier steps in the iteration process, and Wang et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] integrated spatial domain gradients within images with earlier work that concentrated on temporal domain gradients. Global momentum initialization is used by Wang et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] to improve update direction stability.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Gradient Attack-Based Derivation</title>
<p>Since gradient-based adversarial attacks were first introduced, numerous methods have been developed along this baseline. In addition, several methods derived from this baseline have been thoroughly examined, typically combined with the original approaches to create adversarial examples that offer better transferability. As an example, Li et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] successfully used the optimization process of multi-step attacks to produce adversarial examples with higher black-box success rates by predicting induced adversarial losses through linear mapping of intermediate-level discrepancies. In order to create adversarial examples with better transferability against both normally trained and defended models, Long et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] presented a novel spectral simulation attack that applies spectral transformations to the inputs and performs model enhancement in the frequency domain. Large step-size updates were used for adversarial examples by Yang et al. [<xref ref-type="bibr" rid="ref-19">19</xref>], who calculated several samples with small step sizes within each large step and then averaged the gradients of these samples. By reducing the discrepancy between the true update direction and the steepest descent direction, this method improves the transferability of the resulting adversarial examples.</p>
</sec>
<sec id="s2_1_3">
<label>2.1.3</label>
<title>Attacks Based on Feature Destruction</title>
<p>In recent years, attacks on the feature space of models have been extensively studied. For example, Wang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed the Feature Importance-aware Attack (FIA), which focuses on disrupting important object-related features that play a major role in model&#x2019;s decision-making, resulting in adversarial examples with improved transferability. Huang et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] focused on augmenting perturbations on designated layers of the source model to adjust existing adversarial examples for better performance in black-box settings. Zhu et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] refine the gradient by averaging it over several nearby data points, and subsequently modify the update gradient with a decay indicator. These methods show that destroying the high-level semantic features of the model or optimizing the middle layer differences can significantly improve the black box attack effect.</p>
</sec>
<sec id="s2_1_4">
<label>2.1.4</label>
<title>Attacks Based on Input Transformation</title>
<p>Diverse Inputs (DI) attack [<xref ref-type="bibr" rid="ref-23">23</xref>] generates diverse input patterns by using arbitrary changes on the input samples at every iteration, where the random transformations consist of a certain probability to perform random resizing and padding, resulting in adversarial examples that exhibit greater randomness and enhanced transferability. Translation-Invariant (TI) attack [<xref ref-type="bibr" rid="ref-24">24</xref>] approximates the gradient by applying a fixed kernel matrix to the gradient of an untranslated image. Each iteration requires a gradient computation as the image is subtly shifted. Thus produced adversarial examples to deceive another model with higher probability. The Scale-Invariant (SI) attack [<xref ref-type="bibr" rid="ref-15">15</xref>] presents scale invariance by scaling a collection of input images by an element of <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mi>i</mml:mi></mml:msup></mml:math></inline-formula> (i signifies the hyperparameter), and optimizing the gradient of this set of images with the gradient of input images to create adversarial examples with transferability.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Adversarial Defense</title>
<p>The threat posed by adversarial examples to deep neural networks has driven the design of more sophisticated defense mechanisms [<xref ref-type="bibr" rid="ref-25">25</xref>]. Tram&#x00E9;r et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] proposed separating the generation of adversarial examples from model parameter training, aiming to increase perturbation diversity during training while reducing the dimensionality of the adversarial examples. A fast adversarial training technique was proposed by Shafahi et al. [<xref ref-type="bibr" rid="ref-27">27</xref>], which simultaneously updates the model parameters and image perturbations within one iteration, achieving a training speed 3 to 30 times faster than traditional approaches. In their work, Gokhale et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] developed an adversarial training approach that creates novel samples, maximizing the classifier&#x2019;s exposure to the attribute space, all without relying on test domain data. The min-max optimization problem is tackled by this adversarial training approach, which first optimizes the loss from adversarial perturbations in the inner maximization phase and then finds the best model parameters in the outer minimization phase.</p>
<p>Currently, one of the best techniques for increasing model robustness is adversarial training; however, it faces challenges related to increased training costs, particularly on large-scale datasets. To address the high computational costs, recent studies have focused on designing efficient methods to enhance model robustness. Naseer et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] developed a NRP model, which uses self-derived supervision to learn to purify images from adversarial interference. To identify hostile examples, Xu et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] developed two feature squeezing methods: Bit Reduction (Bit Red) and Spatial Smoothing. In response to adversarial inputs, Feature Distillation (FD) [<xref ref-type="bibr" rid="ref-31">31</xref>] was introduced as a defense system utilizing JPEG compression.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>This section begins with a definition of adversarial attacks, followed by an introduction to the baseline gradient-based methods. We also explain the underlying motivation for our research and describe the proposed Stochastic Gradient Accumulation Momentum Iterative Attack method, drawing connections to previous attack approaches.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Explanation of Adversarial Attack</title>
<p>Adversarial attacks involve generating an adversarial example <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> from a clean sample <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>x</mml:mi></mml:math></inline-formula> using a classifier <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>f</mml:mi></mml:math></inline-formula> with parameters <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>. The crafted adversarial example <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> causes the classifier <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>f</mml:mi></mml:math></inline-formula> to produce incorrect classifications, i.e., <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2260;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the output of the deep neural network (DNN), often denoted by <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>y</mml:mi></mml:math></inline-formula>. The loss function of the classifier <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>f</mml:mi></mml:math></inline-formula> is represented as <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The generated adversarial example must satisfy the constraint <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> is the constraint value, and <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> denotes the p-norm distance, with <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>p</mml:mi></mml:math></inline-formula> typically being 0, 2, or <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula>. In this work, consistent with previous studies [<xref ref-type="bibr" rid="ref-11">11</xref>], we set <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi><mml:mo>.</mml:mo></mml:math></inline-formula> The specific definition is given as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2260;</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo>.</mml:mo><mml:mi>t</mml:mi><mml:mo>.</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03F5;</mml:mi></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Gradient-Based Attack</title>
<p>Following the introduction of the Fast Gradient Sign Method (FGSM) [<xref ref-type="bibr" rid="ref-13">13</xref>], iterative refinements in gradient-based adversarial attacks have resulted in the creation of advanced techniques, including S<inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msup><mml:mi>M</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>I-FGSM [<xref ref-type="bibr" rid="ref-12">12</xref>]. All of these methods rely on gradient optimization. The gradient-based optimization family includes FGSM [<xref ref-type="bibr" rid="ref-13">13</xref>], I-FGSM [<xref ref-type="bibr" rid="ref-32">32</xref>], MI-FGSM [<xref ref-type="bibr" rid="ref-14">14</xref>], NI-FGSM [<xref ref-type="bibr" rid="ref-15">15</xref>], VNI-FGSM [<xref ref-type="bibr" rid="ref-11">11</xref>], VMI-FGSM [<xref ref-type="bibr" rid="ref-11">11</xref>], and S<inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msup><mml:mi>M</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>I-FGSM [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p><bold>Fast Gradient Sign Method (FGSM)</bold> [<xref ref-type="bibr" rid="ref-13">13</xref>] as the earliest proposed gradient-based attack, generates adversarial examples by inputting a clean sample <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>x</mml:mi></mml:math></inline-formula> into the network and performing a single update based on the loss function <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>J</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The loss function is only calculated once in the adversarial example. The following is the particular generation process:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula> denotes the gradient with respect to <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>x</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>J</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the loss function, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the function that computes the sign of the gradient <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> is the perturbation magnitude.</p>
<p><bold>Iterative Fast Gradient Sign Method (I-FGSM)</bold> [<xref ref-type="bibr" rid="ref-32">32</xref>] extends the single-step method FGSM to a multi-step attack by introducing a step parameter <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mi>x</mml:mi><mml:mn>0</mml:mn><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>x</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>x</mml:mi></mml:math></inline-formula> is a clean sample, <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> is the perturbation size, and <italic>T</italic> is the iteration count.</p>
<p><bold>Momentum Iterative Fast Gradient Sign Method (MI-FGSM)</bold> [<xref ref-type="bibr" rid="ref-14">14</xref>] introduces the momentum factor <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula> collects the gradient of every iteration of I-FGSM as momentum in the next gradient calculation:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mi>g</mml:mi></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula> is the momentum element, and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> is the sum of the current gradient and <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula> times (<italic>t</italic> &#x2212; 1) the next gradient.</p>
<p><bold>Nesterov Iterative Fast Gradient Sign Method (NI-FGSM)</bold> [<xref ref-type="bibr" rid="ref-15">15</xref>] introduces the idea of Nesterov Gradient Descent (NAG) by replacing all <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref> with <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> when calculating <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to additional strengthen the black-box aggressiveness of MI-FGSM.</p>
<p><bold>Variance momentum Iterative Fast Gradient Sign Method (VMI-FGSM)</bold> [<xref ref-type="bibr" rid="ref-11">11</xref>] steady the updated guidance of the present gradient by incorporating the gradient variance details from the prior round of iterations:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mi>J</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msup><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:math></inline-formula> is a random sample within a specific range of <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>x</mml:mi></mml:math></inline-formula> uniform distribution.</p>
<p><bold>Spatial Momentum Iterative Fast Gradient Sign Method (S<inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msup><mml:mtext>M</mml:mtext><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>I-FGSM)</bold> [<xref ref-type="bibr" rid="ref-12">12</xref>] considers contextual gradient knowledge in various image regions, introducing a momentum accumulation system from the timing to the spatial domain, with the gradient updated as:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>s</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>g</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mo>.</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the random resizing and padding used to transform <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, n denotes the number of transformations in the spatial domain, and <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> denotes the weight of the gradient.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Stochastic Gradient Accumulation Method</title>
<p>Deep neural networks often exhibit complex forms when handling high-dimensional classification tasks due to the highly nonlinear and often high-curvature nature of their decision boundaries. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates the high curvature of Model 1 and Model 2&#x2019;s decision boundaries. To induce misclassification in both Model 1 and Model 2, our goal is to generate an adversarial example <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>x</mml:mi></mml:math></inline-formula> that identifies the optimal attack direction during its creation. The current techniques for generating adversarial examples struggle in transferring attacks to other models because of the pronounced curvature in the decision boundary. In contrast, robust models feature smoother decision boundaries, which significantly weakens the impact of adversarial attacks against defensive and adversarially trained model.</p>
<p>Our goal is to improve the black-box transferability of adversarial examples by refining the attack direction, steering it towards a more aggressive path. Inspired by VMI-FGSM [<xref ref-type="bibr" rid="ref-11">11</xref>], we investigate the gradient information of a uniform distribution around the sample during the iterative process, effectively improving the transferability of the final adversarial example. Building on this, we consider sample feature information closer to the surrounding sample points, which shares similarities with the feature information of the sampling points. By setting the sampling point as the center of a normal distribution, we can sample misleading features near the sample point with maximum probability. Due to the concentration of the normal distribution around the sample point, most sampled points exhibit features that are more closely aligned with the original examples. In adversarial attacks, critical feature variations often occur near the model&#x2019;s decision boundary. The localized focus of the normal distribution increases the likelihood of sampling points that are closer to these critical regions, thereby providing more informative guidance for gradient optimization. This focus reduces deviations in gradient update directions, resulting in smoother and more stable gradient variations, which enhance the generalization and cross-model transferability of adversarial perturbations. Furthermore, the localized sampling characteristic of the normal distribution mitigates the interference of high-curvature decision boundaries on gradient updates, particularly in adversarially trained models, thereby improving the accuracy of gradient update directions and the overall attack effectiveness. In contrast, uniform distribution sampling generates points randomly across the entire sampling range, which may result in sampled points with features that deviate from the original examples, introducing noise and irrelevant information into the optimization process. As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the comparison of the two sampling methods clearly demonstrates that normal sampling outperforms uniform sampling on adversarially trained models. To boost the attack power of the adversarial example, we add gradient information from the original sample as noise to image. This is achieved by performing random sampling around the sample&#x2019;s point based on a normal distribution and aggregating the gradients obtained from this sampling.</p>
<p>Based on this, we present a new attack method called Stochastic Gradient Accumulation Momentum Iterative Attack (SAMI-FGSM). At each iteration, the method combines the gradient information from the earlier step to ensure a more stable gradient direction, effectively smoothing the update direction. Additionally, during the calculation of accumulated gradient information, it incorporates sample gradient information from a specific normal distribution range. The specific implementation of the proposed SAMI-FGSM method is as follows:</p>
<p><bold>Definition:</bold> Given a classifier <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>f</mml:mi></mml:math></inline-formula> with parameters <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> and a loss function <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, the gradient accumulation can be described as follows:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#xA0;</mml:mtext></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msup><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes a d-dimensional normal distribution. After obtaining the random gradient information from the (<italic>t</italic> &#x2212; 1)-th iteration through the above process, this information is used to adjust the gradient of <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> in the <italic>t</italic>-th iteration, thereby stabilizing the gradient update direction more effectively. Specifically, <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> aggregates the gradients of all sampled points in the current iteration, and this accumulation affects the gradient update in the following ways:</p>
<p><bold>Stable direction:</bold> The historical gradient information (<inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>A</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>) is combined with the current gradient to smooth the randomness of the single gradient update and avoid falling into local optimum.</p>
<p><bold>Enhanced generalization:</bold> The gradient mean of multiple sampling points implies the local geometric characteristics of the decision boundary of the model, so that the attack direction is more suitable for the boundary differences of different models.</p>
<p>The complete process of the proposed stochastic gradient accumulation method is described in Algorithm 1, referred to as SAMI-FGSM. The method proposed here demonstrates optimal performance along the primary path of gradient-based attacks and is compatible with various derivative techniques, such as frequency domain attacks [<xref ref-type="bibr" rid="ref-18">18</xref>] and adaptive targeted attacks [<xref ref-type="bibr" rid="ref-33">33</xref>]. Furthermore, the proposed method is compatible with a variety of existing methods such as DIM attacks, TIM attacks, and SIM attacks.</p>
<fig id="fig-8">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-8.tif"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Differences from Existing Attacks</title>
<p>Derivative methods based on gradient attacks primarily originate from FGSM. This section provides an overview of the key gradient-based attack techniques, as illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. If the domain upper bound <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is set to 0, SAMI-FGSM degenerates into MI-FGSM. In the same way, setting the decay factor <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula> to 0 and limiting the number of iterations T to 1 results in these methods returning to the standard FGSM. Moreover, the aforementioned gradient-based attack methods are capable of being integrated with many input transformations, such as DIM, SIM, and TIM, to enhance the transferability of adversarial examples.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Linkages between various adversarial attacks</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-3.tif"/>
</fig>
<p>Finally, the framework design of SAMI-FGSM is inspired by VMI-FGSM and inherits the key ideas in VMI-FGSM in design. While SAMI-FGSM and VMI-FGSM are on the same level, SAMI-FGSM simplifies the approach of VMI-FGSM and reduces computational overhead. The contribution of SAMI-FGSM is not a simple combination of existing techniques, but through normal sampling theory and gradient accumulation mechanism, it solves the fundamental limitations of traditional methods in local optimum trap and defense model attack efficiency. As presented in <xref ref-type="table" rid="table-2">Table 2</xref>, the proposed method generates adversarial examples for Inc-v3 that achieve the highest average attack success rate compared to six other models.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<p>This section begins with a comprehensive description of the datasets and models employed in the experiments. Next, a comprehensive comparison is made between the proposed approach and baseline attacks, focusing on single models and various input transformations. The experimental findings clearly show the advantages of our method. The impact of adversarial examples produced by our approach in comparison to three baseline attacks is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> illustrates the attention heatmaps of the original image and the adversarial examples generated by SAMI-FGSM on the Inc-v3 model. Finally, we discuss the parameter ablation study conducted on the proposed SAMI-FGSM. It is important to remember that the average success attack rates reported in all tables represent black-box attack performance, with (&#x002A;) indicating the results on white-box models.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The original image is shown in the first column, while the sample effect image of the confrontation created by our SAMI-FGSM method and the three baseline attacks on Inc-v3 model is shown in the remaining column. It is obvious that our method&#x2019;s attacks have a larger success rate than the baselines, yet the difference in visualization remains minimal</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Examples of model attention heatmaps generated by Grad-CAM [<xref ref-type="bibr" rid="ref-34">34</xref>] are used for clean images and for SAMI-FGSM generated adversarial examples</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-5.tif"/>
</fig>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup</title>
<p><bold>Dataset.</bold> Based on earlier works [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>], we randomly selected one image from every one of the 1000 categories in the ILSVRC2012 validation set, each selected image was able to be correctly categorized in all models in that paper. We also list other mainstream datasets, as shown in <xref ref-type="table" rid="table-1">Table 1</xref>. In the standardized adversarial attack research, ILSVRC2012 still has the only unified benchmark; Although other large-scale or diverse datasets have more advantages in the number of categories, they have not yet formed a unified evaluation standard in the adversarial attack community or the computational cost is too high.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Mainstream image classification datasets</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dateset</th>
<th>Categories</th>
<th>Typical application scenarios</th>
</tr>
</thead>
<tbody>
<tr>
<td>ILSVRC2012 (ImageNet)</td>
<td>1000</td>
<td>Standard classification, adversarial attack evaluation</td>
</tr>
<tr>
<td>CIFAR-10</td>
<td>10</td>
<td>Small scale classification benchmark</td>
</tr>
<tr>
<td>CIFAR-100</td>
<td>100</td>
<td>A benchmark for fine-grained classification</td>
</tr>
<tr>
<td>ImageNet-21K</td>
<td>11221</td>
<td>Semantic learning, multi-label classification</td>
</tr>
<tr>
<td>OpenImages V4</td>
<td>19794</td>
<td>Integrated classification, detection, and segmentation</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Models.</bold> To better contrast our method with the existing dominant methods, seven naturally trained models are used, four of which are typically trained models are Inception-v3 (Inc-v3) [<xref ref-type="bibr" rid="ref-35">35</xref>], Inception-v4 (Inc-v4) [<xref ref-type="bibr" rid="ref-36">36</xref>], Inception-Resnet-v2 (IncRes-v2) [<xref ref-type="bibr" rid="ref-36">36</xref>], Resnet-v2-101 (Res-101) [<xref ref-type="bibr" rid="ref-37">37</xref>], as well as three adversary-trained models that have been trained using adversarial examples, namely, ens3-adv-Inception-v3 (Inc-v3<sub>ens3</sub>) [<xref ref-type="bibr" rid="ref-26">26</xref>], ens4-Inception-v3 (Inc-v3<sub>ens4</sub>) [<xref ref-type="bibr" rid="ref-26">26</xref>], and ens-adv-Inception-ResNet-v2 (IncRes-v2<sub>ens</sub>) [<xref ref-type="bibr" rid="ref-26">26</xref>]. In the experiment of this paper, every one of these models functioned as a stand-in model to produce adversarial examples. In addition to the above CNNs, we also use transformer-based architectures including ViT [<xref ref-type="bibr" rid="ref-38">38</xref>], PiT [<xref ref-type="bibr" rid="ref-39">39</xref>], Visformer [<xref ref-type="bibr" rid="ref-40">40</xref>], Swin [<xref ref-type="bibr" rid="ref-41">41</xref>]. Besides, we utilized three cutting-edge defensive models to evaluate our strategy&#x2019;s attack performance: Neural Representation Purifier (NRP) [<xref ref-type="bibr" rid="ref-29">29</xref>], Bit-Reduction (Bit-Red) [<xref ref-type="bibr" rid="ref-30">30</xref>], Feature Distillation (FD) [<xref ref-type="bibr" rid="ref-31">31</xref>], Resize and Padding (RP) [<xref ref-type="bibr" rid="ref-42">42</xref>], HGD [<xref ref-type="bibr" rid="ref-43">43</xref>], and RS [<xref ref-type="bibr" rid="ref-44">44</xref>].</p>
<p><bold>Baseline.</bold> We consider eight gradient-based attacks as our baselines, including MI-FGSM, NI-FGSM, VMI-FGSM, S<inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msup><mml:mtext>M</mml:mtext><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>I-FGSM, NEAA [<xref ref-type="bibr" rid="ref-45">45</xref>], NAA [<xref ref-type="bibr" rid="ref-46">46</xref>] and MFAA [<xref ref-type="bibr" rid="ref-47">47</xref>]. Additionally, our method can be paired with various common input transformation attacks to test its compatibility and effectiveness.</p>
<p><bold>Hyperparameters.</bold> The parameters used in our experiments are consistent with those in the baseline attack methods. Specifically, the number of iterations, maximum perturbation, and step size are set to T &#x003D; 10, <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> &#x003D; 16/255, <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> &#x003D; 1.6/255, respectively. For the gradient-based momentum term (decay factor) in the baseline attacks, it is set to 1.0. For the three input transformation methods, the parameter settings are as follows: DIM has a transformation probability of 0.5, SIM involves 5 scale copies, and TIM uses a Gaussian kernel of size 7 <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 7, as established in previous studies. For VM(N)I-FGSM, the settings are consistent with the optimal attack performance reported in [<xref ref-type="bibr" rid="ref-11">11</xref>], where the domain upper bound factor <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is set to 1.5, and the number of samples for variance adjustment N is set to 20. For the method we propose, based on stochastic gradient accumulation, the sample size drawn from normal distribution is N &#x003D; 500, and the upper limit for sampling is set to 5/2.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Attack a Single Model</title>
<p>In this part, we tested several baseline methods and our proposed Stochastic Gradient Accumulation Momentum Iterative Attack (SAMI-FGSM) on individual deep neural network models. <xref ref-type="table" rid="table-2">Table 2</xref> shows that the Inc-v3 model was initially used to generate adversarial examples for the experiments. The adversarial examples were generated on the Inc-v3, Inc-v4, and IRes-v2 models, and then evaluated on eight different models, comprising one white-box model, two models trained without defenses, two adversarially trained defense models, and three models with advanced defensive techniques. <xref ref-type="table" rid="table-3">Table 3</xref> demonstrates that our SAMI-FGSM technique surpasses every baseline method in terms of attack success rates.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Experimental results of SAMI-FGSM with adversarial examples produced by baseline attacks under a single model on each of the seven models. The best results are bold</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Attack</th>
<th>Inc-v3<bold>&#x002A;</bold></th>
<th>Inc-v4</th>
<th>IncRes-v2</th>
<th>Res-101</th>
<th>Inc-v3<sub><bold>ens3</bold></sub></th>
<th>Inc-v3<sub><bold>ens4</bold></sub></th>
<th>IncRes-v2<sub><bold>ens</bold></sub></th>
<th>Average</th>
</tr>
</thead>
<tbody>
<tr>
<td>FGSM</td>
<td>67.2</td>
<td>25.7</td>
<td>26.0</td>
<td>24.5</td>
<td>10.2</td>
<td>10.4</td>
<td>4.5</td>
<td>16.9</td>
</tr>
<tr>
<td>I-FGSM</td>
<td><bold>100.0</bold></td>
<td>20.3</td>
<td>18.5</td>
<td>16.1</td>
<td>4.6</td>
<td>5.2</td>
<td>2.5</td>
<td>11.2</td>
</tr>
<tr>
<td>MI-FGSM</td>
<td><bold>100.0</bold></td>
<td>45.6</td>
<td>42.3</td>
<td>35.8</td>
<td>14.1</td>
<td>12.4</td>
<td>6.2</td>
<td>26.1</td>
</tr>
<tr>
<td>NI-FGSM</td>
<td><bold>100.0</bold></td>
<td>51.5</td>
<td>49.4</td>
<td>40.6</td>
<td>13.0</td>
<td>12.3</td>
<td>6.8</td>
<td>28.9</td>
</tr>
<tr>
<td>VMI-FGSM</td>
<td><bold>100.0</bold></td>
<td>71.4</td>
<td>68.5</td>
<td>60.0</td>
<td>32.7</td>
<td>30.6</td>
<td>17.4</td>
<td>46.8</td>
</tr>
<tr>
<td>VNI-FGSM</td>
<td><bold>100.0</bold></td>
<td>76.8</td>
<td>75.0</td>
<td>64.6</td>
<td>34.5</td>
<td>33.3</td>
<td>19.2</td>
<td>50.6</td>
</tr>
<tr>
<td>NAA</td>
<td>98.1</td>
<td>85.0</td>
<td>82.4</td>
<td>77.1</td>
<td>50.5</td>
<td>50.8</td>
<td>31.5</td>
<td>62.8</td>
</tr>
<tr>
<td>MFAA</td>
<td>97.6</td>
<td>86.5</td>
<td>84.6</td>
<td>78.1</td>
<td>51.9</td>
<td>46.2</td>
<td>32.5</td>
<td>63.3</td>
</tr>
<tr>
<td>NEAA</td>
<td>99.3</td>
<td>88.0</td>
<td><bold>87.2</bold></td>
<td>78.8</td>
<td>51.4</td>
<td>52.6</td>
<td>31.8</td>
<td>64.9</td>
</tr>
<tr>
<td>S<inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msup><mml:mtext>M</mml:mtext><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>I-FGSM</td>
<td>99.8</td>
<td>78.5</td>
<td>76.1</td>
<td>65.5</td>
<td>62.8</td>
<td>61.6</td>
<td>48.0</td>
<td>65.4</td>
</tr>
<tr>
<td><bold>SAMI-FGSM</bold></td>
<td>99.4</td>
<td><bold>88.1</bold></td>
<td>86.6</td>
<td><bold>82.8</bold></td>
<td><bold>71.3</bold></td>
<td><bold>70.3</bold></td>
<td><bold>54.3</bold></td>
<td><bold>75.6</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Typical attack success rates generated by the four baseline attacks and our methods on three models. The best results are bold</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Attack</th>
<th>Inc-v3</th>
<th>Inc-v4</th>
<th>IncRes-v2</th>
<th>Inc-v3<sub><bold>ens4</bold></sub></th>
<th>IncRes-v2<sub><bold>ens</bold></sub></th>
<th>FD</th>
<th>BIT</th>
<th>NRP</th>
<th>Average</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">Inc-v3</td>
<td>MI-FGSM</td>
<td><bold>100.0&#x002A;</bold></td>
<td>49.3</td>
<td>47.9</td>
<td>30.7</td>
<td>18.7</td>
<td>28.4</td>
<td>20.1</td>
<td>16.9</td>
<td>30.3</td>
</tr>
<tr>
<td>NI-FGSM</td>
<td>99.9&#x002A;</td>
<td>54.0</td>
<td>54.2</td>
<td>33.9</td>
<td>23.4</td>
<td>29.7</td>
<td>24.0</td>
<td>20.7</td>
<td>34.3</td>
</tr>
<tr>
<td>VMI-FGSM</td>
<td><bold>100.0&#x002A;</bold></td>
<td>69.0</td>
<td>66.6</td>
<td>48.0</td>
<td>41.7</td>
<td>46.9</td>
<td>37.9</td>
<td>42.8</td>
<td>50.4</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM</td>
<td>99.8&#x002A;</td>
<td>78.5</td>
<td>76.1</td>
<td>61.6</td>
<td>48.0</td>
<td>58.9</td>
<td>47.5</td>
<td>49.1</td>
<td>59.9</td>
</tr>
<tr>
<td>SAMI-FGSM</td>
<td>99.4&#x002A;</td>
<td><bold>88.1</bold></td>
<td><bold>86.6</bold></td>
<td><bold>70.3</bold></td>
<td><bold>54.3</bold></td>
<td><bold>61.5</bold></td>
<td><bold>52.1</bold></td>
<td><bold>53.7</bold></td>
<td><bold>66.7</bold></td>
</tr>
<tr>
<td rowspan="5">Inc-v4</td>
<td>MI-FGSM</td>
<td>43.1</td>
<td>99.3&#x002A;</td>
<td>34.1</td>
<td>21.7</td>
<td>12.8</td>
<td>21.1</td>
<td>17.6</td>
<td>15.4</td>
<td>23.7</td>
</tr>
<tr>
<td>NI-FGSM</td>
<td>46.9</td>
<td><bold>99.8&#x002A;</bold></td>
<td>38.0</td>
<td>21.5</td>
<td>12.7</td>
<td>21.0</td>
<td>19.6</td>
<td>18.3</td>
<td>25.5</td>
</tr>
<tr>
<td>VMI-FGSM</td>
<td>70.8</td>
<td>99.6&#x002A;</td>
<td>61.3</td>
<td>39.9</td>
<td>36.0</td>
<td>33.2</td>
<td>28.6</td>
<td>25.3</td>
<td>42.1</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM</td>
<td>76.0</td>
<td>99.5&#x002A;</td>
<td>67.8</td>
<td>46.5</td>
<td>38.3</td>
<td>39.8</td>
<td>33.6</td>
<td>29.2</td>
<td>47.3</td>
</tr>
<tr>
<td>SAMI-FGSM</td>
<td><bold>86.7</bold></td>
<td>97.9&#x002A;</td>
<td><bold>83.6</bold></td>
<td><bold>71.9</bold></td>
<td><bold>61.0</bold></td>
<td><bold>48.5</bold></td>
<td><bold>43.2</bold></td>
<td><bold>39.6</bold></td>
<td><bold>62.0</bold></td>
</tr>
<tr>
<td rowspan="5">IncRes-v2</td>
<td>MI-FGSM</td>
<td>43.6</td>
<td>36.2</td>
<td><bold>98.8&#x002A;</bold></td>
<td>22.2</td>
<td>18.7</td>
<td>19.9</td>
<td>15.0</td>
<td>16.4</td>
<td>24.6</td>
</tr>
<tr>
<td>NI-FGSM</td>
<td>45.8</td>
<td>39.5</td>
<td>97.0&#x002A;</td>
<td>22.7</td>
<td>19.5</td>
<td>21.8</td>
<td>18.9</td>
<td>19.3</td>
<td>26.8</td>
</tr>
<tr>
<td>VMI-FGSM</td>
<td>68.9</td>
<td>66.2</td>
<td>97.2&#x002A;</td>
<td>47.5</td>
<td>42.7</td>
<td>33.5</td>
<td>29.8</td>
<td>31.7</td>
<td>45.8</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM</td>
<td>73.1</td>
<td>69.3</td>
<td>97.5&#x002A;</td>
<td>52.3</td>
<td>49.8</td>
<td>42.8</td>
<td>36.5</td>
<td>40.1</td>
<td>51.9</td>
</tr>
<tr>
<td>SAMI-FGSM</td>
<td><bold>84.4</bold></td>
<td><bold>81.3</bold></td>
<td>93.5&#x002A;</td>
<td><bold>69.4</bold></td>
<td><bold>67.6</bold></td>
<td><bold>54.5</bold></td>
<td><bold>46.8</bold></td>
<td><bold>50.1</bold></td>
<td><bold>64.8</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Attack with Input Transformations</title>
<p>To increase the efficacy of adversarial attacks, Wang et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] demonstrated through experiments that DI (DIM), SI (SIM), and TI (TIM) attacks can be integrated into a composite transformation approach (DTS). This method, which integrates multiple transformations, can be paired with gradient-based attack techniques, resulting in enhanced attack success rates. Our approach, the Stochastic Gradient Accumulation Momentum Iterative method (SAMI-FGSM), focuses on improving the transferability of adversarial examples. We validate that our approach is just as applicable as classical gradient-based methods by combining this method with composite transformation techniques. <xref ref-type="table" rid="table-4">Table 4</xref> demonstrates that combining our method with different input transformations consistently improves attack performance.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Success rates of black-box attacks were generated on three models using our technique combined with DTS and by the usual four baseline attacks. The best results are bold</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Model</th>
<th align="center">Attack</th>
<th align="center">Inc-v3</th>
<th align="center">Inc-v4</th>
<th align="center">IncRes-v2</th>
<th align="center">Inc-v3<sub><bold>ens4</bold></sub></th>
<th align="center">IncRes-v2<sub><bold>ens</bold></sub></th>
<th align="center">FD</th>
<th align="center">BIT</th>
<th align="center">NRP</th>
<th align="center">Average</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">Inc-v3</td>
<td>MI-FGSM-DTS</td>
<td>99.7&#x002A;</td>
<td>82.9</td>
<td>79.9</td>
<td>71.0</td>
<td>57.3</td>
<td>70.4</td>
<td>44.2</td>
<td>47.9</td>
<td>64.8</td>
</tr>
<tr>
<td>NI-FGSM-DTS</td>
<td><bold>99.9&#x002A;</bold></td>
<td>84.0</td>
<td>81.9</td>
<td>70.9</td>
<td>58.1</td>
<td>70.9</td>
<td>45.0</td>
<td>44.6</td>
<td>65.1</td>
</tr>
<tr>
<td>VMI-FGSM-DTS</td>
<td>99.7&#x002A;</td>
<td>84.0</td>
<td>81.9</td>
<td>78.0</td>
<td>62.9</td>
<td>73.9</td>
<td>51.6</td>
<td>55.7</td>
<td>69.7</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM-DTS</td>
<td>99.7&#x002A;</td>
<td><bold>86.9</bold></td>
<td><bold>87.0</bold></td>
<td>79.7</td>
<td>68.1</td>
<td>78.6</td>
<td>56.9</td>
<td>57.8</td>
<td>73.7</td>
</tr>
<tr>
<td>SAMI-FGSM-DTS</td>
<td>98.8&#x002A;</td>
<td>86.2</td>
<td>86.2</td>
<td><bold>85.0</bold></td>
<td><bold>76.1</bold></td>
<td><bold>84.8</bold></td>
<td><bold>63.1</bold></td>
<td><bold>65.3</bold></td>
<td><bold>78.1</bold></td>
</tr>
<tr>
<td rowspan="5">Inc-v4</td>
<td>MI-FGSM-DTS</td>
<td>84.2</td>
<td>99.7&#x002A;</td>
<td>79.8</td>
<td>64.7</td>
<td>53.0</td>
<td>65.3</td>
<td>36.1</td>
<td>32.5</td>
<td>59.4</td>
</tr>
<tr>
<td>NI-FGSM-DTS</td>
<td>87.0</td>
<td>99.8&#x002A;</td>
<td>79.8</td>
<td>62.9</td>
<td>52.1</td>
<td>66.0</td>
<td>35.3</td>
<td>31.0</td>
<td>59.2</td>
</tr>
<tr>
<td>VMI-FGSM-DTS</td>
<td>88.9</td>
<td><bold>99.9&#x002A;</bold></td>
<td>83.8</td>
<td>72.9</td>
<td>62.3</td>
<td>71.0</td>
<td>42.7</td>
<td>38.9</td>
<td>65.8</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM-DTS</td>
<td>91.9</td>
<td><bold>99.9&#x002A;</bold></td>
<td>87.8</td>
<td>76.0</td>
<td>64.4</td>
<td>76.7</td>
<td>57.2</td>
<td>40.8</td>
<td>70.7</td>
</tr>
<tr>
<td>SAMI-FGSM-DTS</td>
<td><bold>93.1</bold></td>
<td>97.2&#x002A;</td>
<td><bold>90.5</bold></td>
<td><bold>84.3</bold></td>
<td><bold>71.9</bold></td>
<td><bold>80.2</bold></td>
<td><bold>65.8</bold></td>
<td><bold>51.0</bold></td>
<td><bold>76.7</bold></td>
</tr>
<tr>
<td rowspan="5">IncRes-v2</td>
<td>MI-FGSM-DTS</td>
<td>77.0</td>
<td>73.7</td>
<td>97.3&#x002A;</td>
<td>60.3</td>
<td>57.4</td>
<td>67.0</td>
<td>38.9</td>
<td>42.7</td>
<td>59.6</td>
</tr>
<tr>
<td>NI-FGSM-DTS</td>
<td>77.9</td>
<td>74.3</td>
<td>97.1&#x002A;</td>
<td>60.7</td>
<td>57.0</td>
<td>67.5</td>
<td>40.8</td>
<td>44.6</td>
<td>60.4</td>
</tr>
<tr>
<td>VMI-FGSM-DTS</td>
<td>79.2</td>
<td>77.8</td>
<td>97.0&#x002A;</td>
<td>66.3</td>
<td>62.7</td>
<td>70.9</td>
<td>48.8</td>
<td>50.4</td>
<td>65.2</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM-DTS</td>
<td>80.3</td>
<td>79.0</td>
<td><bold>98.1&#x002A;</bold></td>
<td>67.9</td>
<td>64.8</td>
<td>76.9</td>
<td>57.2</td>
<td>53.1</td>
<td>68.5</td>
</tr>
<tr>
<td>SAMI-FGSM-DTS</td>
<td><bold>85.1</bold></td>
<td><bold>82.5</bold></td>
<td>93.3&#x002A;</td>
<td><bold>82.0</bold></td>
<td><bold>82.4</bold></td>
<td><bold>80.2</bold></td>
<td><bold>65.6</bold></td>
<td><bold>63.4</bold></td>
<td><bold>77.3</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Attack an Ensemble of Models</title>
<p>Through the integration of multiple models, Liu et al. [<xref ref-type="bibr" rid="ref-48">48</xref>] demonstrated that such an approach boosts the attack performance of adversarial examples, significantly increasing their transferability. Typically, there exist three kinds of ensemble methods: ensemble at the prediction level, ensemble at the logit level, and ensemble within the loss function. In this work, we utilize the logit ensemble method, where we average the logit outputs from the Inc-v3, Inc-v4, and IncRes-v2 models. By making a slight sacrifice in attack performance in white-box attacks, we gain enhanced transferability in black-box attacks, where our proposed method exhibits optimal performance. We conduct experiments on two adversarially trained defense models, and six models with advanced defensive techniques, in <xref ref-type="table" rid="table-5">Table 5</xref>, the upper section reports the success attack rates (%) of four baseline attacks and our proposed method across three ensemble models, while the lower section shows the results when combined with DTS. Additionally, we combine the composite transformation method (DTS) with the ensemble approach to confirm the generality of our proposed approach. Compared to the four standard gradient-based attack techniques, our approach delivers superior performance, achieving an average success rate of 88.9%.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>The upper part shows the attack success rate (%) of four baseline attacks and our method produced at each of the three integrated models, and the lower part shows the effect of combining the attacks with DTS on this basis. The best results are bold</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Attack</th>
<th>Inc-v3</th>
<th>Inc-v4</th>
<th>IncRes-v2</th>
<th>Inc-v3<sub>ens4</sub></th>
<th>IncRes-v2<sub>ens</sub></th>
<th>FD</th>
<th>BIT</th>
<th>NRP</th>
<th>HGD</th>
<th>RP</th>
<th>RS</th>
<th>Average</th>
</tr>
</thead>
<tbody>
<tr>
<td>MI-FGSM</td>
<td><bold>100.0&#x002A;</bold></td>
<td>99.9&#x002A;</td>
<td><bold>100.0&#x002A;</bold></td>
<td>52.4</td>
<td>43.8</td>
<td>54.2</td>
<td>39.0</td>
<td>30.9</td>
<td>24.8</td>
<td>22.2</td>
<td>30.3</td>
<td>37.2</td>
</tr>
<tr>
<td>NI-FGSM</td>
<td><bold>100.0&#x002A;</bold></td>
<td><bold>100.0&#x002A;</bold></td>
<td><bold>100.0&#x002A;</bold></td>
<td>55.0</td>
<td>45.9</td>
<td>55.1</td>
<td>41.1</td>
<td>33.0</td>
<td>22.3</td>
<td>23.1</td>
<td>30.7</td>
<td>38.3</td>
</tr>
<tr>
<td>VMI-FGSM</td>
<td><bold>100.0&#x002A;</bold></td>
<td><bold>100.0&#x002A;</bold></td>
<td><bold>100.0&#x002A;</bold></td>
<td>78.9</td>
<td>75.1</td>
<td>70.9</td>
<td>56.0</td>
<td>48.6</td>
<td>54.3</td>
<td>50.6</td>
<td>35.6</td>
<td>58.8</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM</td>
<td>99.9&#x002A;</td>
<td>99.9&#x002A;</td>
<td>99.8&#x002A;</td>
<td>85.1</td>
<td>80.4</td>
<td>80.9</td>
<td>62.5</td>
<td>58.1</td>
<td>74.6</td>
<td>61.2</td>
<td>42.1</td>
<td>68.1</td>
</tr>
<tr>
<td>SAMI-FGSM</td>
<td>99.7&#x002A;</td>
<td>98.5&#x002A;</td>
<td>98.1&#x002A;</td>
<td><bold>88.2</bold></td>
<td><bold>82.9</bold></td>
<td><bold>82.8</bold></td>
<td><bold>65.6</bold></td>
<td><bold>61.2</bold></td>
<td><bold>80.5</bold></td>
<td><bold>81.3</bold></td>
<td><bold>43.2</bold></td>
<td><bold>73.2</bold></td>
</tr>
<tr>
<td>MI-FGSM-DTS</td>
<td>99.9&#x002A;</td>
<td><bold>100.0&#x002A;</bold></td>
<td>99.8&#x002A;</td>
<td>93.8</td>
<td>91.4</td>
<td>90.5</td>
<td>73.5</td>
<td>80.8</td>
<td>82.3</td>
<td>83.6</td>
<td>49.7</td>
<td>80.7</td>
</tr>
<tr>
<td>NI-FGSM-DTS</td>
<td><bold>100.0&#x002A;</bold></td>
<td>99.8&#x002A;</td>
<td><bold>100.0&#x002A;</bold></td>
<td>96.5</td>
<td>94.0</td>
<td>90.6</td>
<td>74.5</td>
<td>82.1</td>
<td>87.9</td>
<td>86.5</td>
<td>58.2</td>
<td>83.8</td>
</tr>
<tr>
<td>VMI-FGSM-DTS</td>
<td>99.8&#x002A;</td>
<td>99.9&#x002A;</td>
<td>99.9&#x002A;</td>
<td>95.4</td>
<td>94.8</td>
<td>91.9</td>
<td>78.3</td>
<td>82.9</td>
<td>90.3</td>
<td>91.0</td>
<td>63.6</td>
<td>86.1</td>
</tr>
<tr>
<td>SM<sup>2</sup>I-FGSM-DTS</td>
<td>99.9&#x002A;</td>
<td><bold>100.0&#x002A;</bold></td>
<td><bold>100.0&#x002A;</bold></td>
<td>96.5</td>
<td>95.1</td>
<td>92.9</td>
<td>80.6</td>
<td>84.5</td>
<td>91.9</td>
<td>91.4</td>
<td>67.4</td>
<td>87.5</td>
</tr>
<tr>
<td>SAMI-FGSM-DTS</td>
<td>98.1&#x002A;</td>
<td>98.0&#x002A;</td>
<td>97.1&#x002A;</td>
<td><bold>96.8</bold></td>
<td><bold>96.5</bold></td>
<td><bold>93.1</bold></td>
<td><bold>81.5</bold></td>
<td><bold>84.9</bold></td>
<td><bold>92.4</bold></td>
<td><bold>91.8</bold></td>
<td><bold>74.5</bold></td>
<td><bold>88.9</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Notably, our method consistently outperforms others on both single and ensemble models, demonstrating its effectiveness and highlighting the vulnerability of current defense models.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Attack a Transformer Architecture Model</title>
<p>In order to verify the generalization ability of SAMI-FGSM on Transformer architecture models, four typical vision Transformer models are selected as target models in this section: ViT, PiT, Visformer, and Swin Transformer. The experiment uses Inc-v3 as the source model to generate adversarial samples, and the attack results are shown in <xref ref-type="table" rid="table-6">Table 6</xref>. The average attack success rate of SAMI-FGSM on four Transformer models reaches 53.7%, which is 21.7% and 10.9% higher than the baseline methods VMI-FGSM (32.0%) and VNI-FGSM (42.8%), respectively. Experiments show that SAMI-FGSM has significant advantages on the Transformer architecture model.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Attack success rate (%) on Transformer model based on adversarial examples generated by Inc-v3. The best results are bold</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Attack</th>
<th>ViT</th>
<th>PiT</th>
<th>Visformer</th>
<th>Swin</th>
<th>Average</th>
</tr>
</thead>
<tbody>
<tr>
<td>FGSM</td>
<td>15.0</td>
<td>17.8</td>
<td>26.4</td>
<td>32.7</td>
<td>22.9</td>
</tr>
<tr>
<td>I-FGSM</td>
<td>4.9</td>
<td>10.0</td>
<td>14.6</td>
<td>21.7</td>
<td>12.8</td>
</tr>
<tr>
<td>MI-FGSM</td>
<td>17.2</td>
<td>23.8</td>
<td>33.7</td>
<td>42.5</td>
<td>29.3</td>
</tr>
<tr>
<td>NI-FGSM</td>
<td>16.6</td>
<td>21.5</td>
<td>33.3</td>
<td>43.2</td>
<td>28.7</td>
</tr>
<tr>
<td>VMI-FGSM</td>
<td>23.6</td>
<td>27.8</td>
<td>34.9</td>
<td>41.5</td>
<td>32.0</td>
</tr>
<tr>
<td>VNI-FGSM</td>
<td>26.3</td>
<td>35.9</td>
<td>52.5</td>
<td>56.3</td>
<td>42.8</td>
</tr>
<tr>
<td><bold>SAMI-FGSM</bold></td>
<td><bold>39.1</bold></td>
<td><bold>48.9</bold></td>
<td><bold>58.7</bold></td>
<td><bold>68.1</bold></td>
<td><bold>53.7</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Statistical Significance Test</title>
<p>To ensure the statistical reliability of the results, we performed a paired <italic>t</italic>-test between SAMI-FGSM and the baseline method, and each attack experiment was independently repeated 10 times. The significance test was performed using a two-sample <italic>t</italic>-test with significance level <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> &#x003D; 0.05 to verify whether the performance difference between SAMI-FGSM and baseline methods was significant. As shown in <xref ref-type="table" rid="table-7">Table 7</xref>, The two-sample <italic>t</italic>-test shows that SAMI-FGSM has a significantly higher attack success rate than VMI-FGSM on the defense model. In addition, in the cross-model transfer scenario, the average success rate of SAMI-FGSM is 28.8% higher than that of the baseline method, indicating that its performance improvement is highly stable.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Attack success rate (%) of VMI-FGSM and SAMI-FGSM in 10 independent repeated runs</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Attack</th>
<th>Inc-v3&#x002A;</th>
<th>Inc-v4</th>
<th>IncRes-v2</th>
<th>Res-101</th>
<th>Inc-v3<sub><bold>ens3</bold></sub></th>
<th>Inc-v3<sub><bold>ens4</bold></sub></th>
<th>IncRes-v2<sub><bold>ens</bold></sub></th>
<th>Average</th>
</tr>
</thead>
<tbody>
<tr>
<td>VMI-FGSM</td>
<td>100.0</td>
<td>71.1 &#x00B1; 1.2</td>
<td>68.5 &#x00B1; 1.5</td>
<td>60.2 &#x00B1; 0.6</td>
<td>32.7 &#x00B1; 0.9</td>
<td>30.6 &#x00B1; 0.4</td>
<td>17.4 &#x00B1; 1.1</td>
<td>46.8 &#x00B1; 0.9</td>
</tr>
<tr>
<td><bold>SAMI-FGSM</bold></td>
<td>99.4</td>
<td>88.1 &#x00B1; <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msup><mml:mn>0.7</mml:mn><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>86.7 &#x00B1; <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msup><mml:mn>0.4</mml:mn><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>82.8 &#x00B1; <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msup><mml:mn>1.2</mml:mn><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>71.3 &#x00B1; <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msup><mml:mn>0.8</mml:mn><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>70.3 &#x00B1; <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msup><mml:mn>0.9</mml:mn><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>54.3 &#x00B1; <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:msup><mml:mn>0.3</mml:mn><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>75.6 &#x00B1; <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msup><mml:mn>0.7</mml:mn><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-7fn1" fn-type="other">
<p><inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msup><mml:mtext>Note:</mml:mtext><mml:mrow><mml:mo>&#x2020;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>Indicates that the <italic>p</italic>-value &#x003C; 0.05 is statistically significant compared with the baseline method (VMI-FGSM).</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Evaluation of Disturbance Perceptibility and Verification of Robustness to Input Transformations</title>
<p>We invite 20 subjects to blind test 100 pairs of original/adversarial examples. The experimental results show that less than 10% of the samples are correctly distinguished by the subjects, which further verifies the perceptual imperceptibility of SAM I-FGSM under the constraint of <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> &#x003D; 16/255. In addition, we also tested the robustness of perturbation to Gaussian noise (<inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> &#x003D; 0.1), motion blur (kernel &#x003D; 5 <sub>&#x002A;</sub> 5) and random cropping (20%). As shown in <xref ref-type="table" rid="table-8">Table 8</xref>, the proposed method has strong robustness to noise, blurring and cropping on the premise of maintaining low perception.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Experiments on robustness to input transformations</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Input transformations</th>
<th>Change of attack success rate</th>
</tr>
</thead>
<tbody>
<tr>
<td>Gaussian noise</td>
<td>&#x2212;3.2</td>
</tr>
<tr>
<td>Motion blur</td>
<td>&#x2212;8.7</td>
</tr>
<tr>
<td>Random cropping</td>
<td>&#x2212;8.4</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_8">
<label>4.8</label>
<title>Parameters Ablation Study</title>
<p>This section presents ablation studies on two parameters of SAMI-FGSM to assess its effectiveness. First, we assess the influence of the two parameters, sampling limit <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> and sampling number N, on the attack performance of SAMI-FGSM. To evaluate the influence of these hyperparameters, adversarial examples are created Utilizing Inc-v3 as a source model, with default settings of <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> &#x003D; 5/2 and N &#x003D; 500.</p>
<p><bold>Sampling Limit <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula>:</bold> As shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, we analyze how the sampling limit <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> in the normal distribution affects the performance of black-box transferability. The left plot shows adversarial examples generated using Inc-v3, while the right plot shows results on the Inc-v4 model. The sampling number N is fixed at 500. When <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> &#x003D; 0, SAMI-FGSM degenerates into MI-FGSM, resulting in lower transferability. When <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> &#x003D; 1/5, despite the small sampling limit, SAMI-FGSM exhibits a significant improvement in attack performance. When <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> &#x003D; 5/2, our method achieves optimal attack performance, showing excellent effectiveness even against several defense models. Finally, as <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> increases further, the attack performance of the proposed method gradually decreases. <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> controls the normal sampling limit, and its value needs to maintain a proportional relationship with the perturbation constraint value <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula>. When <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> &#x003D; 5/2, it ensures that the sampling points not only contain local feature disturbances, but also avoids the introduction of irrelevant noise due to the large range. Therefore, <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> &#x003D; 5/2 is chosen for all experiments in this study.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Success rate of transferability attacks on the remaining six models by SAMI-FGSM produced adversarial examples on Inc-v3 or Inc-v4 when <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is changed</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-6.tif"/>
</fig>
<p><bold>Sample Size N:</bold> Next, we examine how the number of samples (N) in the neighborhood affects the transferability of adversarial examples. As seen in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, the left plot represents adversarial examples created using Inc-v3, while the right plot shows outcomes on the Inc-v4 model. The sampling limit <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is set to 5/2 by default. Our approach also degenerates into MI-FGSM when N &#x003D; 0. The attack performance of our method will be significantly affected by the sampling number when N &#x003D; 20, with the effectiveness of the black-box attack increasing significantly as N increases. Our method achieves near-optimal attack performance at N &#x003D; 500. Although there is a slight increase in success attack rates with further increases in N, each iteration requires extensive sampling and gradient computation. As shown in <xref ref-type="table" rid="table-9">Table 9</xref>, when N increases from 500 to 1000, the average attack success rate increases by 1.9%, but the running time increases by 98.2%. When N &#x003E; 500, the success rate increases slowly, while the time cost increases significantly. Thus, a larger N results in higher computational costs. In our experiments, we set N &#x003D; 500 to strike a compromise between attack success and computational efficiency.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Success rate of transferability attacks on this remaining six models by SAMI-FGSM produced adversarial examples on Inc-v3 or Inc-v4 when sample size (N) is changed</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_64896-fig-7.tif"/>
</fig><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Running time and attack success rate under different sample number N (GPU uses one NVIDIA A40)</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>N</th>
<th>Time (s)</th>
<th>Average</th>
</tr>
</thead>
<tbody>
<tr>
<td>500</td>
<td>840</td>
<td>75.6</td>
</tr>
<tr>
<td>800</td>
<td>1339</td>
<td>76.8</td>
</tr>
<tr>
<td>1000</td>
<td>1665</td>
<td>77.5</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In summary, when N &#x003E; 500, the impact of N on black-box attack effectiveness gradually diminishes, while the parameter <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> has an important impact on the success attack rate. Therefore, for all experiments, we choose <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> &#x003D; 5/2 and N &#x003D; 500.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This paper introduces a novel Stochastic Gradient Accumulation Momentum Iterative Attack (SAMI-FGSM) to enhance the transferability of adversarial examples. This approach stabilizes the gradient update direction by calculating the accumulated gradient of random samples during each iteration, effectively avoiding local optima and achieving higher transferability. The attack efficiency of adversarial examples may also be boosted by integrating this method with other optimization-based attack techniques. Extensive experimental results demonstrate that SAMI-FGSM achieves optimal attack performance under both single-model and multi-model settings. Furthermore, combining our method with various input transformations further enhances attack success rate. Finally, ensemble experiments using three models validate the efficacy of SAMI-FGSM and show that it also achieves superior attack performance. Statistical tests further confirm that the performance improvement of SAMI-FGSM does not fluctuate by chance. This highlights the vulnerabilities of current defense models, underscoring the need to develop more robust defense strategies.</p>
<p>Although our experiments are based on the ImageNet dataset, the design principle of SAMI-FGSM has broad applicability and can be extended to other image modalities: medical images usually contain high-resolution local features and low SNR regions [<xref ref-type="bibr" rid="ref-49">49</xref>]. The normal distribution sampling of SAMI-FGSM can focus on the subtle disturbances in the lesion area, and the gradient accumulation mechanism can alleviate the overfitting problem caused by data scarcity. Satellite images have large-scale spatial heterogeneity and multi-spectral characteristics [<xref ref-type="bibr" rid="ref-50">50</xref>]. The local sampling strategy of SAMI-FGSM can specifically perturb the key areas of ground cover classification, and the cumulative gradient can adapt to the complex decision boundaries of the multi-band model. In the future, we plan to conduct experiments on medical images and satellite images to verify the attack effect of SAMI-FGSM against the scene classification and segmentation model.</p>
<p>When improving adversarial attack performance through stochastic gradient accumulation, although the method based on stochastic gradient accumulation significantly enhances black-box transferability, it is still necessary to explore more distribution sampling methods to determine whether the normal sampling process is optimal. In the gradient accumulation process, as the number of samples increases, the attack success rate also gradually increases. However, the reason why the attack performance reaches a peak at a certain number of samples needs further research and discussion. SAMI-FGSM also has scenarios or potential weaknesses that may perform poorly. When the target model employs random preprocessing such as random cropping and image enhancement to obfuscate the gradients, SAMI-FGSM may suffer from interference in the sampled gradient estimation, thereby reducing the attack success rate. Existing experiments mainly focus on CNN classification models. In object detection or segmentation tasks, there are modules such as anchor box mechanism or self-attention in the model structure, and SAMI-FGSM may not be directly applicable or have poor effects. A key limitation of SAMI-FGSM lies in the computational overhead introduced by the normal sampling process. In comparison to previous methods, generating a high number of samples for gradient accumulation dramatically increases the per-iteration cost. To address the computational overhead, future work will explore adaptive sampling techniques that focus on informative regions to reduce unnecessary computations. Parallel and distributed implementations may further accelerate gradient accumulation, enabling scalability for large-scale tasks. At present, the hyperparameters <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> and N of SAMI-FGSM are considered as fixed values, but there may be differences in the optimal values between different models and datasets. In the future, an adaptive tuning framework based on Bayesian optimization or reinforcement learning can be introduced to realize adaptive hyperparameter adjustment in the attack process, so as to improve the attack efficiency and success rate. It is also possible to extend SAMI-FGSM to physical adversarial attacks and multi-modal datasets. In other domains, we will try to adapt SAMI-FGSM to text classification and machine translation tasks in the future. Although experiments show that normal distribution sampling significantly improves the transferability of adversarial examples, its theoretical optimality has not been rigorously proved mathematically. This limitation comes from the high-dimensional non-convex optimization characteristics of adversarial attacks, and its theoretical analysis requires more in-depth functional analysis and probability theory tools. In our future work, we will cooperate with scholars in the mathematical field to give priority to solving this problem and establish a universal theoretical framework for the distributed design of adversarial attacks.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was supported in part by the National Natural Science Foundation (62202118, U24A20241); in part by Major Scientific and Technological Special Project of Guizhou Province ([2024]014, [2024]003); in part by Scientific and Technological Research Projects from Guizhou Education Department (Qian jiao ji [2023]003); in part by Guizhou Science and Technology Department Hundred Level Innovative Talents Project (GCC[2023]018).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Haolang Feng; methodology, Haolang Feng, Yuling Chen, Yang Huang; validation, Yuling Chen, Yang Huang; formal analysis, Yang Huang; investigation, Haolang Feng; resources, Haolang Feng; data curation, Haolang Feng; writing&#x2014;original draft preparation, Haolang Feng; writing&#x2014;review and editing, Haolang Feng; visualization, Xuewei Wang; supervision, Haiwei Sang; project administration, Haiwei Sang; funding acquisition, Xuewei Wang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are openly available in <ext-link ext-link-type="uri" xlink:href="https://github.com/FHL000/SAMI-FGSM">https://github.com/FHL000/SAMI-FGSM</ext-link> (accessed on 21 May 2025).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Krizhevsky</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sutskever</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name></person-group>. <article-title>Imagenet classification with deep convolutional neural networks</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2012</year>;<volume>25</volume>(<issue>6</issue>):<fpage>84</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3065386</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Simonyan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <comment>arXiv:1409.1556. 2014</comment>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>XB</given-names></string-name>, <string-name><surname>Duong</surname> <given-names>CN</given-names></string-name>, <string-name><surname>Xin</surname> <given-names>L</given-names></string-name>, <string-name><surname>Susan</surname> <given-names>G</given-names></string-name>, <string-name><surname>Han-Seok</surname> <given-names>S</given-names></string-name>, <string-name><surname>Luu</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Micron-BERT: BERT-based facial micro-expression recognition</article-title>. In: <conf-name>2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17&#x2013;24; Vancouver, BC, Canada</conf-name>. p. <fpage>1482</fpage>&#x2013;<lpage>92</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Physical-world optical adversarial attacks on 3D face recognition</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2023 Jun 17&#x2013;24; Vancouver, BC, Canada</conf-name>. p. <fpage>24699</fpage>&#x2013;<lpage>708</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VAD: vectorized scene representation for efficient autonomous driving</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision; 2023 Oct 1&#x2013;6; Paris, France</conf-name>. p. <fpage>8340</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jia</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>PL</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name></person-group>. <article-title>DriveAdapter: breaking the coupling barrier of perception and planning in end-to-end autonomous driving</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision; 2023 Oct 1&#x2013;6; Paris, France</conf-name>. p. <fpage>7953</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ning</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hong</surname> <given-names>K</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Ada3D: exploiting the spatial redundancy with adaptive inference for efficient 3D object detection</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision; 2023 Oct 1&#x2013;6; Paris, France</conf-name>. p. <fpage>17728</fpage>&#x2013;<lpage>38</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Su</surname> <given-names>H</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Composite adversarial attacks</article-title>. In: <conf-name>Proceedings of the 2021 AAAI Conference on Artificial Intelligence</conf-name>; <year>2021 Feb 2&#x2013;9</year>; <publisher-loc>Online</publisher-loc>. p. <fpage>8884</fpage>&#x2013;<lpage>92</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Larson</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Towards large yet imperceptible adversarial image perturbations with perceptual color distance</article-title>. In: <conf-name>Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2020 Jun 13&#x2013;19; Seattle, WA, USA</conf-name>. p. <fpage>1039</fpage>&#x2013;<lpage>48</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>H</given-names></string-name></person-group>. <article-title>GLH: from global to local gradient attacks with high-frequency momentum guidance for object detection</article-title>. <source>Entropy</source>. <year>2023</year>;<volume>25</volume>(<issue>3</issue>):<fpage>461</fpage>. doi:<pub-id pub-id-type="doi">10.3390/e25030461</pub-id>; <pub-id pub-id-type="pmid">36981349</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Enhancing the transferability of adversarial attacks through variance tuning</article-title>. In: <conf-name>Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2021 Jun 20&#x2013;25; Nashville, TN, USA</conf-name>. p. <fpage>1924</fpage>&#x2013;<lpage>33</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Improving adversarial transferability with spatial momentum</article-title>. <comment>arXiv:2203.13479. 2022</comment>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Shlens</surname> <given-names>J</given-names></string-name>, <string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Explaining and harnessing adversarial examples</article-title>. <comment>arXiv:1412.6572. 2014</comment>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Su</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Boosting adversarial attacks with momentum</article-title>. In: <conf-name>Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition; 2018 Jun 18&#x2013;23; Salt Lake City, UT, USA</conf-name>. p. <fpage>9185</fpage>&#x2013;<lpage>93</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Song</surname> <given-names>C</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hopcroft</surname> <given-names>JE</given-names></string-name></person-group>. <article-title>Nesterov accelerated gradient and scale invariance for adversarial attacks</article-title>. <comment>arXiv:1908.06281. 2019</comment>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Boosting the transferability of adversarial attacks with global momentum initialization</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>255</volume>(<issue>5</issue>):<fpage>124757</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.124757</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Yet another intermediate-level attack</article-title>. In: <conf-name>Computer Vision&#x2013;ECCV 2020: 16th European Conference, 2020 Aug 23&#x2013;28; Glasgow, UK</conf-name>. p. <fpage>241</fpage>&#x2013;<lpage>57</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Long</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Frequency domain model augmentation for adversarial attack</article-title>. In: <conf-name>European Conference on Computer Vision</conf-name>; <year>2022</year>; <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>. p. <fpage>549</fpage>&#x2013;<lpage>66</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Improving the transferability of adversarial examples via direction tuning</article-title>. <comment>arXiv:2303.15109. 2023</comment>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Feature importance-aware transferable adversarial attacks</article-title>. In: <conf-name>Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision; 2021 Oct 10&#x2013;17; Montreal, QC, Canada</conf-name>. p. <fpage>7639</fpage>&#x2013;<lpage>48</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Katsman</surname> <given-names>I</given-names></string-name>, <string-name><surname>He</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Belongie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lim</surname> <given-names>SN</given-names></string-name></person-group>. <article-title>Enhancing adversarial example transferability with an intermediate level attack</article-title>. In: <conf-name>Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision; 2019 Oct 27&#x2013;Nov 2; Seoul, Republic of Korea</conf-name>. p. <fpage>4733</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sui</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Boosting adversarial transferability via gradient relevance attack</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision; 2023 Oct 1&#x2013;6; Paris, France</conf-name>. p. <fpage>4741</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Su</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Evading defenses to transferable adversarial examples by translation-invariant attacks</article-title>. In: <conf-name>Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2019 Jun 15&#x2013;20; Long Beach, CA, USA</conf-name>. p. <fpage>4312</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Improving transferability of adversarial examples with input diversity</article-title>. In: <conf-name>Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2019 Jun 15&#x2013;20; Long Beach, CA, USA</conf-name>. p. <fpage>2730</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sakurai</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Versatile defense against adversarial attacks on image recognition</article-title>. <comment>arXiv:2403.08170. 2024</comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tram&#x00E8;r</surname> <given-names>F</given-names></string-name>, <string-name><surname>Kurakin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Papernot</surname> <given-names>N</given-names></string-name>, <string-name><surname>Goodfellow</surname> <given-names>I</given-names></string-name>, <string-name><surname>Boneh</surname> <given-names>D</given-names></string-name>, <string-name><surname>McDaniel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Ensemble adversarial training: attacks and defenses</article-title>. <comment>arXiv:1705.07204. 2017</comment>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Shafahi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Najibi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ghiasi</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Dickerson</surname> <given-names>J</given-names></string-name>, <string-name><surname>Studer</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Adversarial training for free! In:  Advances in neural information processing systems</article-title>; <year>2019</year>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1904.12843</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gokhale</surname> <given-names>T</given-names></string-name>, <string-name><surname>Anirudh</surname> <given-names>R</given-names></string-name>, <string-name><surname>Kailkhura</surname> <given-names>B</given-names></string-name>, <string-name><surname>Thiagarajan</surname> <given-names>JJ</given-names></string-name>, <string-name><surname>Baral</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Attribute-guided adversarial training for robustness to natural perturbations</article-title>. In: <conf-name>Proceedings of the 2021 AAAI Conference on Artificial Intelligence</conf-name>; <year>2021 Feb 2&#x2013;9</year>; <publisher-loc>Online</publisher-loc>. p. <fpage>7574</fpage>&#x2013;<lpage>82</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Naseer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hayat</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>FS</given-names></string-name>, <string-name><surname>Porikli</surname> <given-names>F</given-names></string-name></person-group>. <article-title>A self-supervised approach for adversarial robustness</article-title>. In: <conf-name>Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2020 Jun 13&#x2013;19; Seattle, WA, USA</conf-name>. p. <fpage>262</fpage>&#x2013;<lpage>71</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Feature squeezing: detecting adversarial examples in deep neural networks</article-title>. <comment>arXiv:1704.01155. 2017</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Feature distillation: DNN-oriented JPEG compression against adversarial examples</article-title>. In: <conf-name>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <year>2019 Jun 15&#x2013;20</year>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>860</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Kurakin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>S</given-names></string-name></person-group>. <chapter-title>Adversarial examples in the physical world</chapter-title>. In: <source>Artificial intelligence safety and security</source>. <publisher-loc>Boca Raton, FL, USA</publisher-loc>: <publisher-name>Chapman and Hall/CRC</publisher-name>; <year>2018</year>. p. <fpage>99</fpage>&#x2013;<lpage>112</lpage></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>YG</given-names></string-name></person-group>. <article-title>Enhancing the self-universality for transferable targeted attacks</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2023 Jun 17&#x2013;24; Vancouver, BC, Canada</conf-name>. p. <fpage>12281</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Selvaraju</surname> <given-names>RR</given-names></string-name>, <string-name><surname>Cogswell</surname> <given-names>M</given-names></string-name>, <string-name><surname>Das</surname> <given-names>A</given-names></string-name>, <string-name><surname>Vedantam</surname> <given-names>R</given-names></string-name>, <string-name><surname>Parikh</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Grad-cam: visual explanations from deep networks via gradient-based localization</article-title>. In: <conf-name>Proceedings of the 2017 IEEE International Conference on Computer Vision; 2017 Oct 22&#x2013;29; Venice, Italy</conf-name>. p. <fpage>618</fpage>&#x2013;<lpage>26</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Vanhoucke</surname> <given-names>V</given-names></string-name>, <string-name><surname>Ioffe</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shlens</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wojna</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Rethinking the inception architecture for computer vision</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition; 2016 Jun 27&#x2013;30; Las Vegas, NV, USA</conf-name>. p. <fpage>2818</fpage>&#x2013;<lpage>26</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ioffe</surname> <given-names>S</given-names></string-name>, <string-name><surname>Vanhoucke</surname> <given-names>V</given-names></string-name>, <string-name><surname>Alemi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Inception-V4, Inception-ResNet and the impact of residual connections on learning</article-title>. In: <conf-name>AAAI&#x2019;17: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence; 2017 Feb 4&#x2013;9; San Francisco, CA, USA</conf-name>. p. <fpage>4278</fpage>&#x2013;<lpage>84</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition; 2016 Jun 27&#x2013;30; Las Vegas, NV, USA</conf-name>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Beyer</surname> <given-names>L</given-names></string-name>, <string-name><surname>Kolesnikov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Weissenborn</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Unterthiner</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>. <comment>arXiv: 2010.11929. 2020</comment>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Heo</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yun</surname> <given-names>S</given-names></string-name>, <string-name><surname>Han</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chun</surname> <given-names>S</given-names></string-name>, <string-name><surname>Choe</surname> <given-names>J</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>SJ</given-names></string-name></person-group>. <article-title>Rethinking spatial dimensions of vision transformers</article-title>. In: <conf-name>Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision; 2021 Oct 10&#x2013;17; Montreal, QC, Canada</conf-name>. p. <fpage>11936</fpage>&#x2013;<lpage>45</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Niu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Visformer: the vision-friendly transformer</article-title>. In: <conf-name>Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision; 2021 Oct 10&#x2013;17; Montreal, QC, Canada</conf-name>. p. <fpage>589</fpage>&#x2013;<lpage>98</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Swin transformer: hierarchical vision transformer using shifted windows</article-title>. In: <conf-name>Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision; 2021 Oct 10&#x2013;17; Montreal, QC, Canada</conf-name>. p. <fpage>10012</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yuille</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Mitigating adversarial effects through randomization</article-title>. <comment>arXiv:1711.01991. 2017</comment>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Defense against adversarial attacks using high-level representation guided denoiser</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018 Jun 18&#x2013;23; Salt Lake City, UT, USA</conf-name>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cohen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rosenfeld</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kolter</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Certified adversarial robustness via randomized smoothing</article-title>. In: <conf-name>ICML 2019: 36th International Conference on Machine Learning; 2019 Jun 10&#x2013;15; Long Beach, CA, USA</conf-name>. p. <fpage>1310</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ke</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>D</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Min</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Improving the transferability of adversarial examples through neighborhood attribution</article-title>. <source>Knowl Based Syst</source>. <year>2024</year>;<volume>296</volume>:<fpage>111909</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2024.111909</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>JT</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Su</surname> <given-names>Y</given-names></string-name> <etal>et al</etal></person-group>. <article-title>Improving adversarial transferability via neuron attribution-based attacks</article-title>. In: <conf-name>Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2022 Jun 18&#x2013;24; New Orleans, LA, USA</conf-name>. p. <fpage>14993</fpage>&#x2013;<lpage>5002</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ke</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>G</given-names></string-name>, <string-name><surname>Min</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Enhancing the transferability of adversarial attacks via multi-feature attention</article-title>. <source>IEEE Trans Inf Forensics Secur</source>. <year>2025</year>;<volume>20</volume>:<fpage>1462</fpage>&#x2013;<lpage>74</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tifs.2025.3526067</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Song</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Delving into transferable adversarial examples and black-box attacks</article-title>. <comment>arXiv:1611.02770. 2016</comment>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Survey on adversarial attack and defense for medical image analysis: methods and challenges</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>57</volume>(<issue>3</issue>):<fpage>79</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3702638</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chin</surname> <given-names>TJ</given-names></string-name>, <string-name><surname>Law</surname> <given-names>YW</given-names></string-name>, <string-name><surname>Sasdelli</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rajasegaran</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Physical adversarial attacks on an aerial imagery object detector</article-title>. In: <conf-name>Proceedings of the 2022 IEEE/CVF Winter Conference on Applications of Computer Vision; 2022 Jan 3&#x2013;8; Waikoloa, HI, USA</conf-name>. p. <fpage>1796</fpage>&#x2013;<lpage>806</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>