<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">70990</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.070990</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>ResghostNet: Boosting GhostNet with Residual Connections and Adaptive-SE Blocks</article-title>
<alt-title alt-title-type="left-running-head">ResghostNet: Boosting GhostNet with Residual Connections and Adaptive-SE Blocks</alt-title>
<alt-title alt-title-type="right-running-head">ResghostNet: Boosting GhostNet with Residual Connections and Adaptive-SE Blocks</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0009-0001-5570-0034</contrib-id>
<name name-style="western"><surname>Chen</surname><given-names>Yuang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Yong</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>liyong@nudt.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Lin</surname><given-names>Fang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Lv</surname><given-names>Shuhan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Jiang</surname><given-names>Jiaze</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Key Laboratory of CTC &#x0026; IE (Engineering University of PAP), Ministry of Education</institution>, <addr-line>Xi&#x2019;an, 710086</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Graduate Student Brigade, Engineering University of PAP</institution>, <addr-line>Xi&#x2019;an, 710086</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yong Li. Email: <email>liyong@nudt.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>09</day><month>12</month><year>2025</year>
</pub-date>
<volume>86</volume>
<issue>2</issue>
<fpage>1</fpage>
<lpage>18</lpage>
<history>
<date date-type="received">
<day>29</day>
<month>07</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>09</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_70990.pdf"></self-uri>
<abstract>
<p>Aiming at the problem of potential information noise introduced during the generation of ghost feature maps in GhostNet, this paper proposes a novel lightweight neural network model called ResghostNet. This model constructs the Resghost Module by combining residual connections and Adaptive-SE Blocks, which enhances the quality of generated feature maps through direct propagation of original input information and selection of important channels before cheap operations. Specifically, ResghostNet introduces residual connections on the basis of the Ghost Module to optimize the information flow, and designs a weight self-attention mechanism combined with SE blocks to enhance feature expression capabilities in cheap operations. Experimental results on the ImageNet dataset show that, compared to GhostNet, ResghostNet achieves higher accuracy while reducing the number of parameters by 52%. Although the computational complexity increases, by optimizing the usage strategy of GPU cache memory, the model&#x2019;s inference speed becomes faster. The ResghostNet is optimized in terms of classification accuracy and the number of model parameters, and shows great potential in edge computing devices.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Residual connections</kwd>
<kwd>adaptive-SE blocks</kwd>
<kwd>lightweight neural network</kwd>
<kwd>GPU memory usage</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Science and Technology Innovation</funding-source>
<award-id>ZZKY20222304</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the rapid development of artificial intelligence technologies, the demand for lightweight models in edge computing devices is growing urgently [<xref ref-type="bibr" rid="ref-1">1</xref>], especially in scenarios such as industrial monitoring [<xref ref-type="bibr" rid="ref-2">2</xref>], equipment status perception [<xref ref-type="bibr" rid="ref-3">3</xref>], and real-time safety early warning [<xref ref-type="bibr" rid="ref-4">4</xref>], where efficient and low-latency vision models have become a key support for realizing intelligent upgrades. Traditional convolutional neural networks improve performance by increasing network depth, expanding parameter scales, and optimizing growing network structures [<xref ref-type="bibr" rid="ref-5">5</xref>], whereas lightweight models need to achieve high performance computing in resource-constrained environments [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>]. This fundamental difference in design philosophy makes the design of lightweight models more challenging.</p>
<p>Currently, the main methods for model lightweighting include pruning [<xref ref-type="bibr" rid="ref-8">8</xref>], architecture design [<xref ref-type="bibr" rid="ref-6">6</xref>], neural architecture search (NAS) [<xref ref-type="bibr" rid="ref-9">9</xref>], and knowledge distillation [<xref ref-type="bibr" rid="ref-10">10</xref>]. Based on these methods, researchers have designed many classic lightweight models, such as the MobileNet series [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>], ShuffleNet series [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>], SqueezeNet [<xref ref-type="bibr" rid="ref-16">16</xref>], and GhostNet [<xref ref-type="bibr" rid="ref-17">17</xref>], etc. These classic lightweight models typically combine pointwise convolutions and depthwise convolutions to reduce computational complexity and achieve weight reduction. Among them, GhostNet creatively introduced the Ghost module in its architecture, using a combination of primary convolutions and cheap operations instead of pointwise convolutions to achieve a breakthrough in feature generation mechanisms.</p>
<p>However, there are several key issues in this innovative design that deserve in-depth exploration: What is the basis for generating ghost feature maps using the Ghost module? Can all original feature maps generate ghost feature maps? Will some ghost feature maps become noise during feature extraction?</p>
<p>In response to these questions, this paper conducts a new exploration of the method for GhostNet to generate ghost feature maps and designs the ResghostNet model. Its design concept is as follows:
<list list-type="order">
<list-item>
<p>Transform the original Ghost Module into a residual connection [<xref ref-type="bibr" rid="ref-18">18</xref>], named Resghost Module, to retain original input information and reduce the impact of low-quality ghost feature maps on subsequent feature extraction. This improvement ensures that key feature information can be directly passed to deeper network layers.</p></list-item>
<list-item>
<p>The SE module is added to the shortcut connection branch of the Resghost Module to improve the model&#x2019;s expressive ability and overall performance.</p></list-item>
<list-item>
<p>Design an Adaptive SE block and introduce it before cheap operations to enhance the expressive capabilities of high-quality feature maps in cheap operations.</p></list-item>
</list></p>
<p>As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the model architecture designed in this paper is based on the above design methods. Although the computational complexity has increased, the additional computations occur in the cached memory on the GPU during runtime. Meanwhile, the reduced number of parameters alleviates the pressure on the allocated memory on the GPU [<xref ref-type="bibr" rid="ref-19">19</xref>]. Therefore, compared with GhostNet, the inference speed has even become faster.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Architecture of ResghostNet. The leftmost part is the backbone of ResghostNet. In the middle section, the dark-green backgrounds denote Resghost bottlenecks with an SE block, while the light-green backgrounds indicate Resghost bottlenecks without an SE block. The rightmost part is the Resghost Module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-1.tif"/>
</fig>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>GhostNet</title>
<p>Han et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed GhostNet, an efficient lightweight convolutional neural network architecture in 2020, aiming to reduce computational costs through cheap operations while maintaining high performance. As shown in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>, GhostNet consists of 16 Ghost bottlenecks, mainly used for feature extraction, with the Ghost Module as its core component. As shown in <xref ref-type="fig" rid="fig-2">Fig. 2b</xref>, the operation steps of the Ghost Module are to first perform a small number of ordinary convolution operations on the input feature maps to generate some basic feature maps, and then apply a series of cheap operations to these basic feature maps to generate more ghost feature maps. The ingenuity of this design lies in achieving effects similar to pointwise convolutions with fewer computational resources.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The two-layer structure within GhostNet: (<bold>a</bold>) GhostNet bottleneck; (<bold>b</bold>) Ghost Module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-2.tif"/>
</fig>
<p>However, while cheap operations can improve computational efficiency, its processing of basic feature maps is not detailed enough. In the process of generating ghost feature maps, all basic feature maps are not distinguished according to their importance or relevance. This means that some basic feature maps that may contain key information do not receive more detailed processing, while some less important ones may receive excessive attention. This non-selective processing method may affect the quality of feature extraction and potentially negatively impact the overall model performance. Based on these considerations, this paper improves the GhostNet design to optimize the generation process of ghost feature maps.</p>
<p>It should be noted that in recent years, some lightweight models based on Transformers, such as MobileViT [<xref ref-type="bibr" rid="ref-20">20</xref>], MobileFormer [<xref ref-type="bibr" rid="ref-21">21</xref>], and EfficientFormer [<xref ref-type="bibr" rid="ref-22">22</xref>], have achieved excellent performance in image classification tasks by combining the local perception ability of CNNs with the global modeling advantages of Transformers. These models typically adopt a &#x201C;hybrid&#x201D; architecture, using convolutions in the shallow layers to extract local features and introducing lightweight attention mechanisms in the deeper layers to capture long-range dependencies.</p>
<p>Although the above methods have certain advantages in terms of accuracy, they are limited by the higher computational overhead brought by the architecture they rely on. Despite lightweight improvements, they still limit the deployment efficiency on extremely low-power edge devices. In contrast, pure convolutional architectures such as GhostNet, although having a limited receptive field, have higher computational density and better hardware compatibility, and are especially suitable for delay-sensitive application scenarios.</p>
<p>Therefore, this paper chooses to improve the convolutional architecture rather than turning to the Transformer paradigm. We propose to optimize the generation process of &#x201C;ghost feature maps&#x201D; in GhostNet by introducing residual connections and adaptive channel attention mechanisms, thereby improving the quality and selectivity of feature expression without significantly increasing the number of parameters. This design concept aims to balance model accuracy, computational efficiency, and deployment feasibility, and is especially suitable for resource-constrained edge computing environments.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>SE Block</title>
<p>Hu et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed Squeeze-and-Excitation Networks (SENet), an architecture that enhances the representation capabilities of neural networks by modeling the interdependencies between feature channels. The core of this architecture is the SE block, which can be embedded into existing CNNs to improve performance.</p>
<p>In the SE block, S represents the Squeeze operation and E represents the Excitation operation. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the Squeeze operation uses global average pooling (GAP) to compress the spatial dimension of each channel into a single value. Specifically, for an input feature map with a shape of <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>, the Squeeze operation generates a tensor <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>z</mml:mi></mml:math></inline-formula> with a shape of <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, where each element represents the global average value of the corresponding channel. In the Excitation operation, first, a fully-connected layer is used to reduce the number of channels from <italic>C</italic> to <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>C</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>r</mml:mi></mml:math></inline-formula> for dimensionality reduction, and then another fully-connected layer is used to increase the number of channels back to <italic>C</italic>, generating a weight tensor <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:mover><mml:mi>z</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> with the weight values for each channel. After the Squeeze-Excitation processing, the weight tensor <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mover><mml:mi>z</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is multiplied element-by-element with the original input feature map <italic>X</italic> to obtain the recalibrated feature map <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, thereby enhancing useful channels and suppressing less important ones.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Principle of how the squeeze-and-excitation block processes the input tensor</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-3.tif"/>
</fig>
<p>The design of the SE block enables the model to effectively re-calibrate the channel feature responses, thus enhancing the model&#x2019;s representational ability. It can be well applied before the cheap operations of GhostNet to distinguish the importance of basic feature maps. However, the use of global average pooling in the Squeeze operation, while capable of compressing spatial information to capture global information of each channel, also has the limitation of losing local detail information. Therefore, it can also be improved for application.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>GPU Memory Allocation</title>
<p>Literature [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed that using direct metrics is more important than indirect metrics when designing efficient convolutional neural network architectures. For example, examining the model&#x2019;s inference speed is more direct than comparing FLOPs to reflect the model&#x2019;s computational speed. Therefore, the author Ma et al. proposed four principles for GPU use in the field of deep learning in the literature. On this basis, experiments have found that the performance of deep learning models is also closely related to GPU memory allocation. During the training and inference processes, the different proportions of allocated memory and cached memory on the GPU are crucial for whether the model can achieve better performance.</p>
<p>Allocated memory refers to the actual physical memory directly allocated for a model. When a model is loaded onto a GPU for training or inference, the model&#x2019;s parameters, activation values, gradients, etc. need to be stored in the GPU memory, and this part of the memory is reserved specifically for this model and cannot be used by other processes [<xref ref-type="bibr" rid="ref-23">23</xref>]. Therefore, when designing lightweight models, reducing the number of parameters is an effective means reducing the occupied allocated memory.</p>
<p>Cached memory refers to the additional memory cached on the GPU in addition to the memory directly allocated for the current task. This memory usually contains temporary data generated during the model training or inference process. This data can be reused to speed up the calculation without having to read data again from the main memory or other slower storage media. Therefore, it is mainly used to predict and preload data that may be needed, and retain frequently accessed data in the cache to effectively reduce waiting time and improve the overall computational efficiency [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>].</p>
<p>From the operating mechanisms of allocated memory and cached memory, it can be seen that model parameters mainly occupy allocated memory, while data generated during forward or backward propagation mainly occupie cached memory, and the cached memory can improve computational efficiency. Therefore, this paper hypothesizes that by reducing the overall number of model parameters, the allocated memory can be decreased, and the SE blocks, originally residing in the cache memory, can be relocated before the cheap operations. This aims to achieve higher accuracy without any slowdown&#x2014;or even faster&#x2014;inference speed.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Method</title>
<sec id="s3_1">
<label>3.1</label>
<title>Adaptive-SE Block</title>
<p>Although the SE block can basically distinguish the importance of different channels, its method of learning channel importance weights through global average pooling and simple fully-connected layers may not be able to fully capture the dynamic relationships and relative importance of each channel within the entire input feature map. Therefore, the SE block still has certain limitations in capturing the complex interdependencies between channels. Thus, in order to further enhance the model&#x2019;s ability to understand the dependencies between channels, this paper designs a weight self-attention mechanism based on the principle of the attention mechanism [<xref ref-type="bibr" rid="ref-24">24</xref>], and its formula is:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>W</mml:mi><mml:mi>e</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>f</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>z</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:msqrt></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>z</mml:mi></mml:math></inline-formula> is the weight matrix generated after the Squeeze operation, <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msup><mml:mi>z</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula> is the transpose of <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>z</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:msqrt></mml:math></inline-formula> is the scaling factor. The role of this attention mechanism is to adjust the weight values according to the relationship between each element and the global, so as to more accurately reflect the relative importance of each channel within the entire input. Specifically, by calculating the similarity scores and normalizing these scores using the <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:math></inline-formula> function, a probability distribution representing the interrelationships between channels can be obtained. This distribution is used to re-weight the weight matrix after the Squeeze operation to emphasize more important channel features.</p>
<p>The SE block combined with the weight self-attention mechanism is named the Adaptive-SE block in this paper. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, this is the processing process of the Adaptive-SE block for the input feature map. The input feature maps first undergo the Squeeze operation to obtain their weight matrix Subsequently, through the weight self-attention mechanism, each element in the weight matrix adjusts its weight value according to its own relationship with the global, so that important channel features are strengthened. Finally, after the Excitation operation further adjusts the channel weights and uses the <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:math></inline-formula> function to generate the final channel importance weights, these weights are used to re-calibrate the input feature map, thereby outputting the optimized feature map.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Principle of how the Adaptive Squeeze-and-Excitation block processes the input tensor. Compared with the standard SE block, this paper inserts a weight-based self-attention mechanism between the Squeeze and Excitation steps</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-4.tif"/>
</fig>
<p>By introducing the weight self-attention mechanism into the SE block, it can effectively alleviate the problem of single-dimensional feature information extraction caused by only using global average pooling. This enables the model to dynamically adjust the importance weights of each channel based on the current input feature map, rather than simply relying on the fixed results of global average pooling, so as to better capture the key information in the input data.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Resghost Module</title>
<p>In the design of lightweight neural networks, model parameters (which affect allocated memory) and the memory occupation of temporary feature maps during inference (i.e., cached memory) are two key but often confused resource dimensions. Traditional optimizations often focus on reducing FLOPs or the number of parameters, yet overlook the utilization efficiency of GPU cached memory during actual inference. The Resghost Module proposed in this paper was designed from a systemic perspective from the very beginning, aiming to efficiently utilize the feature maps in cached memory for dynamic computation without significantly increasing the number of parameters, thereby achieving a dual improvement in accuracy and speed.</p>
<p>Specifically, although the Adaptive-SE block introduced in the Resghost Module increases the amount of computation, it does not introduce additional learnable parameters and thus exerts minimal pressure on the model&#x2019;s memory allocation. On the contrary, these computations fully leverage the existing feature data in the GPU&#x2019;s high-speed cache, realizing an efficient strategy of &#x201C;trading computation for quality&#x201D;&#x2014;that is, enhancing the feature expression capability through lightweight attention mechanisms without increasing the memory burden during deployment. This design philosophy enables ResghostNet to maintain high-throughput, low-latency inference performance on edge devices.</p>
<p>Although the Adaptive-SE block can effectively identify important feature maps through channel attention mechanisms, solving the problem of how to determine the importance of feature map channels, in the process of generating ghost feature maps, relying only on simple linear operations is likely to lead to information loss and weakening of feature expression ability. This is especially true in deep-network structures, where it may cause gradient vanishing or network degradation. Inspired by the role of residual connections in stabilizing the training process of deep neural networks, this paper fuses the original feature map and the ghost feature map across layers to construct the Resghost Module with an identity mapping path.</p>
<p>As shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, the Resghost Module adopts a dual-branch structure to achieve feature reuse: the first branch uses a standard SE block to re-weight the feature maps, strengthening the semantic information of important channels; the second branch generates ghost feature maps through the Adaptive-SE block and cheap operations.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Internal structure of the Resghost Module, comprising a main branch and a residual connection</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-5.tif"/>
</fig>
<p>Specifically, in the design of the Resghost Module, the original input first passes through a primary convolution layer, which is mainly used to generate basic feature maps in preparation for subsequent operations. Let the input feature map be <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, and after passing through the primary convolution layer, it is transformed into <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup></mml:math></inline-formula>. Since both branches process <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup></mml:math></inline-formula>, the Squeeze operation can be performed on <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>z</mml:mi></mml:math></inline-formula> before the branches to obtain the weight matrix <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup></mml:math></inline-formula>, thereby reducing computational complexity.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>H</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In the cheap operation branch, the weight matrix <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>z</mml:mi></mml:math></inline-formula> first completes the processing steps of weight self-attention and the Excitation operation in the Adaptive-SE block to obtain the feature map <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mi>X</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:math></inline-formula>.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>W</mml:mi><mml:mi>e</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>f</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>z</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:msqrt><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>X</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mi>E</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mi>E</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> refers to the ReLU function, and <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>r</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are the weight matrices for dimensionality reduction and dimensionality increase, respectively. After the important channels are weighted by the Adaptive-SE block, a linear operation is performed on the feature map <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msubsup><mml:mi>X</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:math></inline-formula> to generate the ghost feature map <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:math></inline-formula>.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:math></inline-formula> represents the cheap operation, which is used to generate ghost feature maps.</p>
<p>In the residual connection path, the previously obtained weight matrix <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>z</mml:mi></mml:math></inline-formula> is processed by the Excitation operation to obtain <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mi>X</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:math></inline-formula>.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>X</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mi>E</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mi>E</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mi>z</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Finally, the outputs <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mi>X</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:math></inline-formula> of the two branches are concatenated in the channel dimension to obtain the output <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>The dual path design of the Resghost Module constructs an identity mapping path, ensuring the integrity of the original features and optimizing the stability of gradient propagation, effectively alleviating the issue of gradient diffusion in deep networks. The branch for generating ghost feature maps, combined with the Adaptive-SE block, enhances the semantic feature expression ability in cheap linear operations.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>ResghostNet</title>
<p>Based on the innovative design of the Resghost Module, taking an input image of size 224 <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224 as an example, we propose the ResghostNet architecture as shown in <xref ref-type="table" rid="table-1">Table 1</xref>. This model inherits the lightweight concept of GhostNet and uses the basic structure of MobileNetV3 as a reference framework. However, it reconstructs the core bottleneck block, the ghost bottleneck, into a Resghost bottleneck containing the Resghost Module. The initial layer of the network is a standard 3 <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 3 convolution layer with 16 filters. Subsequently, it consists of three stages grouped according to the resolution of the feature maps, and each stage contains several Resghost bottlenecks. The stage division and layer configuration are as follows: The first stage contains one Ghost bottleneck layer, using a 3 <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 3 depthwise convolution and maintaining input and output channels at 16; the second stage includes two Resghost bottleneck layers, where the first layer performs downsampling with a stride of 2 using a 3 <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 3 depthwise convolution, reducing the channels from 48 to 24, and the second layer maintains the feature map size and expands the channels from 72 to 24; the third stage deploys two Resghost bottleneck layers, where the first layer uses a 5 <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 5 depthwise convolution for downsampling (stride of 2), expanding the channels from 72 to 40 and introducing an Adaptive SE module with a reduction ratio of 0.25, and the second layer maintains the resolution and expands the channels from 120 to 40, retaining the SE module in the Resghost bottleneck backbone. Downsampling in all stages is achieved through the first layer&#x2019;s convolution with a stride of 2 in each stage. This design reduces computational complexity by adjusting the resolution upfront.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Overall architecture of ResghostNet</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Input size</th>
<th>Operator</th>
<th>Output size</th>
<th>SE ratio</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msup><mml:mn>224</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 3</td>
<td>Conv2d 3 <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 3</td>
<td><inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msup><mml:mn>112</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 16</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msup><mml:mn>112</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 16</td>
<td>Ghost bottleneck</td>
<td><inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msup><mml:mn>112</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 16</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msup><mml:mn>112</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 16</td>
<td>Resghost bottleneck</td>
<td><inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msup><mml:mn>56</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 24</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msup><mml:mn>56</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 24</td>
<td>Resghost bottleneck</td>
<td><inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msup><mml:mn>56</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 24</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msup><mml:mn>56</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 24</td>
<td>Resghost bottleneck</td>
<td><inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msup><mml:mn>28</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 40</td>
<td>0.25</td>
</tr>
<tr>
<td><inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msup><mml:mn>28</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 40</td>
<td>Resghost bottleneck</td>
<td><inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msup><mml:mn>28</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 40</td>
<td>0.25</td>
</tr>
<tr>
<td><inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msup><mml:mn>28</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 40</td>
<td>Conv2d 1 <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1</td>
<td><inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msup><mml:mn>28</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msup><mml:mn>28</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280</td>
<td>AvgPool 7 <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 7</td>
<td><inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280</td>
<td>FC</td>
<td>1000</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Compared with the traditional GhostNet, the innovations of ResghostNet are reflected in three aspects: First, it introduces the Adaptive-SE block to replace the standard SE block, which enhances the semantic expression of important channels through dynamic weight allocation and mitigates information loss caused by linear operations when combined with cheap operations. Second, it achieves feature reuse through dual-branch residual connections, balancing computational efficiency and representation capability. Third, it reduces the model&#x2019;s depth and consequently the number of parameters while significantly improving performance.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Related Configurations</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Parameter Settings</title>
<p>The experimental environment of this study is based on the Ubuntu 22.04.5 LTS operating system, using the PyTorch 2.4.1 deep learning framework and accelerated by GPU through CUDA 11.1. The experimental hardware configuration includes an Intel(R) Xeon(R) Gold 6133 CPU @ 2.50 GHz and four NVIDIA RTX 4090 D GPUs (each with 24 GB of GPU memory). RGB images are employed as the input modality, and the optimizer is stochastic gradient descent (SGD). Training is conducted for 200 epochs with a step-wise learning-rate schedule: the learning rate is reduced by 10% at the 92nd and 136th epochs. <xref ref-type="table" rid="table-2">Table 2</xref> summarizes the experimental performance metrics of the ResghostNet model on three benchmark datasets.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Experimental results of ResghostNet on three datasets. Proportion refers to the proportion of cached memory on the GPU</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Top-1 Acc (%)</th>
<th>Parameters (M)</th>
<th>FLOPs (M)</th>
<th>Inference time (<inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>)</th>
<th>Proportion (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CIFAR-10</td>
<td>90.85</td>
<td>0.79</td>
<td>154</td>
<td>148</td>
<td>95.37</td>
</tr>
<tr>
<td>CIFAR-100</td>
<td>83.2</td>
<td>0.93</td>
<td>285</td>
<td>169</td>
<td>95.21</td>
</tr>
<tr>
<td>ImageNet</td>
<td>74.7</td>
<td>2.49</td>
<td>567</td>
<td>202</td>
<td>95.76</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Datasets</title>
<p>The CIFAR-10 and CIFAR-100 datasets used in this study are two classic image datasets widely applied in the field of computer vision. Both of these datasets were compiled by scholars Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton from the University of Toronto in Canada and are part of the CIFAR (Canadian Institute for Advanced Research) series of datasets.</p>
<p>The CIFAR-10 dataset contains 60,000 color images divided into 10 categories, with 6000 images in each category. Among them, 50,000 images are for training and 10,000 are for testing. Each image has a resolution of <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> pixels and uses the RGB color space. The images in CIFAR-10 are rich and diverse in content, covering a wide range of object types such as airplanes, cars, and birds. Moreover, the number of samples in each category is balanced, which helps to avoid the problem of class imbalance. Due to the low image resolution, CIFAR-10 is an ideal choice for training lightweight models and is also very suitable for rapid experiments and model development.</p>
<p>In contrast, the CIFAR-100 dataset contains 60,000 color images divided into 100 categories, with 600 images in each category. Among them, 50,000 images are for training and 10,000 are for testing. Each image has a resolution of <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> pixels and also uses the RGB color space. Different from CIFAR-10, the categories in CIFAR-100 are more fine-grained, covering a wider range of object types such as animals, vehicles, and electrical appliances. The images within each category have a high degree of similarity, which increases the difficulty of the classification task. Due to the low image resolution, CIFAR-100 is also an ideal choice for training lightweight models and is very suitable for tasks that require more precise classification.</p>
<p>The ImageNet dataset comprises more than 14 million manually annotated images spanning over 20,000 categories. Its unprecedented scale and diversity enable models to learn rich feature representations, thereby improving generalization across a wide range of visual recognition tasks. Through the annual ImageNet Large Scale Visual Recognition Challenge (ILSVRC), ImageNet has become an internationally recognized benchmark. Researchers and engineers participating in the challenge evaluate their algorithms on the same training and test splits, facilitating direct comparisons between different methods and driving rapid progress in deep learning. Owing to its broad acceptance and formidable difficulty, many pivotal technical and theoretical breakthroughs have emerged from efforts to tackle ImageNet-related problems. For instance, the success of convolutional neural networks is largely attributed to their outstanding performance in the ImageNet Challenge. Consequently, experimenting on ImageNet has become the &#x201C;gold standard&#x201D; for validating whether new ideas and technologies are truly effective.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Ablation Experiments</title>
<p>To investigate the impact of each component of the model on overall performance and to determine the optimal model structure, we conducted ablation experiments. The ablation experiments were carried out on a single computer, with each model trained and inferred on a separate GPU within the computer, using the CIFAR-10 dataset. In the ablation experiments, we explored the performance of four structural variants of the model shown in the <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. The four structures are as follows: <xref ref-type="fig" rid="fig-6">Fig. 6a</xref>&#x2014;the ResghostNet without the weight self-attention mechanism; <xref ref-type="fig" rid="fig-6">Fig. 6b</xref>&#x2014;the ResghostNet without the Adaptive-SE block; <xref ref-type="fig" rid="fig-6">Fig. 6c</xref>&#x2014;the ResghostNet without the SE block in the shortcut connection branch; <xref ref-type="fig" rid="fig-6">Fig. 6d</xref>&#x2014;the ResghostNet with the channel shuffle operation [<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The four model structures used in the ablation experiments. (<bold>a</bold>) the ResghostNet without the weight self-attention mechanism, (<bold>b</bold>) the ResghostNet without the Adaptive-SE block, (<bold>c</bold>) the ResghostNet without the SE block in the shortcut connection branch, (<bold>d</bold>) the ResghostNet with the channel shuffle operation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-6.tif"/>
</fig>
<p>As shown in <xref ref-type="table" rid="table-3">Table 3</xref>, after removing the Weighted Attention (WA) mechanism and the Adaptive SE block (ASE), the model&#x2019;s Top-1 accuracy dropped to 90.37% and 89.90%, respectively, which is a significant decrease compared to the original ResghostNet&#x2019;s accuracy. This indicates that both the Weighted Attention mechanism and the Adaptive SE block play a crucial role in enhancing the model&#x2019;s expressive power and feature extraction capabilities. Moreover, when the Shortcut connection branch did not integrate the SE block, the model&#x2019;s Top-1 accuracy further decreased to 87.90%. This suggests that the SE block not only improves the quality of feature maps but also effectively promotes the propagation of information through deeper layers of the network. In the model variant that introduced Channel Shuffle, although the operation was designed to optimize the interaction between feature maps and enhance the model&#x2019;s expressive power, the Top-1 accuracy of this variant was only 89.97% in the actual experiment, lower than other variants. This result indicates that Channel Shuffle may not be suitable for the current network architecture, or it needs to be combined with other optimization strategies to realize its potential advantages. From the perspective of inference efficiency, although ResghostNet&#x2019;s floating-point operations (FLOPs) increased compared to GhostNet, the model&#x2019;s inference speed was not significantly affected due to the substantial reduction in the number of parameters. This proves that our assumptions about computational resource allocation in the design phase are valid, that is, by optimizing the use of cached memory on the GPU, higher model accuracy can be achieved without compromising inference efficiency.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Results of ablation experiments. Inference time is measured as the per-image processing duration. Proportion refers to the proportion of cached memory on the GPU</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Top-1 Acc (%)</th>
<th>Parameters</th>
<th>FLOPs (M)</th>
<th>Inference time (<inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>)</th>
<th>Proportion (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>ResghostNet</td>
<td>90.85</td>
<td>789,518</td>
<td>154</td>
<td>148</td>
<td>95.37</td>
</tr>
<tr>
<td>ResghostNet-WA</td>
<td>90.37</td>
<td>789,518</td>
<td>148</td>
<td>172</td>
<td>94.23</td>
</tr>
<tr>
<td>ResghostNet-ASE</td>
<td>89.90</td>
<td>789,518</td>
<td>146</td>
<td>163</td>
<td>93.84</td>
</tr>
<tr>
<td>ResghostNet-SE</td>
<td>87.90</td>
<td>789,518</td>
<td>151</td>
<td>168</td>
<td>94.13</td>
</tr>
<tr>
<td>ResghostNet&#x002B;CS</td>
<td>89.97</td>
<td>789,518</td>
<td>154</td>
<td>167</td>
<td>95.43</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experiments on the CIFAR-10 Dataset</title>
<p>We first conducted image classification experiments on the CIFAR-10 dataset. The experimental results in <xref ref-type="table" rid="table-4">Table 4</xref> fully demonstrate the effectiveness of the proposed ResghostNet and its variant ResghostNet Small in lightweight model design. Overall, by introducing residual connections, SE blocks, and Adaptive-SE blocks, ResghostNet not only significantly improved the model&#x2019;s classification performance but also achieved optimizations in terms of model parameter count and memory usage.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Experimental results of classic lightweight models on the CIFAR-10 dataset. Proportion refers to the proportion of cached memory on the GPU</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Top-1 Acc (%)</th>
<th>Parameters</th>
<th>Inference time (<inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>)</th>
<th>Proportion</th>
</tr>
</thead>
<tbody>
<tr>
<td>ResghostNet(ours)</td>
<td>90.85</td>
<td>789,518</td>
<td>148</td>
<td>95.37</td>
</tr>
<tr>
<td>ResghostNet Small(ours)</td>
<td>90.05</td>
<td>402,028</td>
<td>140</td>
<td>95.26</td>
</tr>
<tr>
<td>GhostNet [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>83.44</td>
<td>3,912,650</td>
<td>193</td>
<td>85.98</td>
</tr>
<tr>
<td>SqueezeNet [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>88.22</td>
<td>746,292</td>
<td>134</td>
<td>97.15</td>
</tr>
<tr>
<td>MobileNetV2<inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula>0.5 [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>82.35</td>
<td>700,778</td>
<td>138</td>
<td>91.25</td>
</tr>
<tr>
<td>MobileNetV2 [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>86.18</td>
<td>2,237,770</td>
<td>147</td>
<td>92.99</td>
</tr>
<tr>
<td>MobileNetV2<inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula>1.5 [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>88.80</td>
<td>4,958,762</td>
<td>156</td>
<td>92.78</td>
</tr>
<tr>
<td>MobileNetV3 Small [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>80.46</td>
<td>1,528,106</td>
<td>171</td>
<td>85.80</td>
</tr>
<tr>
<td>MobileNetV3 Large [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>85.39</td>
<td>4,214,842</td>
<td>176</td>
<td>87.67</td>
</tr>
<tr>
<td>ShuffleNetV2<inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula>0.5 [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>77.35</td>
<td>352,042</td>
<td>169</td>
<td>79.10</td>
</tr>
<tr>
<td>ShuffleNetV2 [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>81.24</td>
<td>1,258,576</td>
<td>170</td>
<td>85.52</td>
</tr>
<tr>
<td>ShuffleNetV2<inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula>1.5 [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>83.05</td>
<td>2,488,874</td>
<td>176</td>
<td>84.09</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In terms of classification accuracy, ResghostNet achieved a Top-1 accuracy of 90.85% and a Top-5 accuracy of 99.61%, significantly outperforming other lightweight models. Although ResghostNet Small had a slightly lower Top-1 accuracy of 90.05% compared to ResghostNet, it still significantly surpassed other classic models. This result indicates that even with a substantial reduction in model parameters, ResghostNet Small can still maintain high classification performance.</p>
<p>In terms of model parameter count, ResghostNet Small demonstrates a significant advantage. It has only 402,028 parameters, much lower than ResghostNet and other classic lightweight models. This substantial reduction in parameters is attributed to the optimized design of the feature extraction layers in ResghostNet&#x2014;by introducing residual connections and Adaptive-SE blocks, ResghostNet effectively reduced the number of feature extraction layers while improving the quality of the feature maps. This not only decreased the model&#x2019;s storage requirements but also provided greater flexibility for deploying on resource-constrained devices.</p>
<p>Regarding inference speed, ResghostNet and ResghostNet Small both demonstrated a good balance. Although ResghostNet has an inference speed of <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mn>148</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> seconds per image, and ResghostNet Small has an inference speed of <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mn>140</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> seconds per image, both models exhibit fast inference speeds. Compared to other models, the ResghostNet series shows how to achieve higher accuracy and more efficient resource utilization without sacrificing too much inference speed.</p>
<p>In terms of memory usage, ResghostNet and ResghostNet Small also demonstrated excellent performance. The proportion of GPU cached memory used by ResghostNet was 95.37%, and that of ResghostNet Small was 95.26%. This indicates that ResghostNet can effectively utilize the GPU memory mechanism by increasing the proportion of cached memory, thereby accelerating model inference speed.</p>
<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> shows the performance of various classic models on the CIFAR-10 dataset. The <italic>x</italic>-axis denotes inference time per image, the <italic>y</italic>-axis denotes accuracy, and the bubble size reflects the model&#x2019;s memory footprint. Bubbles of the same color belong to the same model family; the scaling factor is <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mn>1.5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Performance of various classic models on the CIFAR-10 dataset. The <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis denotes inference time per image, the <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis denotes accuracy, and the bubble size reflects the model&#x2019;s memory footprint. Bubbles of the same color belong to the same model family; the scaling factor is <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mn>1.5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-7.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experiments on the CIFAR-100 Dataset</title>
<p>We validated the effectiveness of ResghostNet and ResghostNet Small on the CIFAR-100 dataset. To ensure the rigor of our conclusions, we conducted additional experiments on the CIFAR-100 dataset. The results in <xref ref-type="table" rid="table-5">Table 5</xref> further demonstrate the effectiveness of ResghostNet and ResghostNet Small in the design of lightweight models.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Experimental results of classic lightweight models on the CIFAR-100 dataset. Proportion refers to the proportion of cached memory on the GPU</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Top-1 Acc (%)</th>
<th>Parameters</th>
<th>Inference time (<inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>)</th>
<th>Proportion (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>ResghostNet (ours)</td>
<td>83.2</td>
<td>927,802</td>
<td>169</td>
<td>95.21</td>
</tr>
<tr>
<td>ResghostNet Small (ours)</td>
<td>80.9</td>
<td>593,967</td>
<td>148</td>
<td>95.17</td>
</tr>
<tr>
<td>GhostNet [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>76.2</td>
<td>4,027,940</td>
<td>206</td>
<td>85.90</td>
</tr>
<tr>
<td>SqueezeNet [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>64.0</td>
<td>746,292</td>
<td>137</td>
<td>97.15</td>
</tr>
<tr>
<td>MobileNetV2 <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 0.5 [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>70.1</td>
<td>816,068</td>
<td>151</td>
<td>90.95</td>
</tr>
<tr>
<td>MobileNetV2 [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>77.7</td>
<td>2,353,060</td>
<td>151</td>
<td>92.83</td>
</tr>
<tr>
<td>MobileNetV2 <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1.5 [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>79.3</td>
<td>5,131,652</td>
<td>155</td>
<td>92.61</td>
</tr>
<tr>
<td>MobileNetV3 Small [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>72.0</td>
<td>1,620,356</td>
<td>159</td>
<td>85.49</td>
</tr>
<tr>
<td>MobileNetV3 Large [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>82.0</td>
<td>4,330,132</td>
<td>166</td>
<td>87.47</td>
</tr>
<tr>
<td>ShuffleNetV2 <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 0.5 [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>68.1</td>
<td>444,292</td>
<td>152</td>
<td>79.28</td>
</tr>
<tr>
<td>ShuffleNetV2 [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>70.0</td>
<td>1,350,826</td>
<td>159</td>
<td>81.53</td>
</tr>
<tr>
<td>ShuffleNetV2 <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1.5 [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>74.1</td>
<td>2,581,124</td>
<td>160</td>
<td>83.82</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In terms of classification accuracy, ResghostNet achieved a Top-1 accuracy of 83.2%, while ResghostNet Small had a Top-1 accuracy of 80.9%. Although ResghostNet Small had a slightly lower Top-1 accuracy compared to ResghostNet, both models significantly outperformed other lightweight models.</p>
<p>For example, GhostNet achieved a Top-1 accuracy of 76.2%, much lower than the ResghostNet series; MobileNetV3 Small achieved a Top-1 accuracy of 72.0%, which is also significantly lower than the ResghostNet series; ShuffleNetV2 <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1.5 achieved a Top-1 accuracy of 74.1%, which is also lower than the ResghostNet series. These results indicate that the ResghostNet can better preserve the original input information and enhance feature expression capabilities, thereby significantly improving the model&#x2019;s classification performance.</p>
<p>In terms of the number of model parameters, ResghostNet Small demonstrates significant advantages. Its number of parameters is only 593,967, which is far lower than that of ResghostNet and other classic lightweight models. For instance, MobileNetV2 <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 0.5 has 816,068 parameters, and ShuffleNetV2 <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1.5 has 2,581,124 parameters. This substantial reduction in the number of parameters is attributed to the optimized design of the feature extraction layers in ResghostNet. By introducing the residual structure and the Adaptive-SE block, ResghostNet effectively reduces the number of feature extraction layers while improving the quality of feature maps. This not only reduces the storage requirements of the model but also provides greater flexibility for deploying the model on resource-constrained devices.</p>
<p>In terms of inference speed, through adaptive adjust and optimization of the model structure, ResghostNet and ResghostNet Small have demonstrated significant performance advantages. Specifically, ResghostNet achieved an inference speed of <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mn>169</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> s per image, while ResghostNet Small further improved this to <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mn>148</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> s per image. This result indicates that the ResghostNet series outperforms most of the compared classic models in terms of inference speed. Notably, SqueezeNet, the fastest model, achieves an inference time of <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mn>134</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> s, yet its Top-1 accuracy is only 64.0%&#x2014;substantially lower than the 83.2% of ResghostNet Small and the 80.9% of ResghostNet Small. This indicates that, while SqueezeNet enjoys a speed advantage, its classification performance cannot rival that of the ResghostNet family.</p>
<p><xref ref-type="fig" rid="fig-8">Fig. 8</xref> shows the performance of various classic models on the CIFAR-100 dataset. The <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis represents inference time per image, the <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis denotes accuracy, and the bubble size indicates the model&#x2019;s memory footprint. Bubbles of the same color belong to the same model family; the scaling factor is <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Performance of various classic models on the CIFAR-100 dataset. The <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis represents inference time per image, the <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis denotes accuracy, and the bubble size indicates the model&#x2019;s memory footprint. Bubbles of the same color belong to the same model family; the scaling factor is <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Experiments on ImageNet Dataset</title>
<p>Finally, we conduct experiments on ImageNet, comparing ResghostNet and ResghostNet Small against the GhostNet family, MobileNetV2 family, MobileNetV3_Large family, and MobileNetV3_Small family&#x2014;representative lightweight models. The comparison results shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref> further demonstrate the effectiveness of ResghostNet and ResghostNet Small in lightweight model design.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Performance comparison of representative models on ImageNet. The <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis indicates the number of parameters, and the <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis denotes Top-1 accuracy. Each line depicts how the accuracy of a lightweight model series evolves as its parameter count increases, with lines of the same color belonging to the same model family</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_70990-fig-9.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-9">Fig. 9</xref> confirms the effectiveness of ResghostNet and ResghostNet Small in lightweight model design. For GhostNet, accuracy rises gradually with increasing parameters, but the gains are modest. MobileNetV2 performs well at low parameter counts yet plateaus as the model grows. MobileNetV3_Large delivers competitive results in the mid-parameter regime, but further increases yield diminishing returns. MobileNetV3_Small starts poorly at very low parameters, yet its accuracy improves markedly once capacity rises. Across all parameter levels, the ResghostNet family maintains superior accuracy, especially at the low-end, where it significantly outperforms every alternative. ResghostNet attains high accuracy with far fewer parameters, underscoring its advantage in lightweight design. As parameters continue to grow, accuracy keeps climbing, demonstrating excellent scalability and robustness. Both ResghostNet and ResghostNet Small thus excel in resource-constrained scenarios, offering strong practical potential and clear benefits for real-world deployment.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>By combining the Ghost module with the residual structure and introducing the Adaptive-SE block, this study has successfully improved the quality of feature maps and significantly reduced the model&#x2019;s depth and the number of parameters. Nevertheless, experimental results on ImageNet show that ResghostNet surpasses GhostNet in accuracy while cutting the parameter count by 52%. These results demonstrate that the design enhancements significantly boosted the classification performance of ResghostNet.</p>
<p>It is worth noting that, despite the increase in computational complexity, by optimizing the usage strategy of GPU cache memory, the inference speed of the model is not significantly affected. This indicates that, on the premise of reasonable allocation of computing resources, ResghostNet can achieve higher model accuracy while maintaining an efficient inference speed.</p>
<p>These results indicate that the ResghostNet series of models provide an efficient and powerful solution for edge computing devices, with broad application prospects. Future research can further explore methods to optimize the balance between computational efficiency and model performance, and conduct in-depth studies on hardware limitations in different application scenarios to further enhance the resource utilization efficiency of the models.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by Science and Technology Innovation Project grant No. ZZKY20222304.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Yuang Chen and Yong Li; methodology, Yuang Chen and Yong Li; software, Yuang Chen; validation, Fang Lin and Yong Li; formal analysis, Yuang Chen; investigation, Fang Lin and Jiaze Jiang; resources, Yong Li; data curation, Shuhan Lv and Jiaze Jiang; writing&#x2014;original draft preparation, Yuang Chen; writing&#x2014;review and editing, Yong Li; visualization, Yuang Chen; supervision, Shuhan Lv; project administration, Yong Li; funding acquisition, Yong Li. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The datasets used in this study are publicly available as follows: 1. CIFAR-10 and CIFAR-100 Dataset: The CIFAR-10 dataset that supports the findings of this study is openly available at <ext-link ext-link-type="uri" xlink:href="https://www.cs.toronto.edu/~kriz/cifar.html">https://www.cs.toronto.edu/~kriz/cifar.html</ext-link> (accessed on 25 September 2025); 2. ImageNet Dataset: The ImageNet dataset that supports the findings of this study is openly available at <ext-link ext-link-type="uri" xlink:href="https://image-net.org/">https://image-net.org/</ext-link>. No restrictions apply to the availability of these datasets, and they can be accessed freely by following the provided URLs.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>W</given-names></string-name></person-group>. <article-title>OpenEI: an open framework for edge intelligence</article-title>. <source>Electr Eng Syst Sci</source>. <year>2019</year>;<volume>33</volume>:<fpage>1840</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icdcs.2019.00182</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Boumaraf</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <string-name><surname>Radi</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Abdelhafez</surname> <given-names>FO</given-names></string-name>, <string-name><surname>Behouch</surname> <given-names>A</given-names></string-name>, <string-name><surname>Awadhi</surname> <given-names>KYAl</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Optimized flare performance analysis through multi-modal machine learning and temporal standard deviation enhancements</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>34362</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3540558</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Boumaraf</surname> <given-names>S</given-names></string-name>, <string-name><surname>Radi</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Abdelhafez</surname> <given-names>FO</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <string-name><surname>Awadhi</surname> <given-names>KYAl</given-names></string-name>, <string-name><surname>Karki</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Vision-based air-flow monitoring in industrial flares system design using deep convolutional neural networks</article-title>. <source>Expert Syst Appl</source>. <year>2025</year>;<volume>272</volume>:<fpage>126733</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2025.126733</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>F</given-names></string-name></person-group>. <article-title>DualCascadeTSF-MobileNetV2: a lightweight violence behavior recognition model</article-title>. <source>Appl Sci</source>. <year>2025</year>;<volume>15</volume>(<issue>7</issue>):<fpage>3862</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app15073862</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sohail</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zahoora</surname> <given-names>U</given-names></string-name>, <string-name><surname>Qureshi</surname> <given-names>AS</given-names></string-name></person-group>. <article-title>A survey of the recent architectures of deep convolutional neural networks</article-title>. <source>Artif Intell Rev</source>. <year>2020</year>;<volume>53</volume>(<issue>8</issue>):<fpage>5455</fpage>&#x2013;<lpage>516</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-020-09825-6</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>F</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Han</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>F</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Review of lightweight deep convolutional neural networks</article-title>. <source>Arch Comput Methods Eng</source>. <year>2024</year>;<volume>31</volume>(<issue>4</issue>):<fpage>1915</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11831-023-09973-w</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chiew</surname> <given-names>TK</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Image recognition based on lightweight convolutional neural network: recent advances</article-title>. <source>Image Vis Comput</source>. <year>2024</year>;<volume>146</volume>:<fpage>105037</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.imavis.2023.105037</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>K</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>HT</given-names></string-name></person-group>. <article-title>Pruning from scratch</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2020</year>;<volume>34</volume>(<issue>7</issue>):<fpage>12273</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v34i07.6817</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Elsken</surname> <given-names>THJ</given-names></string-name>, <string-name><surname>Metzen</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Hutter</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Neural architecture search: a survey</article-title>. <source>J Mach Learn Res</source>. <year>2019</year>;<volume>20</volume>(<issue>30</issue>):<fpage>1</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Knowledge distillation from few samples</article-title>. <comment>Statistics. arXiv:1811.05047. 2018</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Howard</surname> <given-names>AG</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kalenichenko</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Weyand</surname> <given-names>T</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MobileNets: efficient convolutional neural networks for mobile vision applications</article-title>. <comment>arXiv:1704.04861. 2017</comment>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sandler</surname> <given-names>M</given-names></string-name>, <string-name><surname>Howard</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhmoginov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>LC</given-names></string-name></person-group>. <article-title>MobileNetV2: inverted residuals and linear bottlenecks</article-title>. In: <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2018 Jun 18&#x2013;22</conf-name>; <publisher-loc>Salt Lake City, UT, USA</publisher-loc>. p. <fpage>4510</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2018.00474</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Howard</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sandler</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>LC</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Searching for MobileNetV3</article-title>. In: <conf-name>2019 IEEE/CVF International Conference on Computer Vision (ICCV)</conf-name>; <year>2019 Oct 27&#x2013;31</year>; <publisher-loc>Seoul, Republic of Korea</publisher-loc>. p. <fpage>1314</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV.2019.00140</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>ShuffleNet: an extremely efficient convolutional neural network for mobile devices</article-title>. In: <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2018 Jun 18&#x2013;22</conf-name>; <publisher-loc>Salt Lake City, UT, USA</publisher-loc>. p. <fpage>6848</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2018.00716</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>HT</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>ShuffleNet V2: practical guidelines for efficient CNN architecture design</article-title>. In: <conf-name>Proceedings of the European Conference on Computer Vision (ECCV); 2018 Sep 8&#x2013;14</conf-name>; <publisher-loc>Munich, Germany</publisher-loc>. p. <fpage>116</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-01228-1_7</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Albanie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Squeeze-and-excitation networks</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2020</year>;<volume>42</volume>(<issue>8</issue>):<fpage>2011</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2019.2913372</pub-id>; <pub-id pub-id-type="pmid">31034408</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>GhostNet: more features from cheap operations</article-title>. In: <conf-name>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <year>2020 Jun 13&#x2013;19</year>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>1581</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00166</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shafiq</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition: a survey</article-title>. <source>Appl Sci</source>. <year>2022</year>;<volume>12</volume>(<issue>8972</issue>):<fpage>8972</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app12178972</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>AntMan: dynamic scaling on GPU clusters for deep learning</article-title>. In: <conf-name>Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation (OSDI&#x2019;20); 2020 Jul 13&#x2013;15</conf-name>; <publisher-loc>Virtual</publisher-loc>. p. <fpage>357</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1109/msp.2010.134</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Mehta</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rastegari</surname> <given-names>M</given-names></string-name></person-group>. <article-title>MobileViT: Light-weight, general-purpose, and mobile-friendly vision transformer</article-title>. <comment>arXiv:2201.00986. 2022</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>L</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Mobile-former: bridging mobilenet and transformer</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18&#x2013;24</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>13396</fpage>&#x2013;<lpage>405</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>E</given-names></string-name>, <string-name><surname>Evangelidis</surname> <given-names>G</given-names></string-name>, <string-name><surname>Tulyakov</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>EfficientFormer: vision transformers at MobileNet speed</article-title>. <comment>arXiv:2206.01191. 2022</comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name></person-group>. <article-title>SwapAdvisor: pushing deep learning beyond the GPU memory limit via smart swapping</article-title>. In: <conf-name>Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS&#x2019;20); 2020 Mar 9&#x2013;13</conf-name>; <publisher-loc>Lausanne, Switzerland</publisher-loc>. p. <fpage>675</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3377555.3377901</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Attention is all you need</article-title>. In: <conf-name>Advances in Neural Information Processing Systems (NeurIPS); 2017 Dec 4&#x2013;9</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>