<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">52097</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.052097</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An Enhanced GAN for Image Generation</article-title>
<alt-title alt-title-type="left-running-head">An Enhanced GAN for Image Generation</alt-title>
<alt-title alt-title-type="right-running-head">An Enhanced GAN for Image Generation</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Tian</surname><given-names>Chunwei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="aff" rid="aff-3">3</xref><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Gao</surname><given-names>Haoyang</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Pengwei</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Zhang</surname><given-names>Bob</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>bobzhang@um.edu.mo</email></contrib>
<aff id="aff-1"><label>1</label><institution>PAMI Research Group, University of Macau</institution>, <addr-line>Macau, 999078</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Software, Northwestern Polytechnical University</institution>, <addr-line>Xi&#x2019;an, 710129</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>Yangtze River Delta Research Institute, Northwestern Polytechnical University</institution>, <addr-line>Taicang, 215400</addr-line>, <country>China</country></aff>
<aff id="aff-4"><label>4</label><institution>Research &#x0026; Development Institute, Northwestern Polytechnical University</institution>, <addr-line>Shenzhen, 518057</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Bob Zhang. Email: <email>bobzhang@um.edu.mo</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>18</day>
<month>7</month>
<year>2024</year></pub-date>
<volume>80</volume>
<issue>1</issue>
<fpage>105</fpage>
<lpage>118</lpage>
<history>
<date date-type="received">
<day>22</day>
<month>3</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>6</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Tian et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Tian et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_52097.pdf"></self-uri>
<abstract>
<p>Generative adversarial networks (GANs) with gaming abilities have been widely applied in image generation. However, gamistic generators and discriminators may reduce the robustness of the obtained GANs in image generation under varying scenes. Enhancing the relation of hierarchical information in a generation network and enlarging differences of different network architectures can facilitate more structural information to improve the generation effect for image generation. In this paper, we propose an enhanced GAN via improving a generator for image generation (EIGGAN). EIGGAN applies a spatial attention to a generator to extract salient information to enhance the truthfulness of the generated images. Taking into relation the context account, parallel residual operations are fused into a generation network to extract more structural information from the different layers. Finally, a mixed loss function in a GAN is exploited to make a tradeoff between speed and accuracy to generate more realistic images. Experimental results show that the proposed method is superior to popular methods, i.e., Wasserstein GAN with gradient penalty (WGAN-GP) in terms of many indexes, i.e., Frechet Inception Distance, Learned Perceptual Image Patch Similarity, Multi-Scale Structural Similarity Index Measure, Kernel Inception Distance, Number of Statistically-Different Bins, Inception Score and some visual images for image generation.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Generative adversarial networks</kwd>
<kwd>spatial attention</kwd>
<kwd>mixed loss</kwd>
<kwd>image generation</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Science and Technology Development Fund</funding-source>
<award-id>0028/2023/RIA1</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Gusu Innovation and Entrepreneurship</funding-source>
<award-id>ZXL2023170</award-id>
</award-group>
<award-group id="awg3">
<funding-source>TCL Science and Technology Innovation</funding-source>
<award-id>D5140240118</award-id>
</award-group>
<award-group id="awg4">
<funding-source>Guangdong Basic and Applied Basic Research Foundation</funding-source>
<award-id>2021A1515110079</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Due to the development of image vision techniques, image generation techniques have been applied in many fields, i.e., person privacy protection [<xref ref-type="bibr" rid="ref-1">1</xref>] and entertainment [<xref ref-type="bibr" rid="ref-2">2</xref>]. That is, digital devices can use generated face rather than captured unauthorized faces to address personal privacy protection questions [<xref ref-type="bibr" rid="ref-1">1</xref>]. Image generation techniques [<xref ref-type="bibr" rid="ref-2">2</xref>] can generate virtual persons rather than spokesmen in entertainment to save costs. Traditional image generation techniques can use multiple two-dimensional images to recover three-dimensional structures to achieve a transform of multiple images to one aim image [<xref ref-type="bibr" rid="ref-3">3</xref>]. For instance, three-dimensional morphable model uses principal component analysis to decrease texture and facial shape features in low-dimensional space to generate more real face images [<xref ref-type="bibr" rid="ref-3">3</xref>]. However, it requires a lot of varying illuminations, postures, and expressions, which may cause high data collection costs and low of the proposed method. GANs with generating high-quality and diverse images have become popular in image generation [<xref ref-type="bibr" rid="ref-4">4</xref>]. Designed two different encoders in an unsupervised way can learn potential distribution to address attribute entanglement, where output distribution can use adversarial strategy learning to maintain characteristic features of GAN to improve the effect of image generation [<xref ref-type="bibr" rid="ref-5">5</xref>]. Using ResNet feature pyramid as an encoder network can extract style from three features with different scales, and then a mapping network is used to extract learned styles from corresponding images in image generation [<xref ref-type="bibr" rid="ref-6">6</xref>]. Alternatively, an iterative feedback mechanism is used to improve the generation quality of face images to keep a balance between image fidelity and editing ability. To improve the quality of image generation, a residual learning is used in a transfer process to improve the iterative feedback mechanism [<xref ref-type="bibr" rid="ref-7">7</xref>]. Using compute similarity between potential vectors and images to design an adaptive similarity encoder to generate high-fidelity images, which can use existing encoders into different GANs to generate images [<xref ref-type="bibr" rid="ref-8">8</xref>]. To improve performance of image generation without increasing computational costs, multi-layer losses of ID and facial analysis are referred to generate more detailed information for image generation [<xref ref-type="bibr" rid="ref-9">9</xref>]. Alternatively, Xu et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] used a novel hierarchical encoder to extract hierarchical features via input images to improve the effects of generated images. Although mentioned gametic GANs may improve the effects of generated images, they may suffer from challenges from varying scenes.</p>
<p>In this paper, we present an enhanced GAN via improving a generator for image generation termed as EIGGAN. EIGGAN utilizes a spatial attention mechanism to improve the generator in order to extract salient information that can enhance the correctness of the predicted images. To improve the generation effects, parallel residual operations are gathered into a generation network to extract more structural information from the different layers in terms of their relation to context. To make a tradeoff between speed and accuracy in image generation, a mixed loss function is used in a GAN to generate more realistic images. Experiments illustrate that the proposed EIGGAN is competitive in terms of the metrics: Frechet Inception Distance (FID) [<xref ref-type="bibr" rid="ref-11">11</xref>], Learned Perceptual Image Patch Similarity (LPIPS) [<xref ref-type="bibr" rid="ref-12">12</xref>], Multi-Scale Structural Similarity Index Measure (MS-SSIM) [<xref ref-type="bibr" rid="ref-13">13</xref>], Kernel Inception Distance (KID) [<xref ref-type="bibr" rid="ref-14">14</xref>], Number of statistically-Different Bins (NDB) [<xref ref-type="bibr" rid="ref-15">15</xref>] and Inception Score (IS) [<xref ref-type="bibr" rid="ref-16">16</xref>] for image generation.</p>
<p>The contributions of this paper can be summarized as follows:
<list list-type="order">
<list-item><p>A spatial attention mechanism is employed to improve the generation to facilitate more salient information, enhancing the truthfulness of the generated images.</p></list-item>
<list-item><p>Parallel residual operations are used to extract more complementary and structural information with a spatial attention mechanism to improve the effects of image generation, according to its relation to context.</p></list-item>
<list-item><p>A mixed loss is applied to generate more realistic images.</p></list-item></list></p>
<p>The remaining parts of this paper are conducted as follows. <xref ref-type="sec" rid="s2">Section 2</xref> presents related work. <xref ref-type="sec" rid="s3">Section 3</xref> gives the proposed method. <xref ref-type="sec" rid="s4">Section 4</xref> provides experimental analysis and results. <xref ref-type="sec" rid="s5">Section 5</xref> summarizes the whole paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>GANs with strong generative abilities are used for image generation in generation [<xref ref-type="bibr" rid="ref-17">17</xref>]. To improve generation effects, designing novel network architectures can improve GANs as the popular way of image generation [<xref ref-type="bibr" rid="ref-17">17</xref>]. It can be summarized into two kinds: improved common GANs and StyleGANs. The first method usually improves a generator or discriminator to improve the learning abilities of GANs for image generation. For improving a generator, asymmetric network architectures between a generator and discriminator are designed to address image content differences between input and outputs to promote the effects and performance of image-to-image translation [<xref ref-type="bibr" rid="ref-17">17</xref>]. To address native relation effect of semantic and latent space, a combination of a pretrain GAN and prior knowledge is used to extract latent semantic from architecture attributes in shallow layers and in apparent deep layers to improve the quality of generation images [<xref ref-type="bibr" rid="ref-18">18</xref>]. Alternatively, zooming areas of training data in image representation space in a discriminator can easily train a GAN to generate images [<xref ref-type="bibr" rid="ref-19">19</xref>]. To address dense visual alignment questions, a spatial Transformer is used to map random samples to a joint object mode for image generation [<xref ref-type="bibr" rid="ref-20">20</xref>]. To edit more attributions of generated images, the second method based on StyleGANs is proposed.</p>
<p>That is, a StyleGAN compares a learnable intermediate latent space W and standard Gaussian latent space to the reflected distribution of training data and they can effectively code rich semantic information to change the attribution of generated images, i.e., expressions and illuminations in image generation [<xref ref-type="bibr" rid="ref-21">21</xref>]. To better address image generation tasks, the second method uses fine-tuning styleGANs to enhance the quality of generation images [<xref ref-type="bibr" rid="ref-22">22</xref>]. To address inverse mapping questions of high-quality reconstruction, editability and fast reference, hypernetworks are proposed [<xref ref-type="bibr" rid="ref-16">16</xref>]. A two-phase mechanism was conducted as follows. The first phase was used to train an encoder to map an input image to a latent space. The second phase utilized a hypernetwork to recover lost information from the first phase to image editing [<xref ref-type="bibr" rid="ref-22">22</xref>]. A hypernetwork can be used to tune the weights of StyleGAN to better express given images in editable areas from the latent space for image editing [<xref ref-type="bibr" rid="ref-23">23</xref>]. To address domain transfer of image generation, a simple feature match loss is gathered into a StypleGAN to improve generation quality with less computational costs [<xref ref-type="bibr" rid="ref-24">24</xref>]. To avoid the collapse of latent variants in StyleGAN, a class embedding enhancement mechanism is referred to a self-supervised learning based on a latent space to reduce the relation of latent variants to improve the results of image generation [<xref ref-type="bibr" rid="ref-25">25</xref>]. Although these methods can improve the effects of image generation, StyleGANs refer to more training time and higher computational costs. To better generate images, a Progressive Growing of Generative Adversarial Networks (PGGAN) uses a progressively large model to reduce training difficulty to generate smoother and continuous images [<xref ref-type="bibr" rid="ref-26">26</xref>]. Inspired by that, we use a PGGAN to improve a GAN for image generation in this paper.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposal Method</title>
<sec id="s3_1">
<label>3.1</label>
<title>Network Architecture</title>
<p>To better generate high-quality images, we designed an enhanced GAN to improve a generator in image generation (EIGGAN) in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. EIGGAN mainly uses a generator and a discriminator to game to generate high-quality images. To enhance the robustness of gamistic generator and discriminator in varying scenes, an improved generator is developed. To generate more real images, three operations are affected on a generation network. The first operation is that a parallel residual learning operation is used to extract more structural information of different layers in terms of relation of context. To enhance the truthfulness of generating images, a spatial attention mechanism is gathered into this generation network to extract more salient information to improve the effect of image generation. The third operation utilizes a mixed loss function to update the parameters of a GAN to balance speed and accuracy in image generation. More detailed information on the generation network can be shown as follows.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Network architecture of EIGGAN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52097-fig-1.tif"/>
</fig>
<p>The generation network is composed of five components, i.e., Transpose ConvBlock, Residual ConvBlock, 2X Upsampling, Attention ConvBlock, Equalized Conv operation. The Transpose ConvBlock [<xref ref-type="bibr" rid="ref-26">26</xref>] is set to the first layer, which can convert a noisy vector to a matrix. 2X Upsampling is set to the second, fourth and sixth layers to enlarge obtained feature mapping to capture more context information. Enhanced Residual ConvBlock is set to the third and fifth layers to extract more structural information from different layers in terms of relation to context. An enhanced attention ConvBlock is set to the sixth layer to extract more salient information for image generation. An Equalized Conv [<xref ref-type="bibr" rid="ref-26">26</xref>] is used as the last layer to normalize obtained features to accelerate training speed. To clearly express the mentioned process, the following equation can be conducted.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>E</mml:mi><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>U</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>U</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>U</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>z</mml:mi></mml:math></inline-formula> is random noise, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>G</mml:mi></mml:math></inline-formula> is a function of a generation network, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is an output of the generation network. <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula> denotes a function of a Transpose ConvBlock, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula> denotes a function of an enhanced residual ConvBlock. <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>U</mml:mi><mml:mi>p</mml:mi></mml:math></inline-formula> denotes a 2X upsampling operation. <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>E</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula> denotes a function of an Equalized Conv. Its parameters can be updated by a mixed loss function in <xref ref-type="sec" rid="s3_2">Section 3.2</xref>.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Loss Function</title>
<p>To make a tradeoff between performance and speed, a mixed loss is conducted. That is, a mixed loss is composed of normal loss-based GAN [<xref ref-type="bibr" rid="ref-27">27</xref>], non-saturating loss [<xref ref-type="bibr" rid="ref-28">28</xref>] and a combination of R1 [<xref ref-type="bibr" rid="ref-29">29</xref>] and R2 regularization [<xref ref-type="bibr" rid="ref-29">29</xref>]. Non-saturation loss can make a generator more stable in the training process [<xref ref-type="bibr" rid="ref-28">28</xref>]. R1 regularization loss can make a discriminator easier to converge via penalizing gradients of real images to obtain more reliable images [<xref ref-type="bibr" rid="ref-28">28</xref>]. R2 regularization loss can penalty gradients of generating images to easier converge for image generation [<xref ref-type="bibr" rid="ref-29">29</xref>]. Mentioned a loss function is shown as follows.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x03B3;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>&#x03D1;</mml:mi><mml:mo>+</mml:mo><mml:mi>p</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>&#x03C8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>&#x00A0;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a loss of a discriminator. <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes probability of real images, where <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>x</mml:mi></mml:math></inline-formula> is a given real image. <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes probability of generating images, where <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>z</mml:mi></mml:math></inline-formula> is noise. <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>m</mml:mi></mml:math></inline-formula> is a sample minibatch of noise. <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>p</mml:mi><mml:mi>&#x03D1;</mml:mi></mml:math></inline-formula> is distribution of true images. <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>p</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> is distribution of a generator. <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a loss of a generator.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Enhanced Residual ConvBlock</title>
<p>Enhanced Residual ConvBlock is used to extract more structural information of different layers for image generation in terms of relation of context as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. It is composed of three Equalized Conv, two Leaky ReLU and two PixelNorm. Two combinations of Equalized Conv, Leaky ReLU and PixeNorm are stacked to extract more structural information to improve the effect of image generation. To extract more complementary information, an extra Equalized Conv, an output of the first stacked Equalized Conv, Leaky ReLU and PixelNorm and an output of two stacked Equalized Conv, Leaky ReLU and PixelNorm in a parallel way are gathered to extract more structural information, where a residual learning operation denotes a fusion way. It can be expressed as the following equation:</p>
<p><disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is an input of an enhanced residual ConvBlock. <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is output of an enhanced residual ConvBlock. <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>L</mml:mi><mml:mi>R</mml:mi></mml:math></inline-formula> is a function of Leaky ReLU. <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>P</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> is a function of PixelNorm. <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mo>+</mml:mo></mml:math></inline-formula> is a residual learning operation.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Enhanced Attention ConvBlock</title>
<p>An enhanced attention ConvBlock can extract more salient information for image generation. It is composed of four components, i.e., Equalized Conv, Spatial Attention, Leaky ReLU and PixelNorm. It is different from ERresidual Block that spatial attention [<xref ref-type="bibr" rid="ref-30">30</xref>] is set between the first Equalized Conv and Leaky ReLU to extract salient information to improve effect of image generation. This process can be shown as follows:</p>
<p><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>A</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>S</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mspace width="1em" /><mml:mo>+</mml:mo><mml:mrow><mml:mtext>PN</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>LR</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>EC</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>PN</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>LR</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>SA</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>EC</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>I</mml:mtext></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mspace width="1em" /><mml:mo>+</mml:mo><mml:mrow><mml:mtext>EC</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>I</mml:mtext></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is an input of EAttention ConvBlock. <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>A</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is an output of EResidual ConvBlock. <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mrow><mml:mtext>SA</mml:mtext></mml:mrow></mml:math></inline-formula> denotes a function of spatial attention.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Analysis and Results</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>Real image datasets are composed of 202,599 face images from CelebA [<xref ref-type="bibr" rid="ref-31">31</xref>] and 60,000 images with ten different categories, i.e., planes, cars, birds, cats, deer, dogs, frogs, horses, ships and trucks from CIFAR-10 [<xref ref-type="bibr" rid="ref-32">32</xref>]. Each image is 32 &#x00D7; 32. Generating image datasets: Each EIGGAN model can generate 10,000 images.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Settings</title>
<p>Parameters of training a EIGGAN in image generation are listed as follows. Slope parameter of LeakyReLU is 0.2. &#x03BB; of regularization parameter is 1 on CelebA and it is 0.1 on CIFAR-10. Batch size is 64. Epoch number on CelebA is 255. Epoch number on CIFAR-10 is 266. Learning rate on CelebA is 0.002, learning rate on CIFAR-10 is 0.001. &#x03B2;1 &#x003D; 0 and &#x03B2;2 &#x003D; 0.99. Also, parameters can be optimized by Adam [<xref ref-type="bibr" rid="ref-33">33</xref>]. All the codes implemented by PyTorch 1.13.1, and Python of 3.9.13 run on Ubuntu 20.04.3 with AMD EPYC 7502/3.35 GHz, Central Processing Unit (CPU) of 32 cores, 128 RAM. Also, a Graphics Processing Unit (GPU) of Nvidia GeForce GTX 3090 and Nvidia CUDA 12.2 can be used to improve training speed.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experimental Analysis</title>
<p>To improve robustness of GANs in image generation, an enhanced GAN for image generation is proposed. To improve the quality of generating images, an enhanced generation network is conducted. Deep networks rely on a deep network architecture to extract more structural information [<xref ref-type="bibr" rid="ref-34">34</xref>]. To quickly extract key information, a spatial attention mechanism is proposed [<xref ref-type="bibr" rid="ref-30">30</xref>]. It can use different dimensional channels to extract salient information [<xref ref-type="bibr" rid="ref-30">30</xref>]. Inspired by that, a spatial attention mechanism is fused into the third EResidual ConvBlock from a generation network to extract salient information to improve effects in image generation, where its information can be shown in <xref ref-type="sec" rid="s3_4">Section 3.4</xref> Its effectiveness is given in <xref ref-type="table" rid="table-1">Table 1</xref>. That is, a combination of PGGAN and Spatial attention mechanism has obtained a lower FID than that of PGGAN in <xref ref-type="table" rid="table-1">Table 1</xref>. Although deep networks can extract more accurate information to pursue better performance in vision tasks, they may ignore the importance of hierarchical information to limit their better performance and robustness [<xref ref-type="bibr" rid="ref-34">34</xref>]. To overcome the drawbacks, a parallel residual learning operations are set in the third EResidual ConvBlock as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> to extract more structural information from different layers. Its competitive results can be proved by comparing &#x2018;A combination of PGGAN, Spatial attention and Parallel residual learning operations&#x2019; and &#x2018;A combination of PGGAN and Spatial attention in <xref ref-type="table" rid="table-1">Table 1</xref>. A mixed loss function to update parameters of a GAN to balance speed and accuracy in image generation. Non-saturation loss can make a generator more stable in training process. R1 regularization loss can make a discriminator easier to converge via penalizing gradients of real images to obtain more reliable images. R2 regularization loss can penalty gradients of generating images to easier converge for image generation. Effects of R1 for EIGGAN for image generation can be tested by using &#x2018;A combination of PGGAN, Spatial attention, Parallel residual learning operations, Non-saturation and R1&#x2019; and &#x2018;A combination of PGGAN, Spatial attention and Parallel residual learning operations&#x2019; in terms of FID in <xref ref-type="table" rid="table-1">Table 1</xref>. It has an improvement of 2.76 in terms of FID in <xref ref-type="table" rid="table-1">Table 1</xref>, which shows effectiveness of R1 for image generation. Positive effects of a mixed loss can verified by comparing &#x2018;A combination of PGGAN, Spatial attention, Parallel residual learning operations, Non-saturation and R1&#x2019; and &#x2018;A combination of PGGAN, Spatial attention, Parallel residual learning operations and Mixed loss&#x2019; in <xref ref-type="table" rid="table-1">Table 1</xref>. Effectiveness of proposed key techniques can be verified by comparing &#x2018;A combination of PGGAN, Spatial attention, Parallel residual learning operations and Mixed loss&#x2019; and &#x2018;PGGAN&#x2019; in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>FID values of different methods in image generation</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>FID</th>
</tr>
</thead>
<tbody>
<tr>
<td>PGGAN [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>28.13</td>
</tr>
<tr>
<td>A combination of PGGAN and spatial attention</td>
<td>26.01</td>
</tr>
<tr>
<td>A combination of PGGAN, spatial attention and parallel residual learning operations</td>
<td>24.33</td>
</tr>
<tr>
<td>A combination of PGGAN, spatial attention, parallel residual learning operations, non-saturation and R1</td>
<td>21.57</td>
</tr>
<tr>
<td>A combination of PGGAN, spatial attention, parallel residual learning operations and Mixed loss</td>
<td>21.36</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experimental Results</title>
<p>To test the effectiveness of our method, quantitative and qualitative analysis are used to conduct. Quantitative analysis that uses a Deep Convolution GAN (DCGAN) [<xref ref-type="bibr" rid="ref-35">35</xref>], WGAN-GP [<xref ref-type="bibr" rid="ref-36">36</xref>], PGGAN [<xref ref-type="bibr" rid="ref-26">26</xref>], NCSN [<xref ref-type="bibr" rid="ref-37">37</xref>] as comparative methods on CIFAR-10 and CelebA to test performance of the proposed EIGGAN in image generation. We have chosen six indicators, i.e., Frechet Inception Distance (FID) [<xref ref-type="bibr" rid="ref-11">11</xref>], Learned Perceptual Image Patch Similarity (LPIPS) [<xref ref-type="bibr" rid="ref-12">12</xref>], Multi-Scale Structural Similarity Index Measure (MS-SSIM) [<xref ref-type="bibr" rid="ref-13">13</xref>], Kernel Inception Distance (KID) [<xref ref-type="bibr" rid="ref-14">14</xref>], Number of statistically-Different Bins (NDB) [<xref ref-type="bibr" rid="ref-15">15</xref>] and Inception Score (IS) [<xref ref-type="bibr" rid="ref-16">16</xref>] to demonstrate the superiority of our method. All formulas of FID [<xref ref-type="bibr" rid="ref-11">11</xref>], LPIPS [<xref ref-type="bibr" rid="ref-12">12</xref>], MS-SSIM [<xref ref-type="bibr" rid="ref-13">13</xref>], KID [<xref ref-type="bibr" rid="ref-14">14</xref>] and NDB [<xref ref-type="bibr" rid="ref-15">15</xref>] can be given at <ext-link ext-link-type="uri" xlink:href="https://github.com/hellloxiaotian/EIGGAN/blob/main/equation">https://github.com/hellloxiaotian/EIGGAN/blob/main/equation</ext-link> (accessed on 19/03/2024).</p>
<p>As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, we can see that the proposed EIGGAN has obtained the best results in terms of FID, LPIPS, MS-SSIM and KID on CIFAR-10. Also, it has obtained the second result in terms of NDB on CIFAR-10 in <xref ref-type="table" rid="table-2">Table 2</xref>. To verify its robustness, we conduct some extended experiments on CelebA dataset. That is, our EIGGAN has obtained the best performance in terms of FID and LPIPS on CelebA in <xref ref-type="table" rid="table-3">Table 3</xref>. Also, it has obtained the second results in terms of MS-SSIM, KID and NDB in <xref ref-type="table" rid="table-3">Table 3</xref>. According to mentioned illustrations, we can see that the proposed EIGGAN is effective for image generation.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Results of some methods on CIFAR-10 for image generation</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>FID&#x2193;</th>
<th>LPIPS&#x2191;</th>
<th>MS-SSIM&#x2193;</th>
<th>KID&#x2193;</th>
<th>NBD&#x2193;</th>
<th>IS&#x2191;</th>
</tr>
</thead>
<tbody>
<tr>
<td>DCGAN [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>35.05</td>
<td>0.207</td>
<td>0.12147</td>
<td>0.0238</td>
<td><underline>33</underline></td>
<td>6.37</td>
</tr>
<tr>
<td>WGAN-GP [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>29.30</td>
<td><underline>0.217</underline></td>
<td><underline>0.10598</underline></td>
<td>0.0145</td>
<td><bold>18</bold></td>
<td>7.86</td>
</tr>
<tr>
<td>PGGAN [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>28.13</td>
<td>0.214</td>
<td>0.11058</td>
<td>0.0133</td>
<td>68</td>
<td><bold>8.80</bold></td>
</tr>
<tr>
<td>NCSN [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td><underline>25.32</underline></td>
<td>0.195</td>
<td>0.11225</td>
<td><underline>0.0111</underline></td>
<td>81</td>
<td>8.37</td>
</tr>
<tr>
<td>EEIGGAN (ours)</td>
<td><bold>21.03</bold></td>
<td><bold>0.229</bold></td>
<td><bold>0.10501</bold></td>
<td><bold>0.0098</bold></td>
<td><underline>33</underline></td>
<td><underline>8.56</underline></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Results of some methods on CelebA for image generation</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>FID&#x2193;</th>
<th>LPIPS&#x2191;</th>
<th>MS-SSIM&#x2193;</th>
<th>KID&#x2193;</th>
<th>NBD&#x2193;</th>
</tr>
</thead>
<tbody>
<tr>
<td>DCGAN [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>32.99</td>
<td>0.230</td>
<td>0.35644</td>
<td>0.0117</td>
<td>83</td>
</tr>
<tr>
<td>WGAN-GP [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>31.17</td>
<td><underline>0.292</underline></td>
<td><bold>0.26625</bold></td>
<td>0.0073</td>
<td><bold>28</bold></td>
</tr>
<tr>
<td>PGGAN [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td><underline>12.88</underline></td>
<td>0.283</td>
<td>0.28567</td>
<td><bold>0.0010</bold></td>
<td>46</td>
</tr>
<tr>
<td>NCSN [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>42.59</td>
<td>0.118</td>
<td>0.31952</td>
<td>0.0191</td>
<td>86</td>
</tr>
<tr>
<td>EEIGGAN (Ours)</td>
<td><bold>10.62</bold></td>
<td><bold>0.309</bold></td>
<td><underline>0.28250</underline></td>
<td><underline>0.0019</underline></td>
<td><underline>45</underline></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To verify the superiority of our EIGGAN for image generation, we choose newest image generation methods, i.e., StyleGAN2 [<xref ref-type="bibr" rid="ref-38">38</xref>], TransGAN [<xref ref-type="bibr" rid="ref-39">39</xref>], ViTGAN [<xref ref-type="bibr" rid="ref-40">40</xref>], D2WMGAN [<xref ref-type="bibr" rid="ref-41">41</xref>], PFGAN [<xref ref-type="bibr" rid="ref-42">42</xref>] on CIFAR-10 in terms of Information System (IS) to conduct experiments. As mentioned earlier, the FID value is used to evaluate the quality of generated images. In addition, since the CIFAR-10 dataset contains 10 different types of object images, we also tested the IS value to evaluate the type diversity of generated images. To fairly compare the performance of our proposed method for image generation, we conduct some experiments on CIFAR-10. Due to some methods, i.e., StyleGAN2, TransGAN, ViTGAN, D2WMGAN and PFGAN didn&#x2019;t release codes, public index, i.e., FID and IS from related GANs [<xref ref-type="bibr" rid="ref-38">38</xref>&#x2013;<xref ref-type="bibr" rid="ref-42">42</xref>] can be obtained. As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, we can see that StyleGAN2 is superior to our EIGGAN. Also, our method has fewer parameters. Also, our method has obtained better effects than that of other methods for image generation in <xref ref-type="table" rid="table-5">Table 5</xref>. In summary, our EIGGAN is comparative for image generation.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparing FID and IS results of some image generation methods on CIFAR-10</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>FID&#x2193;</th>
<th>IS&#x2191;</th>
</tr>
</thead>
<tbody>
<tr>
<td>StyleGAN2 [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td><bold>8.41</bold></td>
<td><bold>9.16</bold></td>
</tr>
<tr>
<td>TransGAN [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>22.53</td>
<td>8.26</td>
</tr>
<tr>
<td>ViTGAN [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>30.72</td>
<td>8.30</td>
</tr>
<tr>
<td>D2WMGAN [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>34.94</td>
<td>7.31</td>
</tr>
<tr>
<td>PFGAN [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>47.32</td>
<td>7.97</td>
</tr>
<tr>
<td>EEIGGAN (ours)</td>
<td><underline>21.03</underline></td>
<td><underline>8.56</underline></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comparison of parameters for some methods</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>Parameters</th>
</tr>
</thead>
<tbody>
<tr>
<td>DCGAN [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>12.14 M</td>
</tr>
<tr>
<td>WGAN-GP [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>12.14 M</td>
</tr>
<tr>
<td>PGGAN [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>17.77 M</td>
</tr>
<tr>
<td>NCSN [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>30.18 M</td>
</tr>
<tr>
<td>StyleGAN2 [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>28.27 M</td>
</tr>
<tr>
<td>ViTGAN [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>45.37 M</td>
</tr>
<tr>
<td>PFGAN [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>22.19 M</td>
</tr>
<tr>
<td>EIGGAN (Ours)</td>
<td>17.90 M</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Qualitative analysis is composed of two parts: visual generation images and contrastive visual generation images. The first part contains two parts, i.e., ten objects (regarded as planes, cars, birds, cats, deer, dogs, frogs, horses, ships and trucks) and different face images. Also, each object is generated from the CIFAR-10 and each object has eight images with different directions in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. Different face images from the CelebA are generated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. These show that our EIGGAN is effective in visual images in image generation.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Generated images on the CIFAR-10 from our EIGGAN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52097-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Generated face images from our EIGGAN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52097-fig-3.tif"/>
</fig>
<p>The second part conducted some comparative visual images of popular methods and our method for image generation to test the generation ability of our EIGGAN. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, we can see that our EIGGAN is clearer than other methods. As listed in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, we can see that our EIGGAN has more detailed information than that of other methods, i.e., DCGAN, WGAN-GP, PGGAN and NCSN. For instance, DCGAN may cause distorted faces. WGGAN-GP may generate blurred faces. NCSN may generate images of poor colour and luster. According to mentioned illustrations, our method is useful for image generation in terms of qualitative and quantitative analysis.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Generated images of different methods on CIFAR-10</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52097-fig-4a.tif"/><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52097-fig-4b.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Generated images of different methods on CelebA</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_52097-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>In this paper, we improve a generator to enhance GAN for image generation. To enrich the effect of a generator, spatial attention is applied to a generation network to extract salient information, generating more information that is detailed. To enhance the relation of context, parallel residual operations are used to increase more structural information from different layers in image generation. Taking into consideration the speed and accuracy, a mixed loss function is merged into the generation network to produce images that are more realistic. Overall, the proposed method (EIGGAN) has obtained good performance in image generation. As part of our future work, we will improve the performance of image generation according to the image attributes.</p>
</sec>
</body>
<back>
<ack><p>This work was supported in part by the Science and Technology Development Fund, in part by Fundamental Research Funds for the Central Universities.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported in part by the Science and Technology Development Fund, Macao S.A.R (FDCT) 0028/2023/RIA1, in part by Leading Talents in Gusu Innovation and Entrepreneurship Grant ZXL2023170, in part by the TCL Science and Technology Innovation Fund under Grant D5140240118, in part by the Guangdong Basic and Applied Basic Research Foundation under Grant 2021A1515110079.</p>
</sec>
<sec><title>Author Contributions</title>
<p>The first author Chunwei Tian gives the main conception and writes of this paper. The second author Haoyang Gao conducts experiments, visual figures and writes part of this paper. The third author Pengwei Wang conducts part experiments. The fourth author Bob Zhang gives key comments. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The datasets used during the current study are available from the corresponding author on reasonable request.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the
present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Human action image generation with differential privacy</article-title>,&#x201D; in <conf-name>2020 IEEE Int. Conf. Multimedia and Expo (ICME)</conf-name>, <publisher-loc>London, UK</publisher-loc>, <comment>Jul. 6&#x2013;10</comment>, <year>2020</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Shan</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Chai</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Fu</surname></string-name></person-group>, &#x201C;<article-title>Virtual face image generation for illumination and pose insensitive face recognition</article-title>,&#x201D; in <conf-name>2003 Int. Conf. Multimedia and Expo. ICME&#x0027;03</conf-name>, <publisher-loc>Hong Kong, China</publisher-loc>, <publisher-name>IEEE</publisher-name>, <comment>Apr. 6&#x2013;10</comment>, <year>2003</year>, vol. <volume>3</volume>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Blanz</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Vetter</surname></string-name></person-group>, &#x201C;<article-title>Face recognition based on fitting a 3D morphable model</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>25</volume>, no. <issue>9</issue>, pp. <fpage>1063</fpage>&#x2013;<lpage>1074</lpage>, <year>Sep. 2003</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2003.1227983</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Goodfellow</surname></string-name> <etal>et al.</etal>,</person-group> &#x201C;<article-title>Generative adversarial networks</article-title>,&#x201D; <source>Commun. ACM</source>, vol. <volume>63</volume>, no. <issue>11</issue>, pp. <fpage>139</fpage>&#x2013;<lpage>144</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1145/3422622</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Pidhorskyi</surname></string-name>, <string-name><given-names>D. A.</given-names> <surname>Adjeroh</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Doretto</surname></string-name></person-group>, &#x201C;<article-title>Adversarial latent autoencoders</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, <comment>Jun. 13&#x2013;19</comment>, <year>2020</year>, pp. <fpage>14104</fpage>&#x2013;<lpage>14113</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>O.</given-names> <surname>Tov</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Alaluf</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Nitzan</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Patashnik</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Cohen-Or</surname></string-name></person-group>, &#x201C;<article-title>Designing an encoder for stylegan image manipulation</article-title>,&#x201D; in <conf-name>ACM Trans. Graph.</conf-name>, <publisher-loc>New York, NY, USA</publisher-loc>, <year>2021</year>, vol. <volume>40 no. 4</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Alaluf</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Patashnik</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Cohen-Or</surname></string-name></person-group>, &#x201C;<article-title>ReStyle: A residual-based stylegan encoder via iterative refinement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Int. Conf. Comput. Vis.</conf-name>, <publisher-loc>Montreal, QC, Canada</publisher-loc>, <comment>Oct. 10&#x2013;17</comment>, <year>2021</year>, pp. <fpage>6711</fpage>&#x2013;<lpage>6720</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Yu</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Diverse similarity encoder for deep GAN inversion</article-title>,&#x201D; <comment>arXiv preprint arXiv: 2108.10201</comment>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Wei</surname></string-name> <etal>et al.</etal>,</person-group> &#x201C;<article-title>E2Style: Improve the efficiency and effectiveness of StyleGAN inversion</article-title>,&#x201D; <comment>arXiv preprint arXiv:2104.07661</comment>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Yang</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>Generative hierarchical features from synthesizing images</article-title>,&#x201D; <comment>arXiv preprint arXiv:2007.10379</comment>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Heusel</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Ramsauer</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Unterthiner</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Nessler</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Hochreiter</surname></string-name></person-group>, &#x201C;<article-title>Gans trained by a two time-scale update rule converge to a local nash equilibrium</article-title>,&#x201D; in <conf-name>NIPS&#x0027;17: Proc. 31st Int. Conf. Neural Inf. Process. Syst.</conf-name>, <publisher-loc>Long Beach, California, USA</publisher-loc>, <comment>Dec. 4&#x2013;9</comment>, <year>2017</year>, vol. <volume>30</volume>, pp. <fpage>6629</fpage>&#x2013;<lpage>6640</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Isola</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Efros</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Shechtman</surname></string-name>, and <string-name><given-names>O.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>The unreasonable effectiveness of deep features as a perceptual metric</article-title>,&#x201D; in <conf-name>2018 IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <publisher-loc>Salt Lake City, UT, USA</publisher-loc>, <comment>Jun. 18&#x2013;23</comment>, <year>2018</year>, pp. <fpage>586</fpage>&#x2013;<lpage>595</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>E. P.</given-names> <surname>Simoncelli</surname></string-name>, and <string-name><given-names>A. C.</given-names> <surname>Bovik</surname></string-name></person-group>, &#x201C;<article-title>Multiscale structural similarity for image quality assessment</article-title>,&#x201D; in <conf-name>The Thrity-Seventh Asilomar Conf. Signals, Syst. Comput.</conf-name>, <publisher-loc>Pacific Grove, CA, USA</publisher-loc>, <publisher-name>IEEE</publisher-name>, <comment>Nov. 9&#x2013;12</comment>, <year>2003</year>, vol. <volume>2</volume>, pp. <fpage>1398</fpage>&#x2013;<lpage>1402</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Bi&#x0144;kowski</surname></string-name>, <string-name><given-names>D. J.</given-names> <surname>Sutherland</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Arbel</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Gretton</surname></string-name></person-group>, &#x201C;<article-title>Demystifying MMD GANs</article-title>,&#x201D; <comment>arXiv preprint arXiv, abs/1801.01401</comment>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Richardson</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Weiss</surname></string-name></person-group>, &#x201C;<article-title>On GANs and GMMs</article-title>,&#x201D; <comment>arXiv preprint arXiv:1805.12462</comment>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Salimans</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Goodfellow</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zaremba</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Cheung</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Radford</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Improved techniques for training GANs</article-title>,&#x201D; in <conf-name>NIPS&#x0027;16: Proc. 30th Int. Conf. Neural Inf. Process. Syst.</conf-name>, <publisher-loc>Barcelona, Spain</publisher-loc>, <comment>Dec. 5&#x2013;10</comment>, <year>2016</year>, pp. <fpage>2234</fpage>&#x2013;<lpage>2242</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>N.</given-names> <surname>Sebe</surname></string-name></person-group>, &#x201C;<article-title>Asymmetric generative adversarial networks for image-to-image translation</article-title>,&#x201D; <comment>arXiv preprint arXiv:1912.06931</comment>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Jing</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Pan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>He</surname></string-name></person-group>, &#x201C;<article-title>High-resolution face swapping via latent semantics disentanglement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <comment>Jun. 18&#x2013;24</comment>, <year>2022</year>, pp. <fpage>7642</fpage>&#x2013;<lpage>7651</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Liu</surname></string-name> <etal>et al.</etal>,</person-group> &#x201C;<article-title>Improving GAN training via feature space shrinkage</article-title>,&#x201D; in <conf-name>2023 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Vancouver, BC, Canada</publisher-loc>, <comment>Jun. 17&#x2013;24</comment>, <year>2023</year>, pp. <fpage>16219</fpage>&#x2013;<lpage>16229</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Peebles</surname></string-name>, <string-name><given-names>J. Y.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Torralba</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Efros</surname></string-name> and <string-name><given-names>E.</given-names> <surname>Shechtman</surname></string-name></person-group>, &#x201C;<article-title>Gan-supervised dense visual alignment</article-title>,&#x201D; in <conf-name>2022 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <comment>Jun. 18&#x2013;24</comment>, <year>2022</year>, pp. <fpage>13470</fpage>&#x2013;<lpage>13481</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Laine</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Aila</surname></string-name></person-group>, &#x201C;<article-title>A style-based generator architecture for generative adversarial networks</article-title>,&#x201D; in <conf-name>2019 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Long Beach, CA, USA</publisher-loc>, <comment>Jun. 15&#x2013;20</comment>, <year>2019</year>, pp. <fpage>4401</fpage>&#x2013;<lpage>4410</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T. M.</given-names> <surname>Dinh</surname></string-name>, <string-name><given-names>A. T.</given-names> <surname>Tran</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Nguyen</surname></string-name>, and <string-name><given-names>B. S.</given-names> <surname>Hua</surname></string-name></person-group>, &#x201C;<article-title>Hyperinverter: Improving stylegan inversion via hypernetwork</article-title>,&#x201D; in <conf-name>2022 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <comment>Jun. 18&#x2013;24</comment>, <year>2022</year>, pp. <fpage>11389</fpage>&#x2013;<lpage>11398</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Alaluf</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Tov</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Mokady</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Gal</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Bermano</surname></string-name></person-group>, &#x201C;<article-title>Hyperstyle: Stylegan inversion with hypernetworks for real image editing</article-title>,&#x201D; in <conf-name>2022 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <comment>Jun. 18&#x2013;24</comment>, <year>2022</year>, pp. <fpage>18490</fpage>&#x2013;<lpage>18500</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>J. Y.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Choi</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Fix the noise: Disentangling source feature for transfer learning of StyleGAN</article-title>,&#x201D; <comment>arXiv preprint arXiv:2204.14079</comment>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Rangwani</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bansal</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Sharma</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Karmali</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Jampani</surname></string-name> and <string-name><given-names>R. V.</given-names> <surname>Babu</surname></string-name></person-group>, &#x201C;<article-title>Noisytwins: Class-consistent and diverse image generation through stylegans</article-title>,&#x201D; in <conf-name>2023 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Vancouver, BC, Canada</publisher-loc>, <comment>Jun. 17&#x2013;24</comment>, <year>2023</year>, pp. <fpage>5987</fpage>&#x2013;<lpage>5996</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Aila</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Laine</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Lehtinen</surname></string-name></person-group>, &#x201C;<article-title>Progressive growing of gans for improved quality, stability, and variation</article-title>,&#x201D; in <conf-name>The Six Int. Conf. Learn. Represent.</conf-name>, <publisher-loc>Vancouver, BC, Canada</publisher-loc>, <publisher-name>Vancouver Convention Center</publisher-name>, <year>Apr. 30&#x2013;May 3, 2018</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Goodfellow</surname></string-name> <etal>et al.</etal>,</person-group> &#x201C;<article-title>Generative adversarial nets</article-title>,&#x201D; in <conf-name>NIPS&#x0027;14: Proc. 27th Int. Conf. on Neural Inf. Process. Syst.</conf-name>, <publisher-loc>Montreal, Canada</publisher-loc>, <comment>Dec. 8&#x2013;13</comment>, <year>2014</year>, vol. <volume>2</volume>, pp. <fpage>2672</fpage>&#x2013;<lpage>2680</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Goodfellow</surname></string-name></person-group>, &#x201C;<article-title>NIPS 2016 tutorial: Generative adversarial networks</article-title>,&#x201D; <comment>arXiv preprint arXiv:1406.2661</comment>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Mescheder</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Geiger</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Nowozin</surname></string-name></person-group>, &#x201C;<article-title>Which training methods for gans do actually converge?</article-title>,&#x201D; in <conf-name>Proc. 35th Int. Conf. Mach. Learn. Stockholmsm&#x00E4;ssan</conf-name>, <publisher-loc>Stockholm Sweden</publisher-loc>, <comment>Jul. 10&#x2013;15</comment>, <year>2018</year>, vol. <volume>80</volume>, pp. <fpage>3481</fpage>&#x2013;<lpage>3490</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Woo</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>J. Y.</given-names> <surname>Lee</surname></string-name>, and <string-name><given-names>I. S.</given-names> <surname>Kweon</surname></string-name></person-group>, &#x201C;<article-title>CBAM: Convolutional block attention module</article-title>,&#x201D; in <conf-name>ECCV 2018:15th Eur. Conf.</conf-name>, <publisher-loc>Munich, Germany</publisher-loc>, <comment>Sep. 8&#x2013;14</comment>, <year>2018</year>, pp. <fpage>3</fpage>&#x2013;<lpage>19</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Tang</surname></string-name></person-group>, &#x201C;<article-title>Deep learning face attributes in the wild</article-title>,&#x201D; in <conf-name>2015 IEEE Int. Conf. Comput. Vis. (ICCV)</conf-name>, <publisher-loc>Santiago, Chile</publisher-loc>, <comment>Dec. 7&#x2013;13</comment>, <year>2015</year>, pp. <fpage>3730</fpage>&#x2013;<lpage>3738</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Krizhevsk</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Hinton</surname></string-name></person-group>, &#x201C;<article-title>Learning multiple layers of features from tiny images</article-title>,&#x201D; <italic>Handbook Syst.
Autoimmu. Dis.</italic>, vol. 1, no. 4, <fpage>2009</fpage>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>D. P.</given-names> <surname>Kingma</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Ba</surname></string-name></person-group>, &#x201C;<article-title>Adam: A method for stochastic optimization</article-title>,&#x201D; <comment>arXiv preprint arXiv:1412.6980</comment>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Simonyan</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zisserman</surname></string-name></person-group>, &#x201C;<article-title>Very deep convolutional networks for large-scale image recognition</article-title>,&#x201D; in <conf-name>Int. Conf. Learn. Represent. 2014</conf-name>, <publisher-loc>Banff, Canada</publisher-loc>, <year> Apr. 14&#x2013;16, 2014</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Radford</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Metz</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Chintala</surname></string-name></person-group>, &#x201C;<article-title>Unsupervised representation learning with deep convolutional generative adversarial networks</article-title>,&#x201D; in <conf-name>Int. Conf. Learn. Represent. 2015</conf-name>, <publisher-loc>The Hilton San Diego Resort &#x0026; Spa</publisher-loc>, <year>May 7&#x2013;9, 2015</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Gulrajani</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Ahmed</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Arjovsky</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Dumoulin</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Courville</surname></string-name></person-group>, &#x201C;<article-title>Improved training of wasserstein gans</article-title>,&#x201D; in <conf-name>NIPS&#x2019;17:Proc. 31st Int. Conf. Neural Inf. Process. Syst.</conf-name>, <publisher-loc>Long Beach, California, USA</publisher-loc>, <comment>Dec. 4&#x2013;9</comment>, <year>2017</year>, vol. <volume>30</volume>, pp. <fpage>5769</fpage>&#x2013;<lpage>5779</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Song</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Ermon</surname></string-name></person-group>, &#x201C;<article-title>Generative modeling by estimating gradients of the data distribution</article-title>,&#x201D; in <conf-name>NIPS&#x2019;19: Proc. 33rd Int. Conf. Neural Inf. Process. Syst.</conf-name>, <publisher-loc>Vancouver, BC, Canada</publisher-loc>, <comment>Dec. 8&#x2013;14</comment>, <year>2019</year>, vol. <volume>32</volume>, pp. <fpage>11918</fpage>&#x2013;<lpage>11930</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Laine</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Hellsten</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Lehtinen</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Aila</surname></string-name></person-group>, &#x201C;<article-title>Analyzing and improving the image quality of stylegan</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, <comment>Jun. 13&#x2013;19</comment>, <year>2020</year>, pp. <fpage>8110</fpage>&#x2013;<lpage>8119</lpage>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Chang</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>TransGAN: Two pure transformers can make one strong gan, and that can scale up</article-title>,&#x201D; in <conf-name>NIPS&#x2019;21:Proc. 35th Int. Conf. Neural Inf. Process. Syst.</conf-name>, <publisher-loc>Red Hook, NY, USA</publisher-loc>, <comment>Dec. 6&#x2013;14</comment>, <year>2021</year>, vol. <volume>34</volume>, pp. <fpage>14745</fpage>&#x2013;<lpage>14758</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Tu</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>ViTGAN: Training gans with vision transformers</article-title>,&#x201D; in <conf-name>2021 The Ninth Int. Conf. Learn. Represent.</conf-name>, <publisher-loc>Virtual Only Conference</publisher-loc>, <year>May 3&#x2013;7, 2021</year>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Dual discriminator weighted mixture generative adversarial network for image generation</article-title>,&#x201D; <source>J. Ambient Intell. Humaniz. Comput.</source>, vol. <volume>14</volume>, no. <issue>8</issue>, pp. <fpage>10013</fpage>&#x2013;<lpage>10025</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1007/s12652-021-03667-y</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Gan</surname></string-name></person-group>, &#x201C;<article-title>PFGAN: Fast transformers for image synthesis</article-title>,&#x201D; <source>Pattern Recognit. Lett.</source>, vol. <volume>170</volume>, pp. <fpage>106</fpage>&#x2013;<lpage>112</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.patrec.2023.04.013</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>