<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">36688</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.036688</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Mining Fine-Grain Face Forgery Cues with Fusion Modality</article-title>
<alt-title alt-title-type="left-running-head">Mining Fine-Grain Face Forgery Cues with Fusion Modality</alt-title>
<alt-title alt-title-type="right-running-head">Mining Fine-Grain Face Forgery Cues with Fusion Modality</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Peng</surname><given-names>Shufan</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Cai</surname><given-names>Manchun</given-names></name><email>caimanchun@ppsuc.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Lu</surname><given-names>Tianliang</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Xiaowen</given-names></name></contrib>
<aff><institution>People&#x2019;s Public Security University of China</institution>, <addr-line>Beijing 100038</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Manchun Cai. Email: <email>caimanchun@ppsuc.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>27</day><month>3</month><year>2023</year></pub-date>
<volume>75</volume>
<issue>2</issue>
<fpage>4025</fpage>
<lpage>4045</lpage>
<history>
<date date-type="received">
<day>09</day><month>10</month><year>2022</year>
</date>
<date date-type="accepted">
<day>08</day><month>2</month><year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Peng et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Peng et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_36688.pdf"></self-uri>
<abstract>
<p>Face forgery detection is drawing ever-increasing attention in the academic community owing to security concerns. Despite the considerable progress in existing methods, we note that: Previous works overlooked fine-grain forgery cues with high transferability. Such cues positively impact the model&#x2019;s accuracy and generalizability. Moreover, single-modality often causes overfitting of the model, and Red-Green-Blue (RGB) modal-only is not conducive to extracting the more detailed forgery traces. We propose a novel framework for fine-grain forgery cues mining with fusion modality to cope with these issues. First, we propose two functional modules to reveal and locate the deeper forged features. Our method locates deeper forgery cues through a dual-modality progressive fusion module and a noise adaptive enhancement module, which can excavate the association between dual-modal space and channels and enhance the learning of subtle noise features. A sensitive patch branch is introduced on this foundation to enhance the mining of subtle forgery traces under fusion modality. The experimental results demonstrate that our proposed framework can desirably explore the differences between authentic and forged images with supervised learning. Comprehensive evaluations of several mainstream datasets show that our method outperforms the state-of-the-art detection methods with remarkable detection ability and generalizability.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Face forgery detection</kwd>
<kwd>fine-grain forgery cues</kwd>
<kwd>fusion modality</kwd>
<kwd>adaptive enhancement</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Recent studies have shown rapid advances in face forgery techniques [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>], which allow attackers to perform facial area manipulation at a much lower cost. With the remarkable success represented by Deepfakes, the subtle differences between authentic and forged images are indistinguishable. Face forgery&#x2019;s malicious usage may cause serious social problems and political threats. Therefore, developing high-performance detection methods has become a popular research direction.</p>
<p>Face forgery detection technology intends to prevent the harm caused when Deepfakes technology is abused. Such as preventing Deepfakes technology from manipulating elections [<xref ref-type="bibr" rid="ref-5">5</xref>], interfering with media messages [<xref ref-type="bibr" rid="ref-6">6</xref>], creating pornography featuring female celebrities, creating fake accounts, and financial fraud. Some forgeries involving politics are likely to have unpredictable consequences [<xref ref-type="bibr" rid="ref-7">7</xref>]; in December 2020, DeepFake videos featuring Vladimir Putin and Kim Jong-un appeared on social media, exciting discussions about elections and democracy in the United States. Hence high-performance face forgery detection methods have become a hot research concern. The ideal detection model can be applied to most forgery data in the first place and has good detection ability and generalizability to deal with unseen forgeries.</p>
<p>Researchers have developed various methods to detect face forgery employing distinct traces, such as apparent visual artifacts [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>], temporal inconsistencies [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>], and multimodal conflicts [<xref ref-type="bibr" rid="ref-16">16</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>]. These traces are not universal. Existing methods are more demanding on data and less meaningful in real scenarios. In the real Internet, the vast majority of face forgery is presented as images. Thus we put our research on the image-based face forgery detection method. Previous image-based research focused on modifying the network structure or extracting various features. These methods are based on single-modality or physiological features, which could otherwise be more satisfactory in terms of accuracy and generalization. Spatial-based detection methods generally apply modified visual network architectures to face forgery detection, such as capsule networks [<xref ref-type="bibr" rid="ref-19">19</xref>], Xception [<xref ref-type="bibr" rid="ref-20">20</xref>], vision transformers [<xref ref-type="bibr" rid="ref-21">21</xref>], etc. The above methods&#x2019; robustness is susceptible to image post-processing, such as video compression and smoothing. Some image processing methods, such as frequency analysis, have been introduced for highly compressed datasets to face forgery detection. Durall et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] utilized the unnatural spectral distribution generated by the prevalent generative models for detection. Frank et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] found that the generative adversarial network (GAN)-generated images exhibited severe artifacts in the frequency domain. These methods are still single-modal-only, and the upper limit of performance achieved on different datasets is somewhat constrained. Moreover, these single modality-based detection methods fail to explore forgery patterns. These forgery traces extracted depend heavily on the training data and may fail on unseen forgeries.</p>
<p>Supervised face forgery detection methods rely on neural network fitting capabilities for learning. With a narrow gap between network architectures, how to uncover more critical and more generalizable forgery features becomes a problem worth investigating. From how face forgery images are generated [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>], a forgery face often blends two existing faces or is synthesized by deep neural networks (DNNs). This mode has some similarities with image splicing. Both are similar to the blending of two types of images. However, there are obvious signs of tampering at the boundary between the manipulated region and the genuine region of the spliced image.</p>
<p>In contrast, face forgery images represented by Deepfakes tend to have fine-grain forgery cues, such as visual artifacts and unusual noise, resulting in an anomaly in high-frequency regions. For face forgery detection tasks, local cues play a more critical role than global semantics. Unlike image splicing detection, which utilizes boundary information, several advanced manipulation methods [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>] generate local forgery traces, leading to global facial features&#x2019; discriminability suffering from small-scale tampering. Therefore, exploring the universal local forgery traces is the key to the face forgery detection task.</p>
<p>We observe that if only the RGB modality is employed, detailed local properties are prone to be overlooked as the perceptual field increases. Moreover, we assume that the key to exploring the critical local forgery traces is to exploit the inconsistency in details between authentic and forged images. Several works have proposed solutions in response to this phenomenon. Dang et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] utilized attention maps to locate manipulated parts, Chai et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] segmented images into local patches, and Zhao et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] employed multiple spatial attention heads to focus on the image&#x2019;s different regions. Although the above approaches emphasize local features, local forgery features that rely only on color space are fragile for image post-processing. Future solutions need to be more robust and practical in real scenarios. The noise modality, on the other hand, due to its local properties, its introduction helps the model to learn some local anomalies or local forgery traces. Previous works utilizing image noise modality still intrinsically treat the noise modality as an independent complementary feature to enhance the model&#x2019;s accuracy. Zhou et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] leveraged the complementary properties of RGB and noise streams to detect and locate tampered images efficiently. Luo et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] observed that current convolutional neural network (CNN)-based detectors tend to over-fit color textures and proposed introducing multi-scale noise features to improve generalization across multiple benchmark datasets. Fei et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed a learnable adaptive spatial rich model (ASRM) filter to compensate for conventional noise features&#x2019; shortcomings in adaptive. Previous work ignored the correspondence between noise features and RGB features in the spatial domain. We expect to exploit the spatial commonality of the two modal features to guide the model&#x2019;s perception of local forgery cues.</p>
<p>In contrast to the above approaches, which utilize two modalities of global image features to complement each other, our method is expected to learn more about generalizable forgery patterns. We design a novel fusion enhancement method to introduce the noise modality and employ a particular chunking learning approach to enhance the sensitivity to fine-grain face forgery cues. We design a novel fusion enhancement method to introduce the noise modality and employ a particular chunking learning approach to enhance the sensitivity to fine-grain face forgery cues.</p>
<p>Based on these observations of face forgery image properties, the main motivations behind our work are: (1) In this work, we focus on capturing forgery traces from the perspective of fine-grain face forgery cues. Such local semantics with high transferability have better detectability and generalization. In contrast, learning the global features of images is less important. (2) Specifically, unlike the previous view of frequency information as a separate feature stream, we note that noise features contain some fine-grained local anomalies that are often not easily detectable in RGB features. As an inherent property of images, noise features also correspond to RGB features spatially, and the two features can somewhat complement each other. Therefore, we want to fuse noise features to guide the network to notice such local anomalies of forged images and use them as forgery cues for subsequent sensitive block mining. (3) Since deeper features correspond to larger perceptual fields, a deep network is challenging to learn fine-grained noise features adequately. We design a novel adaptive enhancement method for noise features in the fusion modality that can adaptively adjust the magnitude of the enhancement according to data. (4) We employ a novel chunking learning approach to enhance the network&#x2019;s learning of fine-grained face forgery cues. Specifically, given a face image, we select the sensitive blocks that are most important for detection results by aggregating deep feature descriptors. Unlike previous patch-wise learning methods, our approach adaptively learns vital local forgery patterns while ignoring the less critical features and does not require external annotation.</p>
<p>Our contributions can be summarized as follows:</p>
<p>We propose a novel perspective to address the face forgery detection task, aiming at mining fine-grain face forgery cues to learn the difference between authentic and forged images. To end this, we introduce and adaptively enhance the image noise modality utilizing sensitive blocks to ensure the discrimination between genuine and manipulated regions in deep local features.</p>
<p>We propose two functional modules to reveal and locate the deeper forged features. A dual-modality progressive fusion module (DPFM) is designed to explore dual-modal correlations in spatial and channel dimensions in shallow features and fuse them on this basis. Furthermore, a noise adaptive enhancement module (NAEM) is designed to excavate the artifact hidden in the noise feature adaptively.</p>
<p>We design a sensitive patch branch (SPB) shared with the main network parameters to isolate vital subtle forgery traces. SPB selects as input the sensitive blocks corresponding to the most critical windows for classifiers. SPB effectively enhances the learning of the network for forgery cues, which gives the network remarkable detectability and generalization.</p>
<p>We performed a comprehensive evaluation of mainstream face forgery datasets. The experimental results demonstrate our proposed method&#x2019;s effectiveness with the most advanced competitors.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p><bold>Spatial-based forgery detection methods.</bold> Early face forgery techniques tend to generate obvious forged signs. Many manual feature-based methods explore forgery image anomalies in the spatial domain. These manual features include image noise residual [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>], face warping artifacts (FWA) [<xref ref-type="bibr" rid="ref-8">8</xref>], visual artifacts [<xref ref-type="bibr" rid="ref-9">9</xref>], etc. Face X-ray [<xref ref-type="bibr" rid="ref-36">36</xref>] is based on the property of blending boundaries when faces are swapped, significantly generalizing local forged images. However, this method achieved undesirable results in highly compressed or entire synthetic datasets. Given the excellent performance of deep neural networks in computer vision, DNN-based methods have gradually become the mainstream of research. Some works directly applied existing classification networks [<xref ref-type="bibr" rid="ref-37">37</xref>&#x2013;<xref ref-type="bibr" rid="ref-39">39</xref>]. MesoNet [<xref ref-type="bibr" rid="ref-40">40</xref>] utilized a shallow CNN architecture for forgery detection based on mid-level semantics. Bayer et al. [<xref ref-type="bibr" rid="ref-41">41</xref>] developed a new convolutional layer capable of adaptively learning manipulation detection features. Current state-of-the-art methods explore and learn about forged features. Dang et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] decomposed the face forgery detection task into the localization of manipulated regions and detection. Zhao et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] modeled the detection task as a fine-grain image classification task, leading to learning the proposed Multi-attentional Deepfake Detection (MAT) framework for local forgery traces and shallow texture features in manipulated images.</p>
<p>The above spatial-based detection methods tend to modify the network structure. Their detection performance varies widely across diverse datasets. RGB-modal-only Methods can render the detector overfit to method-specific color texture and thus fail to generalize. Furthermore, these methods are highly impacted by image post-processing, such as compression and smoothing masks&#x2014;lack applicability in real scenarios.</p>
<p><bold>Frequency-based forgery detection methods.</bold> The spatial-based detection methods are susceptible to compression rate. There are anomalies, such as distribution differences of high-frequency components and checkerboard artifacts in the synthetic images. Furthermore, researchers have applied many traditional mathematical methods to practical tasks [<xref ref-type="bibr" rid="ref-42">42</xref>&#x2013;<xref ref-type="bibr" rid="ref-45">45</xref>] with impressive results in recent years. Thus, frequency analysis is introduced into the detection task and achieves desired results in highly compressed datasets. Some works utilized digital image processing methods such as Discrete Fourier Transform (DFT) [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>] and Discrete Cosine Transform (DCT) [<xref ref-type="bibr" rid="ref-48">48</xref>] to obtain frequency domain features and detect anomalies. Frequency in Face Forgery Network (F<sup>3</sup>-Net) [<xref ref-type="bibr" rid="ref-49">49</xref>] proposed frequency-aware decomposition and local frequency statistics to obtain forgery information in the frequency domain. Fake Generated Painting Detection via Frequency Analysis (FGPD-FA) [<xref ref-type="bibr" rid="ref-50">50</xref>] performs forgery detection by fusing three distinct frequency domain features.</p>
<p>However, since different forgery generation methods vary dramatically in the frequency domain space, we observe that the accuracy of the frequency-based detection method alone is substantially reduced on unseen datasets. Most existing frequency analysis-based methods directly convert the entire image into a spectrum. Locally tampered faces [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-51">51</xref>] show indistinct anomalies in the global frequency domain. Thus, these methods also suffer from subtle local forgery traces.</p>
<p><bold>Forgery detection methods combine spatial and frequency features.</bold> Due to single-modality limitations, dual-modality-based detection methods are becoming mainstream research [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>]. Spatial-Phase Shallow Learning (SPSL) [<xref ref-type="bibr" rid="ref-53">53</xref>] employed a shallow network to capture the pixel differences in the phase spectrum of the synthetic images. The shallow network makes it difficult for the method to detect subtle forgery traces. Frequency-aware Discriminative Feature Learning (FDFL) [<xref ref-type="bibr" rid="ref-54">54</xref>] designed a single-center loss to improve intra-class compactness and inter-class separability in the embedding space with dual-modal features. Multimodal Contrastive Classification by Locally Correlated Representations (MC-LCR) [<xref ref-type="bibr" rid="ref-55">55</xref>] proposed a novel perspective that aims to amplify the implicit local discrepancies between authentic and forged face images from dual-modality features.</p>
<p>Previous works treat spatial and frequency domain features as two separate feature streams but neglect the existence of correspondence between the two features in terms of location. In essence, they need to explore the forgery cues on dual-modal features further.</p>
<p><bold>Forgery detection methods utilize local receptive fields.</bold> Previous methods of patch-wise training tend to perform even chunking [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-55">55</xref>]. We note that existing patch-wise detection methods ignore the variation in the forged features between patches and lack of adaptivity. To end this, we introduce more flexible activation map-based sensitive patches, which can extract vital features of arbitrary size. The sensitive patch-based detection method can improve our framework&#x2019;s accuracy and generalization. Sensitive patches can enhance the learning of manipulated patterns rather than global features, such as visual artifacts. In addition, such local semantics makes the detector less susceptible to high-level facial image features, achieving better generalization.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>The Proposed Method</title>
<p>Owing to post-processing with compression or smoothing in mainstream datasets, the difference between authentic and forged images is difficult to discriminate by RGB features alone. Face forgery images usually consist of authentic areas as well as forged areas. Noise features are inherent to the specificity of images. The noise features of the post-generated forged region and the genuine region are difficult to match. We, therefore, carefully design two modules to integrate RGB features with noise features fully. Furthermore, extract important local features employing sensitive patches to guide the learning of our framework for fine-grained forgery cues. Our framework is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>In the dual-modality progressive fusion module, dual-modal spatial and channel correlations are mined separately using the spatial feature interaction module and channel attention fusion module. Different levels of noise features are enhanced adaptively in the noise adaptive enhancement module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-1.tif"/>
</fig>
<p>Our framework&#x2019;s overall training and testing process is illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. Our framework contains a novel modal fusion-enhancement process and an adaptive sensitive patch mining-learning process. Specifically, in the training phase, we obtain critical sensitive patches based on fusion modality and input them into a sensitive patch branch shared with the main network parameters. In the testing phase, we directly use the main network for testing.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The overall training and testing process of our framework</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-2.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Dual-Modality Progressive Fusion Module</title>
<p>Most of the existing face forgery generators are based on GAN, where the up-sampling causes anomalies in the frequency domain features of the face forgery images. As an inherent property of the image, the noise features of the manipulated region are often inconsistent with those of the genuine region. The image&#x2019;s post-forged part leaves a unique trace in the noise space, and this location information can correspond well to the RGB space. This property helps our proposed framework to locate high discrepancy regions between authentic and forged images. We thus introduce noise features to guide our framework in mining the differences between authentic and forged images.</p>
<p>In previous work, RGB and noise features were often treated as two separate feature streams and concatenated directly in the high-level features on the channel dimension. However, the RGB and noise features are not wholly unrelated features. There is quite a lot of shared information in these two modalities. Two-Stream networks may weaken the correlation between spatial features and cause redundancy in the network structure. Therefore, we propose a progressive fusion method to fully obtain the spatial and channel features of the two modalities.</p>
<p>As the network deepens, semantic information increases as the reception field increases, and spatial information diminish as the resolution decreases. We need to retain the spatial information in both modalities for sensitive patch mining. We use a progressive fusion strategy on feature maps of different resolutions in the fusion process. To this end, we propose a dual-modality progressive fusion module in the shallow layer to fully fuse the dual-modal features at different scales. Our proposed module consists of two sub-modules: a spatial feature interaction module and a channel attention fusion module (See <xref ref-type="fig" rid="fig-3">Fig. 3</xref>), which explore the spatial and channel correlations between the two modalities separately.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The spatial feature interaction module and the channel attention fusion module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-3.tif"/>
</fig>
<p>In the spatial feature interaction module, let <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the input of RGB features and noise features. First, we obtain the information within each modality by <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Then, stitching in spatial dimensions to obtain the fused modalities <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. On this basis, the fusion attention map <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is obtained:</p>
<p><disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>Q</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2297;</mml:mo><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mo>&#x2297;</mml:mo></mml:math></inline-formula> denotes the multiply operator, the final output is <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>:</p>
<p><disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>B</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2297;</mml:mo><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>B</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2297;</mml:mo><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>B</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> denotes batch normalization, <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mo>+</mml:mo></mml:math></inline-formula> denotes the sum operator. <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup></mml:math></inline-formula> can interact with the features of another modality effectively.</p>
<p>The channel attention fusion module utilizes the attention mechanism to facilitate inter-channel interactions. This module combines information from all channels and determines each channel&#x2019;s significance. Let <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mrow><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the fusion modality feature maps of the input and output:</p>
<p><disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:msubsup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi></mml:mi><mml:mo>&#x2032;</mml:mo></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>&#x2297;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>G</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>G</mml:mi><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Fusion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p><p>where <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>C</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> is the concatenate operator in the channel dimension, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> is a multi-layer perceptron, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>G</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> denotes the global average pooling, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>G</mml:mi><mml:mi>M</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> denotes the global max pooling, and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is the sigmoid activation function. We adopt this progressive fusion method to fully obtain the information of both modalities, using this fusion modality as a basis for mining the differences between authentic and forged images.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Noise Adaptive Enhancement Module</title>
<p>For datasets with insignificant forgery features, subtle noise will play a critical role in forgery cue mining. The neural network has a low learning priority spectral bias for high frequencies. Furthermore, the sizeable perceptual field of high-level features is challenging to extract local noises. To this end, we design a noise adaptive enhancement module (See <xref ref-type="fig" rid="fig-4">Fig. 4</xref>) to excavate the artifact hidden in the noise feature.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The noise adaptive enhancement module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-4.tif"/>
</fig>
<p>In the noise adaptive enhancement module, for the input feature map <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, we first conduct two transformations:</p>
<p><disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>SRM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p><p>where <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Corresponding to Xception, we use the depthwise separable convolution with a <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> kernel. We generate channel-wise statistics <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> by fusing the RGB branch and enhanced noise branch via element-wise summation and global average pooling:</p>
<p><disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>GAP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>H</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Then, the squeeze-excitation operation is used to learn the feature associations on the channel dimension. We obtain the compact feature descriptor <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mrow><mml:mtext>z</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and normalize the soft attention vectors <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of the corresponding channels of two branches via a softmax operator.</p>
<p><disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mtext>z</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>Squeeze</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>softmax</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">(</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>Excitation-RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>Excitation-EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula></p>
<p>The final feature map is obtained through the attention weights of two branches:</p>
<p><disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RGB</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>EHM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the fusion modality feature map after adaptive enhancement, two noise adaptive enhancement modules are inserted between the blocks of the backbone network, preserving and amplifying subtle noise in low-level and high-level features, respectively. More helpful information is provided for subsequent sensitive patch mining.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Sensitive Patch Branch</title>
<p>Across different forgery face generators, local forgery cues tend to have higher generalizability than global structure. The global structure may vary among advanced semantic information about different faces. Nevertheless, local cues have better commonality across different manipulation methods, such as shared visual artifact features. We, therefore, hypothesize that local cues are more conducive to face forgery detection. In previous work, the features of interest to these networks were scattered, with relatively poor generalizability. Some work uses image chunking or masking of specific regions to limit the network&#x2019;s perception field for learning. Such methods may lose some of the features of the raw image or lack adaptivity. In particular, for partial forgery images, the network may have difficulty converging during the training process if many non-forged regions are included in the patches labeled as fake.</p>
<p>We propose a fine-grain forgery cue mining method (See <xref ref-type="fig" rid="fig-5">Fig. 5</xref>) to precisely locate forged regions and use them as sensitive patches to enhance network learning for different forgery cues. Specifically, we use adaptive-sized sliding windows for extracting these sensitive patches. Furthermore, input these patches into the sensitive patch branch that shares parameters with the main network during the training phase. Our framework is ultimately more oriented towards the learning of crucial manipulated patterns.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The sensitive patch branch. The network parameters in the purple box are shared. The blue box shows the extraction process of sensitive patches</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-5.tif"/>
</fig>
<p>We aggregate the high-level feature maps <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> in the channel dimension to obtain the corresponding activation maps <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>:</p>
<p><disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>i</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> feature map of <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the coordinates of the feature descriptor in space. The value of <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> reflects the contribution of the descriptor to the result. To select patches containing more sensitive information, we use the average activation value within the sliding window to define the <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">e</mml:mi></mml:mrow></mml:math></inline-formula> of these patches:</p>
<p><disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">e</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>H</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>W</mml:mi></mml:math></inline-formula> are the height and width of windows, we sort by the <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">e</mml:mi></mml:mrow></mml:math></inline-formula> of all windows and select the first few windows as sensitive patches. We adopt a non-maximum suppression (NMS) between the extracted sensitive patches to mine more forgery cues. The sensitive patches with the highest <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">e</mml:mi></mml:mrow></mml:math></inline-formula> values are eventually used as input to the branch. To thoroughly learn the subtle forgery traces within sensitive patches, we adopt the Circle Loss as the metric loss to improve the intra-class compactness and inter-class separability in the embedding space.</p>
<p>Our objective function contains the cross-entropy loss of the main network and patch branch and the metric loss between sensitive patches:</p>
<p><disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Raw-Softmax</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>raw</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">b</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Patch-Softmax</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>patch</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">b</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="ueqn-16"><label>(16)</label><mml:math id="mml-ueqn-16" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Patch-Metric</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Circle</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mrow><mml:mtext>patch</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mrow><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">b</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow></mml:math></inline-formula> is the raw image&#x2019;s ground truth label, and <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>raw</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the raw image&#x2019;s category probability. <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>patch</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the softmax layer&#x2019;s output of the sensitive patch branch corresponding to the <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>n</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> sensitive patch. <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>P</mml:mi><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mrow><mml:mtext>patch</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the feature embedding of the <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>n</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> sensitive patch. Since the unstable accuracy of the activation map in the first epoch, the total loss is as follows:</p>
<p><disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>total&#xA0;</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Raw-Softmax</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mrow><mml:mtext>&#xA0;if epoch</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Raw-Softmax</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Patch-Softmax</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Patch-Metric</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mrow><mml:mtext>&#xA0;otherwise&#xA0;</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Settings</title>
<p><bold>Datasets.</bold> We adopt five widely-used public datasets in our experiments, i.e., FaceForensics&#x002B;&#x002B; (FF&#x002B;&#x002B;) [<xref ref-type="bibr" rid="ref-56">56</xref>], FaceShifter (FSR) [<xref ref-type="bibr" rid="ref-57">57</xref>], DeepfakeDetection (DFD) [<xref ref-type="bibr" rid="ref-58">58</xref>], Celeb-DF [<xref ref-type="bibr" rid="ref-26">26</xref>], DeeperForensics-1.0 (DF1.0) [<xref ref-type="bibr" rid="ref-59">59</xref>], and WildDeepfake [<xref ref-type="bibr" rid="ref-60">60</xref>] (See <xref ref-type="table" rid="table-1">Table 1</xref>). We uniformly set the ratio of the training and testing sets to <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mn>7</mml:mn><mml:mrow><mml:mo>&#x003A;</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>3</mml:mn></mml:math></inline-formula>. We take the high-quality version (c23) and the low-quality version (c40) of FF&#x002B;&#x002B; and FSR. FF&#x002B;&#x002B; contains four manipulation methods: Deepfakes (DF), Face2Face (F2F) [<xref ref-type="bibr" rid="ref-51">51</xref>], FaceSwap (FSP) [<xref ref-type="bibr" rid="ref-61">61</xref>], and NeuralTextures (NT) [<xref ref-type="bibr" rid="ref-27">27</xref>]. As the level of compression increases, detection of forgery cues can become increasingly challenging.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Specifications of benchmark databases</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Database</th>
<th>Video scale</th>
<th>Manipulation algorithm</th>
</tr>
</thead>
<tbody>
<tr>
<td>FF&#x002B;&#x002B;(c23/c40)</td>
<td>1000 real, 4000 fake</td>
<td>DF, FSP, NT, F2F</td>
</tr>
<tr>
<td>FaceShifter(c23/c40)</td>
<td>1000 fake</td>
<td>FSR</td>
</tr>
<tr>
<td>DFD</td>
<td>363 real, 3068 fake</td>
<td>Improved DF</td>
</tr>
<tr>
<td>Celeb-DF</td>
<td>590 real, 5639 fake</td>
<td>Improved DF</td>
</tr>
<tr>
<td>DF1.0</td>
<td>11,000 fake</td>
<td>DF-VAE</td>
</tr>
<tr>
<td>WildDeepfake</td>
<td>3805 real, 3509 fake</td>
<td>Improved DF</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Data preprocessing.</bold> Since the dominant face forgery dataset today is in video format, some preprocessing is necessary for our task. We extract keyframes for each video every 10 s, for a total of 30. This process can avoid data redundancy while maintaining the richness of the sample images. For these keyframes, we use Retinaface [<xref ref-type="bibr" rid="ref-62">62</xref>] for face extraction and alignment and resize the aligned faces to <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mn>299</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>299</mml:mn></mml:math></inline-formula>. This processing has become a default standard in face forgery detection and facilitates comparing our method with other state-of-the-art works.</p>
<p><bold>Implementation detail.</bold> Xception pre-trained on ImageNet is adopted as the backbone of our framework, which is trained with AdamW optimizer with a learning rate of <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, weight decay of <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and batch size of 8. The cosine decay is with a total of 20 epochs. We obtain the comparison results from their paper and specify our implementation by &#x2020; otherwise.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Ablation Study</title>
<p><bold>Parameter influence.</bold> The number of sensitive patches tagged as <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula>, and the minimum threshold of window size tagged as <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> may affect the final result. Primarily, <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> controls the hierarchy of mined forgery cues and flexibility. Furthermore, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula> controls our framework&#x2019;s ability to extract forgery cues. We conducted an empirical analysis based on the c23 version and the c40 version of FF&#x002B;&#x002B; to study the optimal values of the two hyper-parameters.</p>
<p>We observe that the sensitive patch&#x2019;s window size correlates with <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mrow><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">e</mml:mi></mml:mrow></mml:math></inline-formula> negatively. The window size corresponding to the extracted sensitive patch tends to be around <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula>. If <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> is too small, the percentage of sensitive patches is small and relatively concentrated, which is not conducive to mining forgery cues at different levels. Conversely, it may contain irrelevant regions that lack flexibility and affect the training phase. As <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula> increases, initially sensitive patches may capture more forgery cues, but too large may introduce some authentic background regions to the detriment of the final result. As illustrated in <xref ref-type="fig" rid="fig-6">Figs. 6a</xref> and <xref ref-type="fig" rid="fig-6">6b</xref>, we get the best results when setting <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> to be <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mn>4</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>4</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula> to be 5. Area Under Curve (AUC) reached 0.9961 on the c23 version and 0.9377 on the c40 version.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The detection performances achieve by (a) varying <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> when <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula> is fixed as 5 and (b) varying <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula> when <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula> is fixed as <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mn>4</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>4</mml:mn></mml:math></inline-formula></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-6.tif"/>
</fig>
<p><bold>Components.</bold> To demonstrate the effectiveness of each module, we evaluate the proposed framework and its variants by intra-testing (See <xref ref-type="table" rid="table-2">Table 2</xref>) and cross-testing (See <xref ref-type="table" rid="table-3">Table 3</xref>).</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Ablation study on FF&#x002B;&#x002B; (c40) and FaceShifter (c40). The metric is AUC</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>DF</th>
<th>F2F</th>
<th>FSP</th>
<th>NT</th>
<th>FSR</th>
</tr>
</thead>
<tbody>
<tr>
<td>RGB (Xception)</td>
<td>0.9791</td>
<td>0.9453</td>
<td>0.9441</td>
<td>0.7877</td>
<td>0.9658</td>
</tr>
<tr>
<td>RGB &#x002B; SPB</td>
<td>0.9940</td>
<td>0.9654</td>
<td>0.9862</td>
<td>0.8232</td>
<td>0.9834</td>
</tr>
<tr>
<td>RGB &#x002B; SRM (Baseline)</td>
<td>0.9865</td>
<td>0.9602</td>
<td>0.9778</td>
<td>0.8057</td>
<td>0.9798</td>
</tr>
<tr>
<td>Baseline &#x002B; SPB</td>
<td>0.9956</td>
<td>0.9684</td>
<td>0.9887</td>
<td>0.8415</td>
<td>0.9861</td>
</tr>
<tr>
<td>Baseline &#x002B; SPB &#x002B; DPFM</td>
<td>0.9969</td>
<td>0.9740</td>
<td>0.9901</td>
<td>0.8443</td>
<td>0.9903</td>
</tr>
<tr>
<td>Baseline &#x002B; SPB &#x002B; DPFM &#x002B; NAEM</td>
<td><bold>0.9982</bold></td>
<td><bold>0.9806</bold></td>
<td><bold>0.9943</bold></td>
<td><bold>0.8564</bold></td>
<td><bold>0.9937</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Ablation study from FF&#x002B;&#x002B;(c23) to Celeb-DF. The metric is AUC</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Celeb-DF</th>
</tr>
</thead>
<tbody>
<tr>
<td>RGB (Xception)</td>
<td>0.6527</td>
</tr>
<tr>
<td>RGB &#x002B; SPB</td>
<td>0.7451</td>
</tr>
<tr>
<td>RGB &#x002B; SRM (Baseline)</td>
<td>0.7255</td>
</tr>
<tr>
<td>Baseline &#x002B; SPB</td>
<td>0.7629</td>
</tr>
<tr>
<td>Baseline &#x002B; SPB &#x002B; DPFM</td>
<td>0.7724</td>
</tr>
<tr>
<td>Baseline &#x002B; SPB &#x002B; DPFM &#x002B; NAEM</td>
<td><bold>0.7868</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Taking Xception as the Baseline, RGB means RGB modality as the input, spatial rich model (SRM) means noise modality as the input, and Baseline means fusion modality as the input. SPB, DPFM, and NAEM represent the sensitive patch branch, dual-modality progressive fusion module, and noise adaptive enhancement module.</p>
<p>We obtain the following conclusions from this experiment. First, the sensitive patch branch can effectively improve the performance of our framework by fine-grain feature learning in both RGB and fusion modalities. In addition, the improvement is insignificant if sensitive patches are mined directly based on the fusion modality compared to the RGB-only modality. The experimental results demonstrate that our specifically designed functional modules can better capture and utilize noise information. As shown in the last three rows, the model&#x2019;s performance gradually improves as each module is added, demonstrating the effectiveness of each module.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Comparison with Recent Works</title>
<p><bold>Within-manipulation-method evaluation.</bold> We compare our method with previous detection methods using frequency domain features on FF&#x002B;&#x002B; and FSR datasets. The results are shown in <xref ref-type="table" rid="table-4">Tables 4</xref> and <xref ref-type="table" rid="table-5">5</xref>. Xception, F<sup>3</sup>-Net, SPSL, and Generalizing Face Forgery Detection (GFF) are the most advanced methods for face forgery detection. (Since SPSL is not open source, we are unable to reproduce their method in our comparison experiments)</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Within-method evaluation of five manipulation techniques (c40). The metric is AUC</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>DF</th>
<th>F2F</th>
<th>FSP</th>
<th>NT</th>
<th>FSR</th>
</tr>
</thead>
<tbody>
<tr>
<td>Xception&#x2020; [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>0.9791</td>
<td>0.9453</td>
<td>0.9442</td>
<td>0.7878</td>
<td>0.9657</td>
</tr>
<tr>
<td>F<sup>3</sup>-Net&#x2020; [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>0.9954</td>
<td>0.9713</td>
<td>0.9879</td>
<td>0.8264</td>
<td>0.9873</td>
</tr>
<tr>
<td>GFF&#x2020; [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>0.9945</td>
<td>0.9574</td>
<td>0.9592</td>
<td>0.7981</td>
<td>0.9735</td>
</tr>
<tr>
<td>SPSL [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>0.9850</td>
<td>0.9462</td>
<td>0.9810</td>
<td>0.8049</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Ours</td>
<td><bold>0.9982</bold></td>
<td><bold>0.9806</bold></td>
<td><bold>0.9943</bold></td>
<td><bold>0.8564</bold></td>
<td><bold>0.9937</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Within-method evaluation of five manipulation techniques (c23). The metric is AUC</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>DF</th>
<th>F2F</th>
<th>FSP</th>
<th>NT</th>
<th>FSR</th>
</tr>
</thead>
<tbody>
<tr>
<td>Xception&#x2020; [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>0.9956</td>
<td>0.9951</td>
<td>0.9894</td>
<td>0.9523</td>
<td>0.9910</td>
</tr>
<tr>
<td>F<sup>3</sup>-Net&#x2020; [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>0.9995</td>
<td>0.9989</td>
<td>0.9993</td>
<td>0.9884</td>
<td>0.9967</td>
</tr>
<tr>
<td>GFF&#x2020; [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>0.9942</td>
<td>0.9973</td>
<td>0.9969</td>
<td>0.9691</td>
<td>0.9942</td>
</tr>
<tr>
<td>SPSL [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Ours</td>
<td><bold>0.9999</bold></td>
<td><bold>0.9994</bold></td>
<td><bold>0.9998</bold></td>
<td><bold>0.9933</bold></td>
<td><bold>0.9976</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Our framework outperforms other advanced detection methods in terms of the AUC index on different versions of the five manipulation methods. Especially in the most challenging c40 version, the AUC metrics achieve more significant improvements. The experimental results show that our method can effectively capture the features of the manipulation method and achieve the desired detection results.</p>
<p><bold>Within-database evaluation.</bold> We compared our method with several state-of-the-art methods. The results are shown in <xref ref-type="table" rid="table-6">Table 6</xref>. Our framework can utilize fusion modality to focus more on manipulated patterns rather than global semantic information. It still has considerable advantages in large datasets composed of multiple manipulation methods.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Within-database evaluation on the FF&#x002B;&#x002B; dataset with c23 and c40 versions and WildDeepfake dataset. The metric is accuracy (Acc) and AUC</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th align="center" colspan="2">FF&#x002B;&#x002B; (c23)</th>
<th align="center" colspan="2">FF&#x002B;&#x002B; (c40)</th>
<th align="center" colspan="2">WildDeepfake</th>
</tr>
<tr>
<th/>
<th>Acc</th>
<th>AUC</th>
<th>Acc</th>
<th>AUC</th>
<th>Acc</th>
<th>AUC</th>
</tr>
</thead>
<tbody>
<tr>
<td>Fridrich et al. [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>70.97%</td>
<td>&#x2013;</td>
<td>55.98%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Cozzolino et al. [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>78.45%</td>
<td>&#x2013;</td>
<td>58.69%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Bayar and Stamm [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>82.97%</td>
<td>&#x2013;</td>
<td>66.84%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Rahmouni et al. [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>79.08%</td>
<td>&#x2013;</td>
<td>61.18%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>DSP-FWA [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>&#x2013;</td>
<td>0.5750</td>
<td>&#x2013;</td>
<td>0.6230</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>MesoNet [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>83.10%</td>
<td>&#x2013;</td>
<td>70.47%</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Face X-ray [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>&#x2013;</td>
<td>0.8735</td>
<td>&#x2013;</td>
<td>0.6160</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>SPSL [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>91.50%</td>
<td>0.9532</td>
<td>81.57%</td>
<td>0.8282</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Xception [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>90.88%&#x2020;</td>
<td>0.9347&#x2020;</td>
<td>80.32%&#x2020;</td>
<td>0.8176&#x2020;</td>
<td>78.42%&#x2020;</td>
<td>0.8677&#x2020;</td>
</tr>
<tr>
<td>Add-Net [<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>96.78%</td>
<td>0.9774</td>
<td>87.50%</td>
<td>0.9101</td>
<td>77.01%&#x2020;</td>
<td>0.8365&#x2020;</td>
</tr>
<tr>
<td>F<sup>3</sup>-Net [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>97.52%</td>
<td>0.9810</td>
<td><bold>90.43%</bold></td>
<td>0.9330</td>
<td>80.78%&#x2020;</td>
<td>0.8756&#x2020;</td>
</tr>
<tr>
<td>FDFL [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>96.69%</td>
<td>0.9930</td>
<td>89.00%</td>
<td>0.9240</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>MAT [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>97.60%</td>
<td>0.9927</td>
<td>88.69%</td>
<td>0.9040</td>
<td>82.23%&#x2020;</td>
<td>0.9098&#x2020;</td>
</tr>
<tr>
<td>Ours</td>
<td><bold>97.83%</bold></td>
<td><bold>0.9961</bold></td>
<td>89.62%</td>
<td><bold>0.9377</bold></td>
<td><bold>83.76%</bold></td>
<td><bold>0.9130</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Cross-database evaluation.</bold> In the actual situation, we cannot predict the means of face image forgery. Generalizability is also an essential criterion for evaluating detection models. The generalizability of the model directly affects its practical application value. A comprehensive cross-database evaluation is performed in this section to check the generalizability of our method. Our framework is trained on different versions of the FF&#x002B;&#x002B; dataset and evaluated on DFD, WildDeepfake, Celeb-DF, and DF1.0. In this section, our method is compared with several advanced methods in terms of generalizability.</p>
<p>Specifically, cross-database evaluation is more challenging due to the difference in the distribution of the training and testing sets. The compression level of DF1.0 is c23. Thus, we do not use DF1.0 to evaluate the generalizability of the method when it is trained on the c40 version. From <xref ref-type="table" rid="table-7">Table 7</xref>, we can observe that our method significantly outperforms the rest of the competitors in almost all datasets. This advantage of generalizability can be further extended to about 5% to 6% with low-quality images as the training set, as shown in <xref ref-type="table" rid="table-8">Table 8</xref>.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Cross-database evaluation from FF&#x002B;&#x002B; (c23) to DFD, WildDeepfake, Celeb-DF, and DF1.0. The metric is AUC. Results in gray indicate the within-database performance</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th style="background:#D9D9D9;">FF&#x002B;&#x002B;(c23)</th>
<th>DFD</th>
<th>WildDeepfake</th>
<th>Celeb-DF</th>
<th>DF1.0</th>
</tr>
</thead>
<tbody>
<tr>
<td>Xception [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td style="background:#D9D9D9;">0.9347</td>
<td>0.8413</td>
<td>0.6617</td>
<td>0.6527</td>
<td>0.6824</td>
</tr>
<tr>
<td>EfficientNet-B4 [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td style="background:#D9D9D9;">0.9422</td>
<td>0.8737</td>
<td>0.6140</td>
<td>0.6852</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Face X-ray [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td style="background:#D9D9D9;">0.8735</td>
<td>0.8560</td>
<td>&#x2013;</td>
<td>0.7420</td>
<td>0.7230</td>
</tr>
<tr>
<td>MLDG [<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td style="background:#D9D9D9;">0.9899</td>
<td>0.8814</td>
<td>0.6412</td>
<td>0.7456</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>F<sup>3</sup>-Net [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td style="background:#D9D9D9;">0.9810</td>
<td>0.8610</td>
<td>0.6771</td>
<td>0.7121</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>MAT [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td style="background:#D9D9D9;">0.9927</td>
<td>0.8758</td>
<td>0.7015</td>
<td>0.7665</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>LTW [<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td style="background:#D9D9D9;">0.9917</td>
<td>0.8856</td>
<td>0.6712</td>
<td>0.7714</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Local-relation [<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
<td style="background:#D9D9D9;">0.9946</td>
<td>0.8924</td>
<td>0.6876</td>
<td>0.7826</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>GFF [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td style="background:#D9D9D9;">0.9930</td>
<td>0.9190</td>
<td>&#x2013;</td>
<td><bold>0.7940</bold></td>
<td>0.9380</td>
</tr>
<tr>
<td>Ours</td>
<td style="background:#D9D9D9;"><bold>0.9961</bold></td>
<td><bold>0.9329</bold></td>
<td><bold>0.7213</bold></td>
<td>0.7868</td>
<td><bold>0.9421</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Cross-database evaluation from FF&#x002B;&#x002B; (c40) to DFD, WildDeepfake, and Celeb-DF. The metric is AUC. Results in gray indicate the within-database performance</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th style="background:#D9D9D9;">FF&#x002B;&#x002B;(c40)</th>
<th>DFD</th>
<th>WildDeepfake</th>
<th>Celeb-DF</th>
</tr>
</thead>
<tbody>
<tr>
<td>Xception [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td style="background:#D9D9D9;">0.8176</td>
<td>0.6413</td>
<td>0.6059</td>
<td>0.6218</td>
</tr>
<tr>
<td>Add-Net [<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td style="background:#D9D9D9;">0.9101</td>
<td>0.5736</td>
<td>0.5421</td>
<td>0.5603</td>
</tr>
<tr>
<td>F<sup>3</sup>-Net [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td style="background:#D9D9D9;">0.9330</td>
<td>0.6988</td>
<td>0.6039</td>
<td>0.6877</td>
</tr>
<tr>
<td>MAT [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td style="background:#D9D9D9;">0.9040</td>
<td>0.7418</td>
<td>0.6549</td>
<td>0.6904</td>
</tr>
<tr>
<td><bold>Ours</bold></td>
<td style="background:#D9D9D9;"><bold>0.9377</bold></td>
<td><bold>0.8134</bold></td>
<td><bold>0.7367</bold></td>
<td><bold>0.7830</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Our method achieves superior results on cross-database evaluation. We also note that supervised learning methods inevitably lead the framework to focus on textures generated by specific manipulation methods. It leads to difficulties in achieving the desired results in the cross-manipulation-method evaluation. Thus, how generalizing forgery cues among unknown manipulation methods is also a problem worth investigating in the future. MLDG [<xref ref-type="bibr" rid="ref-63">63</xref>] and LTW [<xref ref-type="bibr" rid="ref-64">64</xref>] represent Meta-Learning for Domain Generalization and Learning-To-Weight, respectively.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Visualizations</title>
<p><bold>Class Activation Mapping (CAM).</bold> We use Grad-CAM to visualize the attention map and thus explore the regions of interest for the detection method, as shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. We compare the differences between our method and Xception class activation mapping. We observe that Xception&#x2019;s attention is scattered and focused on a larger area, sometimes exceeding the scope of the forged area. Furthermore, it is unreasonable that Xception focuses on similar areas for different manipulation methods of images. Our framework can find key forgery cues and focus attention evenly on different manipulation methods separately. Our method reaches better detection capability and generalization.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>The attention maps of Xception and our framework for different kinds of faces</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-7.tif"/>
</fig>
<p><bold>Adaptive Enhancement.</bold> <xref ref-type="fig" rid="fig-8">Fig. 8</xref> illustrates the enhanced region of different level feature maps via the noise adaptive enhancement module. We observe that for low-level feature maps, the adaptive enhancement regions are more discrete and vary considerably depending on the input. As for the high-level feature maps, the proposed module focuses on potentially manipulated regions, such as the nose and mouth. This noise adaptive enhancement mechanism can adaptively find and enhance helpful subtle forgery cues among different levels of feature maps.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Adaptive enhancement of feature maps at different levels</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-8.tif"/>
</fig>
<p><bold>Sensitive Patch.</bold> We visualize the mined sensitive patches in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. The sensitive patches are represented as red, orange, yellow, and green. We observe that in RGB modality, the forged features are mainly contained in the low-frequency part, such as irregular color blocks. Furthermore, the high-frequency part plays a more prominent role in the noise modality, such as abnormal face detail features and visual artifacts. There is a spatial correspondence between the two modalities so that they can complement each other to some extent in the fusion modality.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Sensitive patches on RGB modality and noise modality</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_36688-fig-9.tif"/>
</fig>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>We propose a novel perspective of fine-grain forgery cues mining with fusion modality to address the face forgery detection task. Firstly, a dual-modality progressive fusion module is designed to complement the single-modal features by interacting and fusing the RGB and noise modalities at different scales. A noise adaptive enhancement module is subsequently designed to enhance the subtle noise features in the feature maps of different levels. Learning key manipulated patterns is achieved by mining subtle forgery traces in the sensitive patch branch. Experiments demonstrate that our method has considerable accuracy and generalization advantages. The visualization of class activation mappings, feature maps, and sensitive patches reveals the intrinsic mechanism of our method and explains its effectiveness.</p>
</sec>
</body>
<back>
<sec><title>Funding Statement</title>
<p>This study is supported by the Fundamental Research Funds for the <funding-source>Central Universities of PPSUC</funding-source> under Grant <award-id>2022JKF02009</award-id>.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study..</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Kemelmacher-Shlizerman</surname></string-name></person-group>, &#x201C;<article-title>Transfiguring portraits</article-title>,&#x201D; <source>ACM Transactions on Graphics (TOG)</source>, vol. <volume>35</volume>, no. <issue>4</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M. R.</given-names> <surname>Koujan</surname></string-name>, <string-name><given-names>M. C.</given-names> <surname>Doukas</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Roussos</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Zafeiriou</surname></string-name></person-group>, &#x201C;<article-title>Head2head: Video-based neural head synthesis</article-title>,&#x201D; in <conf-name>Proc. of the 15th IEEE Int. Conf. on Automatic Face and Gesture Recognition (FG)</conf-name>, <conf-loc>Buenos Aires, Argentina</conf-loc>, pp. <fpage>16</fpage>&#x2013;<lpage>23</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Pumarola</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Agudo</surname></string-name>, <string-name><given-names>A. M.</given-names> <surname>Martinez</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sanfeliu</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Moreno-Noguer</surname></string-name></person-group>, &#x201C;<article-title>Ganimation: Anatomically-aware facial animation from a single image</article-title>,&#x201D; in <conf-name>Proc. of the European Conf. on Computer Vision (ECCV)</conf-name>, <publisher-loc>Munich, Germany</publisher-loc>, pp. <fpage>818</fpage>&#x2013;<lpage>833</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Suwajanakorn</surname></string-name>, <string-name><given-names>S. M.</given-names> <surname>Seitz</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Kemelmacher-Shlizerman</surname></string-name></person-group>, &#x201C;<article-title>Synthesizing obama: Learning lip sync from audio</article-title>,&#x201D; <source>ACM Transactions on Graphics (ToG)</source>, vol. <volume>36</volume>, no. <issue>4</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><collab>MIT TR</collab>, &#x201C;<article-title>Deepfake Putin is here to warn Americans about their self-inflicted doom</article-title>,&#x201D; <year>2020</year>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.technologyreview.com/2020/09/29/1009098/ai-deepfake-putin-kim-jong-un-us-election/">https://www.technologyreview.com/2020/09/29/1009098/ai-deepfake-putin-kim-jong-un-us-election/</ext-link></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><collab>CNN</collab>, &#x201C;<article-title>Deepfake&#x2019; Queen delivers alternative Christmas speech, in warning about misinformation</article-title>,&#x201D; <year>2020</year>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.cnn.com/2020/12/25/uk/deepfake-queen-speech-christmas-intl-gbr">https://www.cnn.com/2020/12/25/uk/deepfake-queen-speech-christmas-intl-gbr</ext-link></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><collab>BuzzFeed</collab>, &#x201C;<article-title>How to spot a deep-fake like the Barack Obama-Jordan Peele video</article-title>,&#x201D; <year>2018</year>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.buzzfeed.com/craigsilverman/obama-jordan-peele-deepfake-video-debunk-buzzfeed">https://www.buzzfeed.com/craigsilverman/obama-jordan-peele-deepfake-video-debunk-buzzfeed</ext-link></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Lyu</surname></string-name></person-group>, &#x201C;<article-title>Exposing deepfake videos by detecting face warping artifacts</article-title>,&#x201D; <comment>arXiv:1811.00656</comment>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Matern</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Riess</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Stamminger</surname></string-name></person-group>, &#x201C;<article-title>Exploiting visual artifacts to expose deepfakes and face manipulations</article-title>,&#x201D; in <conf-name>Proc. of the 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW)</conf-name>, <publisher-loc>Waikoloa Village, HI, USA</publisher-loc>, pp. <fpage>83</fpage>&#x2013;<lpage>92</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Schwarcz</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Chellappa</surname></string-name></person-group>, &#x201C;<article-title>Finding facial forgery artifacts with parts-based detectors</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>933</fpage>&#x2013;<lpage>942</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Masi</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Killekar</surname></string-name>, <string-name><given-names>R. M.</given-names> <surname>Mascarenhas</surname></string-name>, <string-name><given-names>S. P.</given-names> <surname>Gurudatt</surname></string-name> and <string-name><given-names>W.</given-names> <surname>AbdAlmageed</surname></string-name></person-group>, &#x201C;<article-title>Two-branch recurrent network for isolating deepfakes in videos</article-title>,&#x201D; in <conf-name>Proc. of the European Conf. on Computer Vision (ECCV)</conf-name>, <publisher-loc>Glasgow, UK</publisher-loc>, pp. <fpage>667</fpage>&#x2013;<lpage>684</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Local relation learning for face forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the AAAI Conf. on Artificial Intelligence</conf-name>, pp. <fpage>1081</fpage>&#x2013;<lpage>1088</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Liang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Gao</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Fsspotter: Spotting face-swapped video by spatial and temporal clues</article-title>,&#x201D; in <conf-name>Proc. of the 2020 IEEE Int. Conf. on Multimedia and Expo (ICME)</conf-name>, <publisher-loc>London, UK</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Ru</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Bita-Net: Bi-temporal attention network for facial video forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the 2021 IEEE Int. Joint Conf. on Biometrics (IJCB)</conf-name>, <publisher-loc>Shenzhen, China</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Sabir</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jaiswal</surname></string-name>, <string-name><given-names>W.</given-names> <surname>AbdAlmageed</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Masi</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Recurrent convolutional strategies for face manipulation detection in videos</article-title>,&#x201D; <source>Interfaces (GUI)</source>, vol. <volume>3</volume>, no. <issue>1</issue>, pp. <fpage>80</fpage>&#x2013;<lpage>87</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Agarwal</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Farid</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Fried</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Agrawala</surname></string-name></person-group>, &#x201C;<article-title>Detecting deep-fake videos from phoneme-viseme mismatches</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>660</fpage>&#x2013;<lpage>661</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Chugh</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Dhall</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Subramanian</surname></string-name></person-group>, &#x201C;<article-title>Not made for each other-audio-visual dissonance-based deepfake detection and localization</article-title>,&#x201D; in <conf-name>Proc. of the 28th ACM Int. Conf. on Multimedia</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>439</fpage>&#x2013;<lpage>447</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Mittal</surname></string-name>, <string-name><given-names>U.</given-names> <surname>Bhattacharya</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Chandra</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Bera</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Manocha</surname></string-name></person-group>, &#x201C;<article-title>Emotions don&#x2019;t lie: An audio-visual deepfake detection method using affective cues</article-title>,&#x201D; in <conf-name>Proc. of the 28th ACM Int. Conf. on Multimedia</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>2823</fpage>&#x2013;<lpage>2832</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H. H.</given-names> <surname>Nguyen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yamagishi</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Echizen</surname></string-name></person-group>, &#x201C;<article-title>Capsule-forensics: Using capsule networks to detect forged images and videos</article-title>,&#x201D; in <conf-name>Proc. of the 2019 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>, <publisher-loc>Brighton, UK</publisher-loc>, pp. <fpage>2307</fpage>&#x2013;<lpage>2311</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Ju</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zheng</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Locally GAN-generated face detection based on an improved Xception</article-title>,&#x201D; <source>Information Sciences</source>, vol. <volume>572</volume>, no. <issue>11</issue>, pp. <fpage>16</fpage>&#x2013;<lpage>28</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D. A.</given-names> <surname>Coccomini</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Messina</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Gennaro</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Falchi</surname></string-name></person-group>, &#x201C;<article-title>Combining efficientnet and vision transformers for video deepfake detection</article-title>,&#x201D; in <conf-name>Proc. of the Int. Conf. on Image Analysis and Processing</conf-name>, <publisher-loc>Lecce, UK</publisher-loc>, pp. <fpage>219</fpage>&#x2013;<lpage>229</lpage>, <year>2022</year>. </mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Durall</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Keuper</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Keuper</surname></string-name></person-group>, &#x201C;<article-title>Watch your up-convolution: CNN based generative deep neural networks are failing to reproduce spectral distributions</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Seattle, WA</publisher-loc>, pp. <fpage>7890</fpage>&#x2013;<lpage>7899</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Frank</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Eisenhofer</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Sch&#x00F6;nherr</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Fischer</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kolossa</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Leveraging frequency analysis for deep fake image recognition</article-title>,&#x201D; in <conf-name>Proc. of the Int. Conf. on Machine Learning</conf-name>, pp. <fpage>3247</fpage>&#x2013;<lpage>3258</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Mirsky</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Lee</surname></string-name></person-group>, &#x201C;<article-title>The creation and detection of deepfakes: A survey</article-title>,&#x201D; <source>ACM Computing Surveys (CSUR)</source>, vol. <volume>54</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>41</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Tolosana</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Vera-Rodriguez</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Fierrez</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Morales</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Ortega-Garcia</surname></string-name></person-group>, &#x201C;<article-title>Deepfakes and beyond: A survey of face manipulation and fake detection</article-title>,&#x201D; <source>Information Fusion</source>, vol. <volume>64</volume>, no. <issue>1</issue>, pp. <fpage>131</fpage>&#x2013;<lpage>148</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Qi</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Lyu</surname></string-name></person-group>, &#x201C;<article-title>Celeb-DF (v2): A new dataset for DeepFake Forensics</article-title>,&#x201D; <comment>arXiv:1909.12962</comment>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Thies</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zollh&#x00F6;fer</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Nie&#x00DF;ner</surname></string-name></person-group>, &#x201C;<article-title>Deferred neural rendering: Image synthesis using neural textures</article-title>,&#x201D; <source>ACM Transactions on Graphics (TOG)</source>, vol. <volume>38</volume>, no. <issue>4</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Dang</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Stehouwer</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>A. K.</given-names> <surname>Jain</surname></string-name></person-group>, &#x201C;<article-title>On the detection of digital face manipulation</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>5781</fpage>&#x2013;<lpage>5790</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Chai</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Bau</surname></string-name>, <string-name><given-names>S. N.</given-names> <surname>Lim</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Isola</surname></string-name></person-group>, &#x201C;<article-title>What makes fake images detectable? understanding properties that generalize</article-title>,&#x201D; in <conf-name>Proc. of the European Conf. on Computer Vision (ECCV)</conf-name>, <publisher-loc>Glasgow, UK</publisher-loc>, pp. <fpage>103</fpage>&#x2013;<lpage>120</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Multi-attentional deepfake detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>2185</fpage>&#x2013;<lpage>2194</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>V. I.</given-names> <surname>Morariu</surname></string-name> and <string-name><given-names>L. S.</given-names> <surname>Davis</surname></string-name></person-group>, &#x201C;<article-title>Learning rich features for image manipulation detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Salt Lake City, UT, USA</publisher-loc>, pp. <fpage>1053</fpage>&#x2013;<lpage>1061</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yan</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Generalizing face forgery detection with high-frequency features</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>16317</fpage>&#x2013;<lpage>16326</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Fei</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Dai</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Xia</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Learning second order local anomaly for general face forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, pp. <fpage>20270</fpage>&#x2013;<lpage>20280</lpage>, <year>2022</year>. </mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Fridrich</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Kodovsky</surname></string-name></person-group>, &#x201C;<article-title>Rich models for steganalysis of digital images</article-title>,&#x201D; <source>IEEE Transactions on Information Forensics and Security</source>, vol. <volume>7</volume>, no. <issue>3</issue>, pp. <fpage>868</fpage>&#x2013;<lpage>882</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Cozzolino</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Poggi</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Verdoliva</surname></string-name></person-group>, &#x201C;<article-title>Recasting residual-based local descriptors as convolutional neural networks: An application to image forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the 5th ACM Workshop on Information Hiding and Multimedia Security</conf-name>, <publisher-loc>Philadelphia, PA, USA</publisher-loc>, pp. <fpage>159</fpage>&#x2013;<lpage>164</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Bao</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Chen</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Face x-ray for more general face forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>5001</fpage>&#x2013;<lpage>5010</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Chollet</surname></string-name></person-group>, &#x201C;<article-title>Xception: Deep learning with depthwise separable convolutions</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Honolulu, HI, USA</publisher-loc>, pp. <fpage>1251</fpage>&#x2013;<lpage>1258</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Tan</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Le</surname></string-name></person-group>, &#x201C;<article-title>Efficientnet: Rethinking model scaling for convolutional neural networks</article-title>,&#x201D; in <conf-name>Proc. of the Int. Conf. on Machine Learning</conf-name>, <publisher-loc>Long Beach, California, USA</publisher-loc>, pp. <fpage>6105</fpage>&#x2013;<lpage>6114</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Rahmouni</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Nozick</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yamagishi</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Echizen</surname></string-name></person-group>, &#x201C;<article-title>Distinguishing computer graphics from natural images using convolution neural networks</article-title>,&#x201D; in <conf-name>Proc. of the 2017 IEEE Int. Workshop on Information Forensics and Security (WIFS)</conf-name>, <publisher-loc>Rennes, France</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Afchar</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Nozick</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yamagishi</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Echizen</surname></string-name></person-group>, &#x201C;<article-title>Mesonet: A compact facial video forgery detection network</article-title>,&#x201D; in <conf-name>Proc. of the 2018 IEEE Int. Workshop on Information Forensics and Security (WIFS)</conf-name>, <publisher-loc>Hong Kong, China</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Bayar</surname></string-name> and <string-name><given-names>M. C.</given-names> <surname>Stamm</surname></string-name></person-group>, &#x201C;<article-title>A deep learning approach to universal image manipulation detection using a new convolutional layer</article-title>,&#x201D; in <conf-name>Proc. of the 4th ACM Workshop on Information Hiding and Multimedia Security</conf-name>, <publisher-loc>Vigo, Galicia, Spain</publisher-loc>, pp. <fpage>5</fpage>&#x2013;<lpage>10</lpage>, <year>2016</year>. </mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. M. S.</given-names> <surname>Mahdy</surname></string-name></person-group>, &#x201C;<article-title>A numerical method for solving the nonlinear equations of Emden-Fowler models</article-title>,&#x201D; <source>Journal of Ocean Engineering and Science</source>, vol. <volume>88</volume>, no. <issue>16</issue>, pp. <fpage>3406</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. M. S.</given-names> <surname>Mahdy</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Lotfy</surname></string-name> and <string-name><given-names>A. A.</given-names> <surname>El-Bary</surname></string-name></person-group>, &#x201C;<article-title>Use of optimal control in studying the dynamical behaviors of fractional financial awareness models</article-title>,&#x201D; <source>Soft Computing</source>, vol. <volume>26</volume>, no. <issue>7</issue>, pp. <fpage>3401</fpage>&#x2013;<lpage>3409</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. M. S.</given-names> <surname>Mahdy</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Higazy</surname></string-name></person-group>, &#x201C;<article-title>Numerical different methods for solving the nonlinear biochemical reaction model</article-title>,&#x201D; <source>International Journal of Applied and Computational Mathematics</source>, vol. <volume>5</volume>, no. <issue>6</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>17</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. M. S.</given-names> <surname>Mahdy</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Lotfy</surname></string-name>, <string-name><given-names>A.</given-names> <surname>El-Bary</surname></string-name> and <string-name><given-names>I. M.</given-names> <surname>Tayel</surname></string-name></person-group>, &#x201C;<article-title>Variable thermal conductivity and hyperbolic two-temperature theory during magneto-photothermal theory of semiconductor induced by laser pulses</article-title>,&#x201D; <source>The European Physical Journal Plus</source>, vol. <volume>136</volume>, no. <issue>6</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>21</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Durall</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Keuper</surname></string-name>, <string-name><given-names>F. J.</given-names> <surname>Pfreundt</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Keuper</surname></string-name></person-group>, &#x201C;<article-title>Unmasking deepfakes with simple features</article-title>,&#x201D; <comment>arXiv:1911.00686</comment>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Karaman</surname></string-name> and <string-name><given-names>S. F.</given-names> <surname>Chang</surname></string-name></person-group>, &#x201C;<article-title>Detecting and simulating artifacts in gan fake images</article-title>,&#x201D; in <conf-name>Proc. of the 2019 IEEE Int. Workshop on Information Forensics and Security (WIFS)</conf-name>, <publisher-loc>Delft, The Netherlands</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Frank</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Eisenhofer</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Sch&#x00F6;nherr</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Fischer</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kolossa</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Leveraging frequency analysis for deep fake image recognition</article-title>,&#x201D; in <conf-name>Proc. of the Int. Conf. on Machine Learning</conf-name>, pp. <fpage>3247</fpage>&#x2013;<lpage>3258</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Qian</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Yin</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Sheng</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Shao</surname></string-name></person-group>, &#x201C;<article-title>Thinking in frequency: Face forgery detection by mining frequency-aware clues</article-title>,&#x201D; in <conf-name>Proc. of the European Conf. on Computer Vision (ECCV)</conf-name>, <publisher-loc>Glasgow, UK</publisher-loc>, pp. <fpage>86</fpage>&#x2013;<lpage>103</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Lu</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Fake generated painting detection via frequency analysis</article-title>,&#x201D; in <conf-name>Proc. of the 2020 IEEE Int. Conf. on Image Processing (ICIP)</conf-name>, <publisher-loc>Abu Dhabi, United Arab Emirates</publisher-loc>, pp. <fpage>1256</fpage>&#x2013;<lpage>1260</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Thies</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zollhofer</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Stamminger</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Theobalt</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Niessner</surname></string-name></person-group>, &#x201C;<article-title>Face2face: Real-time face capture and reenactment of rgb videos</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, pp. <fpage>2387</fpage>&#x2013;<lpage>2395</lpage>, <year>2016</year>. </mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>Y. T.</given-names> <surname>Gao</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Xiao</surname></string-name></person-group>, &#x201C;<article-title>Sstnet: Detecting manipulated faces through spatial, steganalysis and temporal features</article-title>,&#x201D; in <conf-name>Proc. of the 2020 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>, <publisher-loc>Barcelona, Spain</publisher-loc>, pp. <fpage>2952</fpage>&#x2013;<lpage>2956</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>He</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>772</fpage>&#x2013;<lpage>781</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Frequency-aware discriminative feature learning supervised by single-center loss for face forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>6458</fpage>&#x2013;<lpage>6467</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Jin</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Cui</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Multimodal contrastive classification by locally correlated representations for effective face forgery detection</article-title>,&#x201D; <source>Knowledge-Based Systems</source>, vol. <volume>250</volume>, no. <issue>13</issue>, pp. <fpage>109088</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Rossler</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Cozzolino</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Verdoliva</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Riess</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Thies</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Faceforensics&#x002B;&#x002B;: Learning to detect manipulated facial images</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Int. Conf. on Computer Vision</conf-name>, <publisher-loc>Seoul, South Korea</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>11</lpage>, <year>2019</year>. </mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Bao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Wen</surname></string-name></person-group>, &#x201C;<article-title>Faceshifter: Towards high fidelity and occlusion aware face swapping</article-title>,&#x201D; <comment>arXiv:1912.13457</comment>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="other"><collab>Deepfakedetection</collab>, <comment>Accessed: 2020-05-10</comment>. <italic>Available:</italic> <ext-link ext-link-type="uri" xlink:href="https://ai.googleblog.com/2019/09/contributing-data-to-deepfakedetection.html">https://ai.googleblog.com/2019/09/contributing-data-to-deepfakedetection.html</ext-link>. <year>2020</year>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Qian</surname></string-name> and <string-name><given-names>C. C.</given-names> <surname>Loy</surname></string-name></person-group>, &#x201C;<article-title>Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>2889</fpage>&#x2013;<lpage>2898</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Zi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Ma</surname></string-name> and <string-name><given-names>Y. G.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>Wilddeepfake: A challenging real-world dataset for deepfake detection</article-title>,&#x201D; in <conf-name>Proc. of the 28th ACM Int. Conf. on Multimedia</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>2382</fpage>&#x2013;<lpage>2390</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="other"><collab>Faceswap</collab>, <comment>Accessed: 2020-05-10</comment>. <italic>Available:</italic> <ext-link ext-link-type="uri" xlink:href="https://github.com/MarekKowalski/FaceSwap">https://github.com/MarekKowalski/FaceSwap</ext-link>. <year>2020</year>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Ververas</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Kotsia</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Zafeiriou</surname></string-name></person-group>, &#x201C;<article-title>Retinaface: Single-shot multi-level face localisation in the wild</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, pp. <fpage>5203</fpage>&#x2013;<lpage>5212</lpage>, <year>2020</year>. </mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Y. Z.</given-names> <surname>Song</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Hospedales</surname></string-name></person-group>, &#x201C;<article-title>Learning to generalize: Meta-learning for domain generalization</article-title>,&#x201D; in <conf-name>Proc. of the AAAI Conf. on Artificial Intelligence</conf-name>, <conf-loc>Phoenix, Arizona, USA</conf-loc>, <volume>32</volume>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Domain general face forgery detection by learning to weight</article-title>,&#x201D; in <conf-name>Proc. of the AAAI Conf. on Artificial Intelligence</conf-name>, pp. <fpage>2638</fpage>&#x2013;<lpage>2646</lpage>, <year>2021</year>. </mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Local relation learning for face forgery detection</article-title>,&#x201D; in <conf-name>Proc. of the AAAI Conf. on Artificial Intelligence</conf-name>, pp. <fpage>1081</fpage>&#x2013;<lpage>1088</lpage>, <year>2021</year>. </mixed-citation></ref>
</ref-list>
</back>
</article>