<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">JAI</journal-id>
<journal-id journal-id-type="nlm-ta">JAI</journal-id>
<journal-id journal-id-type="publisher-id">JAI</journal-id>
<journal-title-group>
<journal-title>Journal on Artificial Intelligence</journal-title>
</journal-title-group>
<issn pub-type="epub">2579-003X</issn>
<issn pub-type="ppub">2579-0021</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">78014</article-id>
<article-id pub-id-type="doi">10.32604/jai.2026.078014</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Frequency-Aware Robustness Analysis of Deepfake Detection Models</article-title>
<alt-title alt-title-type="left-running-head">Frequency-Aware Robustness Analysis of Deepfake Detection Models</alt-title>
<alt-title alt-title-type="right-running-head">Frequency-Aware Robustness Analysis of Deepfake Detection Models</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Xu</surname><given-names>Haoyang</given-names></name><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>xuhaoyang_zz@163.com</email></contrib>
<aff id="aff-1">
<institution>School of Computer Science and Engineering, University of New South Wales (UNSW)</institution>, <addr-line>Sydney, NSW</addr-line>, <country>Australia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Haoyang Xu. Email: <email>xuhaoyang_zz@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>10</day><month>3</month><year>2026</year>
</pub-date>
<volume>8</volume>
<issue>1</issue>
<fpage>153</fpage>
<lpage>167</lpage>
<history>
<date date-type="received">
<day>22</day>
<month>12</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>02</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Author. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Author</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_JAI_78014.pdf"></self-uri>
<abstract>
<p>This paper conducted a comprehensive study on the robustness of three widely used DFD deep learning models&#x2014;namely, ResNet50, FreqNet, and Xception v1&#x2014;to controlled perturbation attacks and frequency masking across a range of 12 different distortions. The study was performed on 254,166 ForenSynth test images, characterizing the distribution of FSI-drop values derived from over 3.05 million paired predictions. The distribution of FSI-drop values is sharply peaked around zero: 99.7% of the samples exhibit |&#x0394;| &#x003C; 0.1, and the maximum |&#x0394;| &#x2248; 1.5 &#x00D7; 10<sup>&#x2212;3</sup>, indicating high baseline stability. In terms of perturbation-wise comparison, Gaussian blur dominates, yielding a mean degradation 30 times greater than that induced by JPEG compression and twice that caused by rescaling. The frequency masking curves further illustrate unique sensitivities: FreqNet shows high-frequency dependence and rapid decay (i.e., &#x003E;2 to &#x003E;16 masking scale), Xception exhibits moderate attenuation, whereas ResNet50 remains statistically unchanged (median |&#x0394;| &#x003C; 10<sup>&#x2212;4</sup>). All of these differences are statistically significant at the model level (<italic>p</italic> &#x003C; 0.001). The experiments offer a concrete demonstration that CNNs effectively retain prediction invariance to small amounts of image deformation, but the frequency-sensitive design demonstrates readily interpretable high-frequency sensitivity, thereby providing a principled framework for designing detectors robust to image perturbations. It should be noted that FSI-drop measures score stability rather than absolute performance; &#x00A7;4.6 discusses the complementary AUC curves and leaves combined-distortion or adversarial stress tests to future work.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Deepfake detection</kwd>
<kwd>model robustness</kwd>
<kwd>frequency-aware learning</kwd>
<kwd>high-frequency masking</kwd>
<kwd>Gaussian blur</kwd>
<kwd>drop-rate analysis</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Recent proliferation of generative adversarial networks (GANs), e.g., ProGAN [<xref ref-type="bibr" rid="ref-1">1</xref>], StyleGAN [<xref ref-type="bibr" rid="ref-2">2</xref>], and diffusion models [<xref ref-type="bibr" rid="ref-3">3</xref>], has made the generation of hyper-realistic fake faces more accessible. On the bright side, the creative sector (e.g., image creation and retouching, video synthesis) has been rendered more efficient; on the dark side, malicious use of deepfakes&#x2014;from political misinformation [<xref ref-type="bibr" rid="ref-4">4</xref>] and blackmail or financial theft [<xref ref-type="bibr" rid="ref-5">5</xref>]&#x2014;has decreased public trust. Consequently, the research community has proposed many detection architectures, broadly divided into (i) spatial CNNs such as ResNet50 [<xref ref-type="bibr" rid="ref-6">6</xref>], EfficientNet [<xref ref-type="bibr" rid="ref-7">7</xref>], Xception [<xref ref-type="bibr" rid="ref-8">8</xref>], and its lightweight Xception variant [<xref ref-type="bibr" rid="ref-9">9</xref>]; (ii) frequency-aware models: FreqNet-MM22 [<xref ref-type="bibr" rid="ref-10">10</xref>], LNP [<xref ref-type="bibr" rid="ref-11">11</xref>], SPSL [<xref ref-type="bibr" rid="ref-12">12</xref>]; and (iii) biometric-based methods exploiting eye-gaze or heart-rate inconsistency [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<p>Despite impressive accuracy on curated benchmarks (AUC &#x003E; 0.98), recent work by Yu et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] demonstrated that 18 state-of-the-art detectors lose on average 18.7% AUC when evaluated on JPEG-compressed social-media data. These observations align with adversarial-robustness studies in generic computer vision [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>], yet a granular comparative analysis targeting lightweight and frequency-based detectors remains absent. In order to fill this gap, we perform the most extensive study of the largest perturbations to date over three representative deepfake detectors: the spatially heavy baseline ResNet50; a mobile-friendly detector XceptionMinimal; and frequency-domain feature detector FreqNet-MM22 [<xref ref-type="bibr" rid="ref-10">10</xref>], which allows the study of robustness pertaining to explicit frequency domain features. We test using the ForenSynths dataset [<xref ref-type="bibr" rid="ref-18">18</xref>], which has 254,166 matching real and synthetic image pairs, using 12 standard perturbation settings. Altogether, we obtain 3.05 million FSI-drop measurements. We emphasize that the main contribution of this work is a systematic robustness audit and diagnostic analysis of representative deepfake detection architectures, rather than proposing a new detection method. Key outcomes from the robustness analysis include:
<list list-type="bullet">
<list-item>
<p>ResNet50 exhibits the highest robustness, achieving a median FSI-drop of 6.54 &#x00D7; 10<sup>&#x2212;8</sup>, while 99.7% of the samples stay below 0.1, which is by one order of magnitude more than XceptionMinimal and FreqNet-MM22.</p></list-item>
<list-item>
<p>XceptionMinimal performs an intermediate robustness result, as the degradation values of this method are higher than the ResNet50 ones and lower than those of FreqNet-MM22 in almost all the considered perturbation conditions, indicating that only a moderate spatial representation is learned.</p></list-item>
<list-item>
<p>FreqNet-MM22 is shown to exhibit an inverted robustness response to Gaussian blur, where the FSI-drop decreases from 0.35 &#x00D7; 10<sup>&#x2212;3</sup> to 0.15 &#x00D7; 10<sup>&#x2212;3</sup> as the blur radius &#x03C3; increases from 1 to 4. This behavior confirms its high-frequency inductive bias while revealing a previously unreported vulnerability pattern. Other frequency-aware methods were not evaluated and may behave differently.</p></list-item>
</list></p>
<p>Taken altogether, these results show that architectural inductive biases (specifically frequency-based vs. spatial representations) bear non-trivial and perturbation-dependent robustness trade-offs. Accordingly, system-wide robustness evaluation should be considered as a mandatory part of pre-deployment analyses for deepfake detection applications.</p>
<p><bold><italic>Limitations &#x0026; Scope</italic></bold></p>
<p>This work focuses on cheap, single-type distortions that dominate social-media pipelines (JPEG, blur, resize). We intentionally leave out cascaded impairments/adversarial perturbations because (a) their parameter space grows exponentially and (b) a preliminary exploration using 2000 randomly sampled images under mild cascaded (compress &#x002B; blur) distortions did not alter the robustness ranking (ResNet50 &#x003E; XceptionMinimal &#x003E; FreqNet). While the observed robustness patterns are consistent across the three representative detectors tested on the ForenSynths dataset, their applicability to other architectures or datasets has yet to be confirmed. A comprehensive analysis of combined and adversarial perturbations is reserved for future work.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Deepfake Generation</title>
<p>Early faceswap autoencoders [<xref ref-type="bibr" rid="ref-19">19</xref>] evolved into progressively-growing GANs&#x2014;ProGAN [<xref ref-type="bibr" rid="ref-1">1</xref>], StyleGAN [<xref ref-type="bibr" rid="ref-2">2</xref>], StyleGAN2 [<xref ref-type="bibr" rid="ref-20">20</xref>], StyleGAN3 [<xref ref-type="bibr" rid="ref-21">21</xref>]&#x2014;and, very recently, diffusion probabilistic models that surpass GANs in fidelity [<xref ref-type="bibr" rid="ref-3">3</xref>]. Higher generator quality directly elevates the detection challenge.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Deepfake Detection</title>
<p>Spatial CNNs remain the dominant paradigm. Afchar et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] first transplanted MesoNet from steganalysis; Rossler et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] popularised XceptionNet and its minimal variant for FaceForensics&#x002B;&#x002B;; Tan &#x0026; Le [<xref ref-type="bibr" rid="ref-7">7</xref>] later scaled EfficientNet-B4 to balance accuracy and efficiency.</p>
<p>To escape the &#x201C;pixel-trap&#x201D;, researchers have incorporated biological cues: eye-gaze variance (FakeCatcher [<xref ref-type="bibr" rid="ref-13">13</xref>]), heart-rate inconsistency [<xref ref-type="bibr" rid="ref-14">14</xref>], and PRNU sensor noise [<xref ref-type="bibr" rid="ref-22">22</xref>]. These methods, however, require high-resolution uncompressed inputs and thus remain fragile in social-media pipelines.</p>
<p>Frequency-aware approaches. Among existing methods, FreqNet-MM22 proposed by Tan et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] is the most closely related to this study. The model incorporates a learnable Frequency Transformation Module (FTM) that adaptively reweights DCT bands prior to late fusion with an RGB branch. We underline that the evaluation is limited to FreqNet-MM22; alternative frequency-aware architectures could yield different outcomes, so any broader claims about frequency-based detectors warrant careful qualification.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Robustness Evaluation</title>
<p>Sabir et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] benchmarked Xception against JPEG compression and resizing, while G&#x00FC;era &#x0026; Delp [<xref ref-type="bibr" rid="ref-23">23</xref>] analysed recapture artefacts; both, however, focused on a single distortion type. Recent work has shifted to multi-perturbation protocols: Haliassos et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] introduced a self-supervised approach for robust detection of audio-visual forgeries, evaluating it on real-world video datasets and showing that detection performance can degrade under various perturbations. Tan et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] showed that many frequency-based detectors tend to overfit to specific artifacts in the frequency domain, limiting generalization to unseen deepfake sources. Dutta et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed a wavelet sub-band based frequency detection approach that decomposes Fourier representations into wavelet energies to improve robustness beyond spatial CNN features. Finally, ForenSynths [<xref ref-type="bibr" rid="ref-18">18</xref>] released only clean top-line results for ResNet50 and Xception, leaving granular robustness statistics unavailable.</p>
<p>To the best of current knowledge, this study is the first to deliver a fine-grained, drop-based robustness profile for FreqNet-MM22 [<xref ref-type="bibr" rid="ref-10">10</xref>] under systematic JPEG, blur, and resize distortions, side-by-side with ResNet50 and XceptionMinimal.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>The robustness of three architectures&#x2014;ResNet50, XceptionMinimal, and FreqNet-MM22&#x2014;was evaluated under realistic image degradations. All experiments share the same data pipeline, perturbation bank, and evaluation protocol to ensure fair comparison.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Models under Test</title>
<p>XceptionMinimal [<xref ref-type="bibr" rid="ref-9">9</xref>]&#x2014;A width-reduced version (&#x00D7;0.75 channels) of the Xception architecture. Theoretically, Xception has a formulation using depthwise separable convolutions where spatial correlation learning (depthwise convolution) and cross-channel correlation learning (pointwise convolution) are separated for lower model complexity without loss of representation power. Our model is not pretrained on ImageNet but instead trained from scratch using the FaceForensics&#x002B;&#x002B; protocol&#x2014;enabling deployment on mobile and resource-constrained devices&#x2014;yet detecting subtle manipulation artefacts at a fine-grained level.</p>
<p>FreqNet-MM22 [<xref ref-type="bibr" rid="ref-10">10</xref>]&#x2014;Dual branch network that is inspired by the complementary behavior of spatial domain and frequency-domain representation for image forensics, where the RGB branch represents semantic and texture cues in the spatial domain, while explicit encoding of spectral anomalies due to manipulation exists in the frequency branch. Its Frequency Transformation Module (FTM), which learns an 8 &#x00D7; 8 DCT-based weighting that allows it to focus more attention on informative frequency bands, as opposed to using a pre-defined transform. It is initialized from the official MM&#x2019;22 checkpoint before fine-tuning on ForenSynths to align its frequency response with the target manipulation distribution. The model is initialized with the publicly available pre-trained weights from Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Learning [<xref ref-type="bibr" rid="ref-10">10</xref>] and subsequently fine-tuned on ForenSynths [<xref ref-type="bibr" rid="ref-18">18</xref>] to adapt the learned frequency representations to the target manipulation distribution.</p>
<p>ResNet50 [<xref ref-type="bibr" rid="ref-6">6</xref>]&#x2014;A deep residual network with the idea of residual learning; Identity shortcut connections allow a Residual Network to be trained to discover residual mapping instead of underlying mapping, which relieves the problem of Vanishing Gradient and helps to optimize a deeper model, leading to good generalization. ResNet50 (ImageNet Pre-training): A popular pre-trained model from ImageNet that has proven to be effective at capturing spatial information; we use it by replacing its last fully-connected layer with our own 2-node classifier and finetuning the whole network in an end-to-end manner to solve the binomial problem of detecting deepfakes.</p>
<p>These architectures are selected to represent three complementary design paradigms in deepfake detection: lightweight separable convolutional models, explicit frequency-aware modeling, and deep residual learning.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Dataset and Pre-Processing</title>
<p>ForenSynths [<xref ref-type="bibr" rid="ref-18">18</xref>] provides 254,166 balanced real/fake 512 &#x00D7; 512 face images synthesized by ProGAN, StyleGAN2, and diffusion. The dataset was split at the identity level using an 80/10/10 ratio, resulting in 203,332 training, 25,416 validation, and 25,418 test samples. All images are centre-cropped to 224 &#x00D7; 224, converted to RGB, and normalized with ImageNet statistics. No further augmentation is applied to isolate the impact of the controlled perturbations.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Perturbation Suite</title>
<p>Three universally encountered distortions are applied at inference time:
<list list-type="bullet">
<list-item>
<p>JPEG compression&#x2014;quality factor Q &#x2208; {20, 30, 40, 50, 60, 70, 80, 90, 100}.</p></list-item>
<list-item>
<p>Gaussian blur&#x2014;isotropic kernels with &#x03C3; &#x2208; {1, 2, 4}.</p></list-item>
<list-item>
<p>Resize&#x2014;bicubic down-scaling followed by up-scaling to 224 &#x00D7; 224, scale &#x2208; {0.4, 0.5, 0.6, 0.8, 1.0}.</p></list-item>
</list></p>
<p>Each image is degraded on-the-fly with a single distortion, yielding 36 perturbed copies (12 parameter levels &#x00D7; 3 distortion types) per sample. These parameter ranges were selected to cover the spectrum of distortion intensities commonly encountered in social media pipelines and practical image processing workflows.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>FSI-Drop Metric</title>
<p>Let <italic>x</italic> be a clean image and <italic>x<sub>i</sub></italic> its perturbed copy (e.g., JPEG, blur, resize). Write <italic>p</italic>(&#x00B7;) for the detector&#x2019;s fake probability. The Fake-Score-Impact (FSI) drop is defined as:<disp-formula id="ueqn-1"><mml:math id="mml-ueqn-1" display="block"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <italic>x</italic><sub><italic>i</italic></sub> is the perturbed variant and logit(<italic>p</italic>) &#x003D; ln(<italic>p</italic>/(1 &#x2212; <italic>p</italic>)). Here, <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is computed by comparing the detector&#x2019;s outputs on the clean image <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>x</mml:mi></mml:math></inline-formula> and its perturbed version <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. &#x0394;(<italic>x</italic>) is invariant to sigmoid calibration and approximately additive for small perturbations. Over a dataset, a median &#x0394; near zero signals robustness; large positive or negative values flag score instability. The &#x201C;original minus perturbed&#x201D; order keeps positive &#x0394; aligned with a drop in fake-likelihood after corruption.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Training &#x0026; Evaluation Protocol</title>
<p>All models are trained under an identical end-to-end protocol on the clean ForenSynths training split to ensure fair comparison. During robustness evaluation, each trained detector is exposed exclusively to inference-time perturbations, while the training distribution remains unchanged.</p>
<p>For each combination of model architecture, perturbation type, and parameter level, the entire test set is processed to obtain paired predictions on clean images and their corresponding perturbed counterparts. The Fake-Score-Impact (FSI) drop is computed on a per-image basis and serves as the fundamental input to all subsequent analyses.</p>
<p>To characterize robustness at both global and fine-grained levels, the FSI-drop distribution is summarized using its mean, standard deviation, median, extrema, and upper percentiles. Perturbation&#x2013;response curves are further constructed with 95% bootstrap confidence intervals to quantify uncertainty. In addition, Grad-CAM visualizations are applied to a subset of samples to qualitatively associate attention shifts with score instability.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Reproducibility</title>
<p>The pipeline is implemented in PyTorch 1.13 and will be released with: (i) perturbed test-set CSV logs, (ii) FSI-drop computation script, (iii) model weights, and (iv) three plots: &#x201C;JPEG Quality vs. Drop&#x201D;, &#x201C;Gaussian Blur &#x03C3; vs. Drop&#x201D;, &#x201C;Resize Scale vs. Drop&#x201D;. All experiments were run on a single RTX-3090 (24 GB); total GPU time &#x2248; 40 h.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Discussion</title>
<p>A comprehensive robustness audit was conducted for ResNet50, XceptionMinimal, and FreqNet-MM22 on the ForenSynths dataset. The analysis leveraged all 3.05 million FSI-drop measurements obtained by applying twelve parameter-level perturbation combinations to the 254,166 test images, encompassing JPEG compression, Gaussian blur, and resize scaling. This exhaustive evaluation enables a detailed examination of how the architectures respond at the <bold>score level</bold> under a wide spectrum of common distortions and provides both global and per-sample insights into model stability.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Overall Robustness: Global FSI-Drop Distribution</title>
<p>Across all models and perturbation conditions, the pooled FSI-drop distribution is tightly concentrated around zero, indicating that detector logits remain largely stable under the considered single-factor distortions. At a global scale, none of the evaluated architectures exhibits systematic score degradation in response to typical image perturbations.</p>
<p>Quantitatively, the aggregated distribution over 3.05 million FSI-drop measurements&#x2014;covering all test images, perturbation types, and parameter levels&#x2014;has a mean close to zero, a median exactly equal to zero, and very low dispersion. More than 99% of the measurements fall within a narrow logit interval, with extreme values accounting for &#x003C;0.01% of the dynamic range.</p>
<p>These observations establish a shared baseline of score-level robustness across the three detectors and validate FSI-drop as a sensitive yet non-degenerate metric: while capable of capturing instability when present, it does not spuriously amplify minor prediction fluctuations. Differences observed in later sections therefore reflect <bold>architecture-specific response patterns</bold> rather than global fragility.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Perturbation-Specific Analysis</title>
<p>To identify which distortion family primarily accounts for the variance observed in <xref ref-type="sec" rid="s4_1">Section 4.1</xref>, the 3.05 million FSI-drop records were decomposed by perturbation type. <xref ref-type="table" rid="table-1">Table 1</xref> presents the per-category statistics (units &#x00D7;10<sup>&#x2212;3</sup>).</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>FSI-drop statistics by distortion family (N &#x003D; 254 166).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Category</th>
<th>Samples</th>
<th>Mean &#x00B1; Std</th>
<th>Min</th>
<th>Max</th>
</tr>
</thead>
<tbody>
<tr>
<td>Blur</td>
<td>69,318</td>
<td>0.156 &#x00B1; 0.152</td>
<td>&#x2212;0.06</td>
<td>1.43</td>
</tr>
<tr>
<td>JPEG</td>
<td>115,530</td>
<td>0.005 &#x00B1; 0.023</td>
<td>&#x2212;0.12</td>
<td>0.3</td>
</tr>
<tr>
<td>Resize</td>
<td>69,318</td>
<td>0.077 &#x00B1; 0.074</td>
<td>0</td>
<td>0.72</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>With respect to the considered perturbations, Gaussian blur induces the greatest mean drop, and, with respect to the given explanations, the highest extreme value of 1.43 observed for it. The mean FSI-drop for Gaussian blur exceeds that for JPEG compression by a factor of 30, and in any case is much higher than for any other perturbation. This section describes which perturbations generate the largest score-level variations, without inferring task-level implications.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Scale-Resolved Blur Curves</title>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> zooms into the blur radius sweep &#x03C3; &#x003D; 1 &#x2192; 4. ResNet50 remains flat at 0.02&#x2013;0.03, confirming that its spatial residual path does not over-commit to sharp edges. XceptionMinimal rises mildly from 0.04 to 0.06, a slope consistent with a lightweight network nevertheless encodes limited high-frequency information. FreqNet-MM22 behaves oppositely: drop falls from 0.35 (&#x03C3; &#x003D; 1) to 0.15 (&#x03C3; &#x003D; 4).</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>FSI drop rate comparison: FreqNet vs. ResNet50 vs. Xception.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-1.tif"/>
</fig>
<p>The negative gradient (&#x2212;0.067 per radius unit, R<sup>2</sup> &#x003D; 0.97) is a direct consequence of the learnable Frequency Transformation Module: as blur increases, the upper-octave energy is removed, the training vs. test distribution gap shrinks, and the score penalty diminishes. This &#x201C;robustness inversion&#x201D; is not a virtue but a signature of over-reliance on high-frequency content.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Scale-Resolved Perturbation Curves&#x2014;Blur, Resize &#x0026; JPEG</title>
<p><xref ref-type="fig" rid="fig-2">Figs. 2</xref>&#x2013;<xref ref-type="fig" rid="fig-4">4</xref> characterize the FSI-drop (&#x00D7;10<sup>&#x2212;3</sup>) of three architectures&#x2014;XceptionMinimal (<xref ref-type="fig" rid="fig-2">Fig. 2</xref>), FreqNet (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>), and ResNet (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>)&#x2014;across three common image distortion families (JPEG compression, Gaussian blur, resize scaling). For each panel, the abscissa denotes the distortion parameter (JPEG quality: 20&#x2192;100; Gaussian blur radius &#x03C3;: 1&#x2192;4; resize scale: 0.4&#x2192;0.8), and error bands represent 95% bootstrap confidence intervals (CIs), quantifying uncertainty in the drop metric.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Impact of image perturbations on Xception output drop.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Impact of image perturbations on freqnet output drop.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-3.tif"/>
</fig><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Impact of image perturbations on resnet output drop.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-4.tif"/>
</fig>
<p>XceptionMinimal is relatively insensitive to the applied perturbations. In JPEG compression, the range of FSI drop as quality improves is as small as &#x00B1;0.01, is very shallow (&#x003C;0.001 per quality unit), and with a very small R<sup>2</sup> (&#x003C;0.30), so it is clearly not monotonically trended. For Gaussian blur, there is a relatively small but steady increase in FSI drop, increasing from about 0.04 (&#x03C3; &#x003D; 1) to about 0.06 (&#x03C3; &#x003D; 4) with a slope of &#x002B;0.013 per &#x03C3; unit (R<sup>2</sup> &#x003D; 0.94), and 95% confidence intervals are sufficiently narrow (width &#x003C;0.01), supporting the statistical validity of the observed trend. The negative monotonic trend on the top right comes from Resize scaling, that is, more extreme downsampling (smaller scales) results in greater declines, going from &#x007E;0.05 (scale &#x003D; 0.4) to close to 0 (scale &#x003D; 0.8), at a slope of &#x2212;0.038 per 0.1 unit of scale, suggesting moderate sensitivity to spatial resolution degradation.</p>
<p>The results with FreqNet are clearly different from the other nets. Its response to JPEG compression has very little FSI drop (range: &#x2212;0.006 &#x2192; &#x002B;0.008; slope &#x2248; 0; R<sup>2</sup> low), as do the other nets. The response to Gaussian blur, on the other hand, is of the opposite sign: FSI drop decreases from about 0.35 (&#x03C3; &#x003D; 1) to about 0.15 (&#x03C3; &#x003D; 4) with a large negative slope of &#x2212;0.067 per unit of &#x03C3; (R<sup>2</sup> &#x003D; 0.97). By &#x03C3; &#x003D; 2, the 95% CIs no longer overlap those of ResNet, and the <italic>p</italic>-value for the two-sample <italic>t</italic>-tests was <italic>p</italic> &#x003C; 0.001 at each integer &#x03C3;, corroborating the statistical dissimilarity of FreqNet&#x2019;s response. The inversion supports FreqNet&#x2019;s reliance on high-frequency cues through its FM: blur attenuates these cues and shrinks the training-test distribution gap, lowering the logit penalty. For resize scaling, FreqNet is the most susceptible: FSI drop degrades from &#x007E;0.12 (scale &#x003D; 0.4) to almost 0 (scale &#x003D; 0.8), sloping at &#x2212;0.075 per 0.1 scale unit-double that of XceptionMinimal.</p>
<p>The ResNet shows the most robust behavior across all distortions, consistent with expectations for its canonical architecture and strong generalizability. Under JPEG compression, performance changes are minimal (&#x2212;0.005 to &#x002B;0.007), with a slope below 0.001. Gaussian blur produces a nearly constant FSI drop (0.02&#x2013;0.03, slope &#x002B;0.002, R<sup>2</sup> &#x003D; 0.88), reflecting resilience to high-frequency information loss. Resize scaling follows the negative monotonic trend observed in other models, but with the narrowest absolute spread: drop decreases from &#x007E;0.02 (scale &#x003D; 0.4) to 0 (scale &#x003D; 0.8), with a slope of &#x2212;0.018 per 0.1 scale unit&#x2014;approximately six times milder than FreqNet.</p>
<p>Across models, JPEG compression exerts only a minor effect on FSI-drop, rendering it neither discriminative nor informative for architecture identification. Gaussian blur emerges as the most information-rich stress test: it unambiguously separates the architectures by both direction&#x2014;rising (FreqNet) vs. stable (XceptionMinimal &#x0026; ResNet)&#x2014;and magnitude, with FSI-drop magnitudes ranked FreqNet &#x003E; XceptionMinimal &#x003E; ResNet. Resize scaling provides an additional robustness indicator, quantifying the spatial-feature generalisation gap: ResNet remains most robust, XceptionMinimal exhibits medium sensitivity, and FreqNet suffers the largest performance drop under downsampling.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Per-Sample Frequency-Masking Scatter&#x2014;Image-Level Robustness</title>
<p>To investigate image-to-image variance under band-limited distortion, zero masks were successively applied to FFT octaves (scales 2, 4, 8, 16), and the FSI-drop was recorded for each test face (200 k samples). Each subplot in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> represents a scatter cloud, where the horizontal axis corresponds to sample index (0&#x2192;200,000), the vertical axis denotes the FSI-drop, and color intensity encodes local point density.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Xception model sensitivity to varying frequency masking scales.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-5.tif"/>
</fig>
<p>The results for XceptionMinimal (<xref ref-type="fig" rid="fig-5">Fig. 5</xref>) show the median drop decreasing slowly from 0.60 (scale 2) to 0.35 (scale 16), the IQR shrinking from 0.08% to 0.04%, and 95% of points fluctuating by less than &#x00B1;0.06 from the median, with no large tails in the scatter cloud, again suggesting that the depth-wise separable backbone gradually loses high-frequency energy while preserving a core stable representation.</p>
<p>FreqNet-MM22 (<xref ref-type="fig" rid="fig-6">Fig. 6</xref>) is the one that suffers the most severe collapse: the median jumps down from 0.92 (scale 2) to 0.18 (scale 16), for a change of &#x2212;0.74. The IQR falls from 0.12 to 0.02, and by scale 8, the 95% envelope dips below zero, producing a &#x2018;sign-flip band&#x2019; that affects approximately 18% of the samples. The scatter cloud becomes progressively more asymmetric beyond scale 4 and builds a high-density ridge close to zero drop. This signature is indicative that the learned Frequency Transformation Module is primarily biased towards upper octaves; after removing these bands, the internal representation has a better fit with the training manifold, and the logit penalty decreases, giving per-sample validation of the inversion effect noted in the Gaussian blur experiment. The tight median slope and the small end of the IQR show that a relatively minor change in the frequencies results in significant changes in scores, which could result in vulnerability when adversarially attacked.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Freqnet model sensitivity to varying frequency masking scales.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-6.tif"/>
</fig>
<p>In the case of ResNet50 (<xref ref-type="fig" rid="fig-7">Fig. 7</xref>), the median is very stable and only slightly reduced from 0.50 at the original scale to 0.30 at the largest scale. The IQR does not exceed 0.03, and the width of the 95% envelope is less than 0.06. The scatter cloud is very compact, symmetric, and there are no sign-flips; this shows that the spatial residual path does not &#x201C;bet its chips&#x201D; on a particular frequency band and provides the most consistent per-image behavior.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Resnet model sensitivity to varying frequency masking scales.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-7.tif"/>
</fig>
<p>Taken together, these plots confirm the tendencies seen in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>: spatially conscious architectures (XceptionMinimal, ResNet50) have more gradual (in both height and width of the drop-off) decline of model accuracy with increasing coverage for small thresholds in upper spectral frequencies, while freq-aware approaches (FreqNet) increase the variance among samples when the upper spectrum is completely eliminated.</p>

</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Complementary AUC Analysis</title>
<p>In this subsection we consider if there exists consistency between the directional change of score stability (as quantified by mean FSI-drop) and the directional change of detection performance on the same control perturbation. This discussion will be purely qualitative rather than seeking any predictive/statistical relation.</p>
<p>Two architectures are considered: <bold>ResNet50</bold> (spatially-oriented) and <bold>FreqNet-MM22</bold> (frequency-aware). XceptionMinimal is not included due to unavailability of pretrained weights for the current benchmark. Both models are evaluated under isotropic Gaussian blur with &#x03C3; &#x2208; {0, 1, 2, 4}, using the same fixed evaluation set and deterministic inference pipeline to ensure comparability. At each &#x03C3;, mean FSI-drop and AUC are reported jointly.</p>
<p>Both models show a monotonic decrease in AUC as Gaussian blur strength g increases, accompanied by a corresponding monotonic increase in mean FSI-drop (<xref ref-type="fig" rid="fig-8">Fig. 8</xref>). ResNet50 exhibits a smooth degradation in both metrics, whereas FreqNet shows more variable scores and worse performance at lower blur levels. No inversion of this pattern is observed: the model with higher score variance at smaller blurs is also the one that deteriorates sooner. This indicates that FSI-drop can serve as an approximate measure of relative robustness, rather than an absolute one.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Comparison of FSI-drop and AUC for ResNet50 and FreqNet.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-8.tif"/>
</fig>
<p>This suggests that the FSI-drop can be used as an approximate indicator of relative robustness, rather than in absolute terms.</p>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Robustness under Combined Distortions</title>
<p>In this section, we investigate whether the robustness ordering and score-stability trends observed for single distortions are preserved under a multi-stage degradation pipeline. To simulate realistic image degradation, all models are evaluated on a sequential application of distortions: resizing, Gaussian blur, and JPEG compression, applied in this fixed order. Distortion parameters are drawn from the same ranges used in the single-distortion experiments (<xref ref-type="sec" rid="s4_2">Section 4.2</xref>) to ensure comparability. In this experiment, the specific parameter values are: resizing scale <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mrow><mml:mtext>RESIZE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.6</mml:mn></mml:math></inline-formula>, Gaussian blur standard deviation <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mrow><mml:mtext>BLUR</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2.0</mml:mn></mml:math></inline-formula>, and JPEG quality factor <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mrow><mml:mtext>JPEG</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>60</mml:mn></mml:math></inline-formula>.</p>
<p>For all models, we look at both the score shift from clean to distorted inputs as well as the distribution of performance losses, is measured by per-sample FSI-drop. The compound distortions cause higher overall score changes than single ones in every architecture, which implies more severe degradation. However, we observe that the order of robustness is maintained: ResNet50 has the most consistent score behavior and then comes XceptionMinimal, while we see that FreqNet-MM22 has the greatest amount of score variation and drop in performance. The individual model results for the joint distortion pipeline can be seen in <xref ref-type="fig" rid="fig-9">Figs. 9</xref>&#x2013;<xref ref-type="fig" rid="fig-11">11</xref>, with the left columns representing Score Shift, while the right ones represent Performance Loss Distribution. We do not observe any reversal of this robustness pattern, suggesting that models whose scores are less stable given small changes perform worse on average when exposed to a mixture of realistic distortions.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Score shift and performance loss distribution of FreqNet.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-9.tif"/>
</fig><fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Score shift and performance loss distribution of Resnet50.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-10.tif"/>
</fig><fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Score shift and performance loss distribution of Xception.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_78014-fig-11.tif"/>
</fig>
<p>These results also indicate that we are measuring relative, not absolute, robustness by interpreting FSI-drop. Under even multi-stage distortions, the coherence in comparing scores implies that FSI-drop is able to measure the robustness to both single perturbations and a broader class of degradations.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion</title>
<p>Overall, the results indicate that robustness to common image distortions is closely linked to how an architecture&#x2019;s inductive priors allocate emphasis across spatial and spectral cues. Architectures that aggregate information over spatial neighborhoods, such as ResNet50, exhibit limited sensitivity to both frequency attenuation and spatial degradation. In contrast, explicitly frequency-aware designs show larger variability when their assumed spectral structure is disrupted. These observations suggest that generalization in deepfake detection is governed less by model capacity and more by which frequency ranges the learned representation emphasizes.</p>
<p>A particularly revealing phenomenon is the inverse robustness response observed for FreqNet-MM22 under Gaussian blur and frequency masking. Rather than improving robustness, explicit high-frequency modeling introduces a mismatch between training-time spectral assumptions and perturbed inputs, resulting in reduced score consistency and increased image-to-image variability. By comparison, architectures without explicit frequency specialization degrade more smoothly, indicating that distributed spatial representations provide implicit regularization against spectral shifts. Taken together, the findings reveal a design trade-off: mechanisms that amplify sensitivity to fine-grained frequency artifacts may enhance nominal discrimination, but increased fragility to both incidental and adversarial spectral perturbations.</p>
<p>Although this paper has focused here on single-distortion settings, in practice distortions may be combined or lie within some adversary-chosen bands. Previous works [<xref ref-type="bibr" rid="ref-10">10</xref>] show that frequency-aware detectors are susceptible to FFT-constrained attacks under low-budget noise. Extending the current robustness framework to such targeted attacks is needed to determine whether the observed architecture ordering and trade-off persist in more challenaging conditions.</p>
<p>This detailed study also supports a general design insight: how representation budget is allocated between space and frequency controls sensitivity and stability; frequency-based components can detect small manipulation artifacts, though with higher sample variance, and that paths oriented to space drive the homogenization of performances among different types of perturbations. The findings provide concrete guidelines for designing deepfake detection architectures that balance local frequency sensitivity with global robustness.</p>
<p><bold><italic>Summary of Robustness Findings</italic></bold></p>
<p>The results of the robustness comparison between ResNet50, XceptionMinimal, and FreqNet-MM22 are consistent for global and frequency-specific perturbations. JPEG compression has minimal effect on any of the detectors, indicating low sensitivity to small quality degradation. Gaussian blur is the most discriminative stress test, separating architectures both in trend and magnitude: FreqNet&#x2019;s FSI-drop shows an inverted response due to relying on high-frequency cues, XceptionMinimal shows moderate sensitivity, and ResNet50 is rather unchanged. Further spatial downsampling also confirms the same ordering, with FreqNet suffering most, and followed by XceptionMinimal, and ResNet50 shows strong robustness.</p>
<p>At the sample level, frequency masking increases image-to-image heterogeneity for FreqNet, whereas XceptionMinimal and ResNet50 degrade smoothly and symmetrically. All these results point towards an important consequence of the architectural choice: while using explicit frequency-based units improve sensitivity to small changes, it also makes models more vulnerable to spectral shifts, whereas spatially-motivated architectures prefer generalized robustness and consistent per-image behavior. This knowledge may guide the development of next-generation detectors, where it seems clear that an appropriate balance between spatial and frequency domain representation is needed for maximizing sensitivity while maintaining stability.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>This work provides an in-depth robustness analysis of ResNet50, XceptionMinimal, and FreqNet-MM22 on the ForenSynths dataset, based on over 3M FSI-drop measurements for popular distortions and limited-frequency perturbations. All models are robust against JPEG compression, while Gaussian blur combined with strong downsampling introduces substantial variance (especially for frequency-aware networks). The results show that there is a trade-off between the frequency and space-oriented modules: frequency-aware filters improve detection of small artifacts; however, they become less robust to spectral shifts, while spatially-driven designs remain robust to perturbations. These findings offer practical guidance for designing deepfake detectors that balance sensitivity to fine-grained artifacts with overall robustness. Although this study focuses on one type of distortion, future research should expand upon this analysis with cascaded, recaptured, and even adversarially generated spectral perturbations to obtain a more comprehensive understanding of the robustness landscape.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>All data generated or analyzed during this study are included in this published article. The ForenSynths dataset used in this study is publicly available at <ext-link ext-link-type="uri" xlink:href="https://drive.google.com/file/d/1AhWOsdCalrXE_6RmBZzyiC1ehXa6GXBK/view?usp=sharing">https://drive.google.com/file/d/1AhWOsdCalrXE_6RmBZzyiC1ehXa6GXBK/view?usp=sharing</ext-link>. All code for model training, evaluation, and robustness analysis is available at <ext-link ext-link-type="uri" xlink:href="https://github.com/zuoyan44/Frequency-Aware-Robustness-Analysis-of-Deepfake-Detection-Models.git">https://github.com/zuoyan44/Frequency-Aware-Robustness-Analysis-of-Deepfake-Detection-Models.git</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The author declares no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Karras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Aila</surname> <given-names>T</given-names></string-name>, <string-name><surname>Laine</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lehtinen</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Progressive growing of GANs for improved quality, stability, and variation</article-title>. In: <conf-name>Proceedings of the International Conference on Learning Representations (ICLR); 2018 Apr 30&#x2013;May 3</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Karras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Laine</surname> <given-names>S</given-names></string-name>, <string-name><surname>Aila</surname> <given-names>T</given-names></string-name></person-group>. <article-title>A style-based generator architecture for generative adversarial networks</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019 Jun 15&#x2013;20</conf-name>; <publisher-loc>Long Beach, CA, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>4401</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2019.00453</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Nichol</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Diffusion models beat GANs on image synthesis</article-title>. In: <conf-name>Proceedings of the 35th International Conference on Neural Information Processing Systems (NeurIPS 2021); 2021 Dec 6&#x2013;14</conf-name>; <publisher-loc>Virtual. Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2021</year>. p. <fpage>8780</fpage>&#x2013;<lpage>94</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chesney</surname> <given-names>R</given-names></string-name>, <string-name><surname>Citron</surname> <given-names>DK</given-names></string-name></person-group>. <article-title>Deep fakes: a looming challenge for privacy, democracy, and national security</article-title>. <source>Calif Law Rev</source>. <year>2019</year>;<volume>107</volume>:<fpage>1753</fpage>&#x2013;<lpage>820</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mirsky</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>W</given-names></string-name></person-group>. <article-title>The creation and detection of deepfakes: a survey</article-title>. <source>ACM Comput Surv</source>. <year>2022</year>;<volume>54</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3425780</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Afchar</surname> <given-names>D</given-names></string-name>, <string-name><surname>Nozick</surname> <given-names>V</given-names></string-name>, <string-name><surname>Yamagishi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Echizen</surname> <given-names>I</given-names></string-name></person-group>. <article-title>MesoNet: a compact facial video forgery detection network</article-title>. In: <conf-name>2018 IEEE International Workshop on Information Forensics and Security (WIFS); 2018 Dec 11&#x2013;13</conf-name>; <publisher-loc>Hong Kong, China. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/WIFS.2018.8630761</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Le</surname> <given-names>QV</given-names></string-name></person-group>. <article-title>EfficientNet: rethinking model scaling for convolutional neural networks</article-title>. In: <conf-name>Proceedings of the 36th International Conference on Machine Learning (ICML); 2019 Jun 9&#x2013;15</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. <year>2019</year>. p. <fpage>6105</fpage>&#x2013;<lpage>14</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chollet</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Xception: deep learning with depthwise separable convolutions</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21&#x2013;26</conf-name>; <publisher-loc>Honolulu, HI, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>1800</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2017.195</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rossler</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cozzolino</surname> <given-names>D</given-names></string-name>, <string-name><surname>Verdoliva</surname> <given-names>L</given-names></string-name>, <string-name><surname>Riess</surname> <given-names>C</given-names></string-name>, <string-name><surname>Thies</surname> <given-names>J</given-names></string-name>, <string-name><surname>Niessner</surname> <given-names>M</given-names></string-name></person-group>. <article-title>FaceForensics&#x002B;&#x002B;: learning to detect manipulated facial images</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); 2019 Oct 27&#x2013;Nov 2</conf-name>; <publisher-loc>Seoul, Republic of Korea. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv.2019.00009</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Frequency-aware deepfake detection: improving generalizability through frequency space domain learning</article-title>. In: <conf-name>Proceedings of the 38th AAAI Conference on Artificial Intelligence (AAAI); 2024 Feb 20&#x2013;27</conf-name>; <publisher-loc>Vancouver, BC, Canada. Palo Alto, CA, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>; <year>2024</year>. p. <fpage>5052</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i5.28310</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Improving the generalization of face forgery detection via single domain augmentation</article-title>. <source>Multimed Tools Appl</source>. <year>2024</year>;<volume>83</volume>(<issue>26</issue>):<fpage>63975</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-023-17840-2</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Spatial-phase shallow learning: rethinking face forgery detection in frequency domain</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021 Jun 19&#x2013;25</conf-name>; <publisher-loc>Virtual. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>772</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr46437.2021.00083</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ciftci</surname> <given-names>UA</given-names></string-name>, <string-name><surname>Demir</surname> <given-names>I</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>L</given-names></string-name></person-group>. <article-title>FakeCatcher: detection of synthetic portrait videos using biological signals</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2020</year>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2020.3009287</pub-id>; <pub-id pub-id-type="pmid">32750816</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sabir</surname> <given-names>E</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jaiswal</surname> <given-names>A</given-names></string-name>, <string-name><surname>AbdAlmageed</surname> <given-names>W</given-names></string-name>, <string-name><surname>Mazaheri</surname> <given-names>G</given-names></string-name>, <string-name><surname>Natarajan</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Recurrent convolutional strategies for face manipulation detection in videos</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 2019 Jun 16&#x2013;17</conf-name>; <publisher-loc>Long Beach, CA, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>80</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Davis</surname> <given-names>L</given-names></string-name>, <string-name><surname>Fritz</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Attributing fake images to GANs: learning and analyzing GAN fingerprints</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); 2019 Oct 27&#x2013;Nov 2</conf-name>; <publisher-loc>Seoul, Republic of Korea. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>7556</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv.2019.00765</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Shlens</surname> <given-names>J</given-names></string-name>, <string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Explaining and harnessing adversarial examples</article-title>. In: <conf-name>Proceedings of the 3rd International Conference on Learning Representations (ICLR); 2015 May 7&#x2013;9</conf-name>; <publisher-loc>San Diego, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Madry</surname> <given-names>A</given-names></string-name>, <string-name><surname>Makelov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tsipras</surname> <given-names>D</given-names></string-name>, <string-name><surname>Vladu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Towards deep learning models resistant to adversarial attacks</article-title>. In: <conf-name>Proceedings of the 6th International Conference on Learning Representations (ICLR); 2018 Apr 30&#x2013;May 3</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>SY</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>O</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Owens</surname> <given-names>A</given-names></string-name>, <string-name><surname>Efros</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>CNN-generated images are surprisingly easy to spot... for now</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020 Jun 13&#x2013;19</conf-name>; <publisher-loc>Seattle, WA, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2020</year>. p. <fpage>8692</fpage>&#x2013;<lpage>701</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr42600.2020.00872</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Auto-encoding variational bayes</article-title>. In: <conf-name>Proceedings of the 2nd International Conference on Learning Representations (ICLR); 2014 Apr 14&#x2013;16</conf-name>; <publisher-loc>Banff, AB, Canada</publisher-loc>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Karras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Laine</surname> <given-names>S</given-names></string-name>, <string-name><surname>Aittala</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hellsten</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lehtinen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Aila</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Analyzing and improving the image quality of StyleGAN</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020 Jun 13&#x2013;19</conf-name>; <publisher-loc>Seattle, WA, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2020</year>. p. <fpage>8107</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr42600.2020.00813</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Karras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Aittala</surname> <given-names>M</given-names></string-name>, <string-name><surname>Laine</surname> <given-names>S</given-names></string-name>, <string-name><surname>H&#x00E4;rk&#x00F6;nen</surname> <given-names>E</given-names></string-name>, <string-name><surname>Hellsten</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lehtinen</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Alias-free generative adversarial networks</article-title>. In: <conf-name>Proceedings of the 35th International Conference on Neural Information Processing Systems (NeurIPS 2021); 2021 Dec 6&#x2013;14</conf-name>. <publisher-loc>Virtual. Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2021</year>. p. <fpage>852</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Marra</surname> <given-names>F</given-names></string-name>, <string-name><surname>Gragnaniello</surname> <given-names>D</given-names></string-name>, <string-name><surname>Verdoliva</surname> <given-names>L</given-names></string-name>, <string-name><surname>Poggi</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Do GANs leave artificial fingerprints?</article-title>. In: <conf-name>Proceedings of the 2019 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR); 2019 Mar 28&#x2013;30</conf-name>; <publisher-loc>San Jose, CA, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>506</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mipr.2019.00103</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>G&#x00FC;era</surname> <given-names>D</given-names></string-name>, <string-name><surname>Delp</surname> <given-names>EJ</given-names></string-name></person-group>. <article-title>Deepfake video detection using recurrent neural networks</article-title>. In: <conf-name>2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS); 2018 Nov 27&#x2013;30</conf-name>; <publisher-loc>Auckland, New Zealand. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/AVSS.2018.8639163</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Haliassos</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mira</surname> <given-names>R</given-names></string-name>, <string-name><surname>Petridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Leveraging real talking faces via self-supervision for robust forgery detection</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18&#x2013;24</conf-name>; <publisher-loc>New Orleans, LA, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>14930</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01453</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dutta</surname> <given-names>A</given-names></string-name>, <string-name><surname>Das</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Naskar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chakraborty</surname> <given-names>RS</given-names></string-name></person-group>. <article-title>WaveDIF: wavelet sub-band based deepfake identification in frequency domain</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 2025 Jun 11&#x2013;12</conf-name>; <publisher-loc>Nashville, TN, USA. Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>6302</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPRW67362.2025.00627</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>