<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">53740</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.053740</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>IMTNet: Improved Multi-Task Copy-Move Forgery Detection Network with Feature Decoupling and Multi-Feature Pyramid</article-title>
<alt-title alt-title-type="left-running-head">IMTNet: Improved Multi-Task Copy-Move Forgery Detection Network with Feature Decoupling and Multi-Feature Pyramid</alt-title>
<alt-title alt-title-type="right-running-head">IMTNet: Improved Multi-Task Copy-Move Forgery Detection Network with Feature Decoupling and Multi-Feature Pyramid</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Huan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Hong</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Jiang</surname><given-names>Zhongyuan</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>jiangzydj@163.com</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Qian</surname><given-names>Qing</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Long</surname><given-names>Yong</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Information, Guizhou University of Finance and Economics</institution>, <addr-line>Guiyang, 550025</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>College of Big Data Statistics, Guizhou University of Finance and Economics</institution>, <addr-line>Guiyang, 550025</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Zhongyuan Jiang. Email: <email>jiangzydj@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>12</day><month>9</month><year>2024</year></pub-date>
<volume>80</volume>
<issue>3</issue>
<fpage>4603</fpage>
<lpage>4620</lpage>
<history>
<date date-type="received">
<day>09</day>
<month>5</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>8</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_53740.pdf"></self-uri>
<abstract>
<p>Copy-Move Forgery Detection (CMFD) is a technique that is designed to identify image tampering and locate suspicious areas. However, the practicality of the CMFD is impeded by the scarcity of datasets, inadequate quality and quantity, and a narrow range of applicable tasks. These limitations significantly restrict the capacity and applicability of CMFD. To overcome the limitations of existing methods, a novel solution called IMTNet is proposed for CMFD by employing a feature decoupling approach. Firstly, this study formulates the objective task and network relationship as an optimization problem using transfer learning. Furthermore, it thoroughly discusses and analyzes the relationship between CMFD and deep network architecture by employing ResNet-50 during the optimization solving phase. Secondly, a quantitative comparison between fine-tuning and feature decoupling is conducted to evaluate the degree of similarity between the image classification and CMFD domains by the enhanced ResNet-50. Finally, suspicious regions are localized using a feature pyramid network with bottom-up path augmentation. Experimental results demonstrate that IMTNet achieves faster convergence, shorter training times, and favorable generalization performance compared to existing methods. Moreover, it is shown that IMTNet significantly outperforms fine-tuning based approaches in terms of accuracy and <italic>F</italic><sub>1</sub>.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Image copy-move detection</kwd>
<kwd>feature decoupling</kwd>
<kwd>multi-scale feature pyramids</kwd>
<kwd>passive forensics</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Guizhou Provincial Science and Technology Project</funding-source>
<award-id>QKH-Basic-ZK[2021]YB311</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Youth Science and Technology Talent Growth Project of Guizhou Provincial Education Department</funding-source>
<award-id>QJH-KY-ZK[2021]132</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Guizhou Provincial Science and Technology Project</funding-source>
<award-id>QKH-Basic-ZK[2021]YB319</award-id>
</award-group>
<award-group id="awg4">
<funding-source>National Natural Science Foundation of China (NSFC)</funding-source>
<award-id>61902085</award-id>
</award-group>
<award-group id="awg5">
<funding-source>Key Laboratory Program of Blockchain and Fintech of Department of Education of Guizhou Province</funding-source>
<award-id>2023-014</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Malicious image forgeries, including techniques such as copy-move, splicing, and removal, can severely undermine the credibility and integrity of digital images. The copy-move tampering method, being one of the most prevalent image tampering techniques, is highly concealable and presents significant challenges for detection. Moreover, the characteristics (such as the saturation, light source, and noise) of the tampered areas can be adapted easily without affecting the original image properties. Therefore, CMFD methods have attracted the attention of forensic science scholars.</p>
<p>The traditional CMFD methods can be classified into two main categories: block-based methods and keypoint-based methods. In the block-based methods, discrete cosine transform (DCT) technologies are generally used in image processing because of their energy compaction properties. The work in [<xref ref-type="bibr" rid="ref-1">1</xref>] analyzes the exhaustive search algorithms and proposes a block matching detection method based on DCT. This work is one of the landmark methods in CMFD methods. Each block contains 64 (8 &#x00D7; 8) features and any two feature vectors that are within a certain range should be matched to determine the duplicated regions. However, the proposed method cannot detect the small duplicated regions and the detection precision is dissatisfactory because of its lexicographically sorting algorithm. The study in [<xref ref-type="bibr" rid="ref-2">2</xref>] improves the method of [<xref ref-type="bibr" rid="ref-1">1</xref>] by reducing the number of features to a quarter that is located in the low frequency parts. However, the detection accuracy is unsatisfactory since some DCT coefficients that are located in the intermediate frequency parts are truncated. A Discrete Cosine Transformation (DCT) and Singular Value Decomposition (SVD) based technique is proposed to detect the copy-move image forgery in [<xref ref-type="bibr" rid="ref-3">3</xref>], the combination of DCT and SVD makes the proposed scheme robust against compression, geometric transformations, and noise. A hybrid method is reported to classify copy-move and splicing images based on the texture information of images in the spatial domain [<xref ref-type="bibr" rid="ref-4">4</xref>]. The proposed method divides the image into equal blocks to get scale-invariant features. The tampered image regions can be detected by matching the scale-invariant features. The proposed method is robust to most regular signal processing type attacks. However, it is less effective against some geometric transformation-type attacks.</p>
<p>In the keypoint-based methods, the underlying principle is that modifications made to the image, such as copy-move operations, will alter the local features and consequently impact the distribution and characteristics of the detected keypoints. By analyzing the changes in the keypoint-based representations, these techniques aim to identify the presence of tampering in the image. Amerini et al. in [<xref ref-type="bibr" rid="ref-5">5</xref>] propose a novel methodology based on a scale invariant features transform (SIFT) method. The proposed method can be used to determine whether a copy-move attack has occurred in an image. It also can be used to recover the geometric transformation that is used to perform cloning technologies. Furthermore, the proposed method can be used to individuate the altered areas and estimate the geometric transformation parameters with high reliability. The work in [<xref ref-type="bibr" rid="ref-6">6</xref>] reports an improved SIFT structure with inherent scaling invariance that is designed to enhance the capability of extracting effective keypoints in the homogeneous region. Zhong et al. in [<xref ref-type="bibr" rid="ref-7">7</xref>] analyze the structure and excavate the inherent characteristics of local descriptors (SURF) for feature extraction in the coarse and smooth regions. Subsequently, the proposed method utilizes kernel features for coarse feature matching to reduce matching costs. Following this, a smaller set of candidate keypoints is identified, which are then used in conjunction with complete features to conduct fine keypoint matching in order to identify suspicious candidate keypoint pairs. A method is proposed in [<xref ref-type="bibr" rid="ref-8">8</xref>] to find and locate the duplicated and pasted portions of a manipulated image by using the combination of Hessian and Raw patch features. In the proposed method, a parallelism condition is applied together with a random sample consensus method to eliminate mismatches. The proposed method was shown to be effective by obtaining high <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> scores in images that are attacked with noise, JPEG compression, and scaling operations.</p>
<p>Traditional image tampering detection methods have faced growing challenges in keeping pace with the rapidly evolving landscape of image forgery techniques. In response, the increased pervasiveness of deep learning technologies has led to the development of contemporary copy-move forgery detection algorithms that predominantly leverage specialized neural network architectures, which are purposefully designed and trained for the task of image tampering identification. It is proposed in [<xref ref-type="bibr" rid="ref-9">9</xref>] that an end-to-end approach called BusterNet can identify the source and target regions by detecting image similarity through parallel branching. However, this method requires high accuracy on both branches. The work in [<xref ref-type="bibr" rid="ref-10">10</xref>] reports a serial branching network that is used to improve the drawbacks of BusterNet. The reported network consists of a copy-move similarity detection network and a source/target region distinguishment network. The branching network is simpler and more accurate compared with the BusterNet. However, generalization was powerless. The study in [<xref ref-type="bibr" rid="ref-11">11</xref>] proposes the Dense-InceptionNet that combines DenseNet and InceptionNet, by utilizing multiscale information and dense features. The study presented in [<xref ref-type="bibr" rid="ref-12">12</xref>] introduces the Spatial Pyramidal Attention Network which is designed to capture inter-block relationships across multiple scales through a pyramidal structure of locally self-attentive blocks. However, this approach demonstrates diminished effectiveness at lower image resolutions and overlooks the interplay between high-dimensional and low-dimensional features. The research in [<xref ref-type="bibr" rid="ref-13">13</xref>] reports a deep learning method for forgery detection at both image and pixel levels. In this method, authors used a pre-trained deep model with a global average pooling (GAP) layer instead of default fully connected layers to detect forgery. The GAP layer creates a good dependency between the feature maps and the classes. The study in [<xref ref-type="bibr" rid="ref-14">14</xref>] proposes the Laterally Linked Pixel (LLP) algorithm, which utilizes two-dimensional arrays and a single layer derived from a unit-linking pulsed neural network to detect copied regions. The method employs kernel tricks to identify multiple manipulations within a single forged image. The accuracy obtained through the LLP algorithm is about 90% and further forgery detection is improved based on optimized kernel selections in classification algorithm.</p>
<p>While extensive experimental evaluations have demonstrated the satisfactory performance of the aforementioned specialized neural network-based copy-move forgery detection methods, such approaches inherently compromise the broader generalizability of the underlying network architecture. This runs counter to the primary objective of neural networks, which is to learn robust, generalizable representations that can be effectively applied across a diverse range of related tasks and domains. Consequentially, the following problems arise: (1) Deep neural networks (DNNs) require a substantial quantity of high-quality labeled datasets [<xref ref-type="bibr" rid="ref-15">15</xref>]. (2) DNNs need hardware with higher computing capacity. (3) As a data-driven algorithm, DNNs produce individual results based on various types of data. However, collecting an enormous amount of data does not provide a comprehensive representation. (4) DNNs aim to create a general model that could cater to the needs of different users, environments, and devices. Consequently, there remains an urgent need to adapt and refine the general model to address personalized tasks effectively [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>Transfer learning is concerned with the transfer of knowledge across domains, leveraging prior experience as a bridge to facilitate the adaptation from one scenario to another. Among its various subfields, feature decoupling stands out as a significant subclass, demonstrating a broad spectrum of applications. Feature decoupling is an approach to designing neural network architectures and training processes that aim to make the features learned from the network independent or uncorrelated from each other. The main purpose of feature decoupling is to improve the generalization ability and interpretability of the model. A feature decoupled training pipeline for describe-then-detect is designed for weakly supervised local feature learning [<xref ref-type="bibr" rid="ref-17">17</xref>]. Additionally, an introduced line-to-window search strategy enhances descriptor learning by explicitly utilizing camera pose information, and attained state-of-the-art performance across various tasks. The work in [<xref ref-type="bibr" rid="ref-18">18</xref>] reports a multi-scale single image deraining network, called the feature decoupling and reorganization network, which introduces a dilated pyramid split attention module to decouple input features and reorganize extracted features.</p>
<p>The aforementioned work exemplifies practical applications of feature decoupling, which often entails specific modifications to network architecture, such as the incorporation of regularization techniques or the addition of terms to the loss function to promote feature independence. By employing feature decoupling, the model is anticipated to acquire more robust and distinguishable feature representations, thereby enhancing its performance on novel and previously unseen data.</p>
<p>Image semantic information can be broadly categorized into three layers: the visual, object, and conceptual layers. The visual layer, often referred to as the shallow feature layer, encompasses detailed attributes such as color and shape. The object layer, or mid-level feature, usually contains object attributes. The conceptual layer, known as the high-level feature contains abundant semantic information. ResNet-50 [<xref ref-type="bibr" rid="ref-19">19</xref>] is specifically designed for image classification tasks. However, the tasks of CMFD and image classification have distinct focuses on the image, which necessitate the network to possess diverse feature extraction capabilities and representation capabilities. Meanwhile, quantitative analysis is performed experimentally on the pre-trained ResNet-50 to achieve structural risk minimization. It is worth mentioning that the 3461 model is a variant of the ResNet-50. It indicates that the ConvNet Layer of four modules in ResNet-50 is repeated 3, 4, 6, and 1 times, respectively.</p>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> illustrates the comparison of the ResNet-50 deep module (fourth module) possessing varying numbers of convolutional layers, where the 3 4 6 1, 3 4 6 3 and 3 4 6 6 represent the numbers of stacks of residual blocks within the module, respectively. Multiple classes of models with slight variations were trained while keeping the first three modules frozen. The comparative analysis of visualization outcomes derived from the three distinct models suggests that increasing the depth of the network architecture enhances the ability to extract complex details. Nevertheless, based on the pertinent graphical representation, it can be inferred that integrating more complex structures into the model is inappropriate for tackling the CMFD task.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Contrast analysis. The first three module weights are frozen based on a pre-trained ResNet-50. The effect of the model is analyzed while the convolution layer of the fourth module is changed. Source and tampered areas are indicated in blue and gray</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-1.tif"/>
</fig>
<p>To address the aforementioned limitations, this paper introduces a feature decoupling approach within the context of transfer learning, leveraging knowledge migration to enhance model performance. When the source and target domains are similar, the main challenge in feature decoupling is determining which layers of knowledge in the source domain network should be fixed or fine-tuned [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
<p>In the proposed scheme, the relationship between the CMFD task and the number of deep feature repetition layers in the pre-trained ResNet-50 is explored. ResNet-50 is introduced for several compelling reasons: (1) the incorporation of identity shortcut connections within the residual branching structure (RBS) optimizes the process of backpropagation, thereby rendering ResNet-50, a simple and highly efficient method. (2) ResNet-50 employs a limited number of optimization techniques effectively, thereby minimizing potential interferences. Furthermore, the architecture incorporates two types of RBS. RBS-1 utilizes a step size of 2, which significantly reduces the output size and mitigates the risk of overfitting. On the other hand, RBS-2 serves as the major module in the ResNet network, primarily focused on enhancing the representational capabilities. (3) In addition, in the field of passive forensics, the scarcity of high-quality datasets leads to low model generalization ability and accuracy rates.</p>
<p>The objective of this paper is to streamline the multi-task framework of ResNet-50 to mitigate disaster forgetting and enhance task efficiency. To accomplish this, the multi-task objective is simplified into a single-task objective. Specifically, the focus is on maximizing the utilization of the solution space for the image classification task through pertinent experiments. Our goal is to enhance the model performance on the CMFD task when the image classification task has already reached its optimum. This paper proposes an improved multi-task copy-move forgery detection network by using a multi-feature pyramid module (MFPM) and features decoupling across tasks. The main contributions of this study are as follows:
<list list-type="order">
<list-item><p>An optimization problem is abstracted to establish a link between image classification and CMFD based on ResNet-50 using feature decoupling, wherein the relationship between the deep structure and task is illustrated during optimization.</p></list-item>
<list-item><p>This paper provides a quantitative demonstration of the similarities between image classification tasks and CMFD. Moreover, it addresses the challenge of limited high-quality data in CMFD through the transfer of frozen weights and retraining of the model.</p></list-item>
<list-item><p>The CMFD field saw the first introduction of the MFPM. It utilizes three matching maps to detect suspicious regions and enhances localization accuracy through the application of Feature Pyramid Networks and an optimized bottom-up pathway.</p></list-item>
</list></p>
<p>The rest of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> describes the stages of the proposed method and briefly explains every step. <xref ref-type="sec" rid="s3">Section 3</xref> presents databases that are being used for experiments and an experimental setup and results in discussions for the proposed architecture. It offers tables and figures related to results calculated using the proposed architecture. At last, <xref ref-type="sec" rid="s4">Section 4</xref> provides conclusions for this study.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Proposed Method</title>
<p>To address the challenges in dataset building and training due to the difficulty of collecting high-quality annotated collections of tampered images in CMFD, IMTNet is proposed in this paper, the diagram is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. It comprises two components: copy-move forgery detection algorithm based on feature decoupling and tampered image localization algorithm named MFPN, which effectively leverages both high and low-dimensional image information.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The diagram of IMTNet for copy-move forgery detection and localization. (a) Improved ResNet-50 backbone. (b) MFPN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-2.tif"/>
</fig>
<sec id="s2_1">
<label>2.1</label>
<title>Copy-Move Forgery Detection Algorithm Based on Feature Decoupling</title>
<p><bold>ResNet-50 based on transfer learning.</bold> Due to the difficulty of collecting datasets in CMFD, the current number of datasets available is limited. Furthermore, the varying shallow weights obtained during different training epochs can potentially disrupt correlation ablation experiments. To tackle the aforementioned challenges, our proposal involves the adoption of a feature decoupling approach and enhancements to the network structure of the pre-trained ResNet-50. The primary objective is to maximize the leverage of the source domain as a feature extraction network for the target domain.</p>
<p>The pre-trained ResNet-50 combined with the feature decoupling approach empowers the IMTNet to acquire highly relevant implicit expression features. Furthermore, feature decoupling is employed as a regularization technique to mitigate dissimilarities between the source and target marginal distributions. This mathematical description can be formulated as
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>Y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>S</mml:mi></mml:math></inline-formula> indicates the source domain and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the sample space in the source domain, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>Y</mml:mi></mml:math></inline-formula> denote the joint feature space and the corresponding label space.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>in which <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>T</mml:mi></mml:math></inline-formula> denotes target domain and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the sample space in the target domain. Besides, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the number of samples in the source domain. There are also <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> samples in the target domain. <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>S</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>T</mml:mi></mml:math></inline-formula> have different probability distributions, transfer learning transfers knowledge from <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>S</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>T</mml:mi></mml:math></inline-formula> to execute specific tasks on <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>T</mml:mi></mml:math></inline-formula>, the transfer process is shown as
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:munder><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:munder><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></disp-formula>in which <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the predicted label after knowledge transfer. From this, feature decoupling is realized by transferring the features of classification to CMFD. Therefore, the features of CMFD can be extracted to verify the integrity of an image.</p>
<p><bold>Optimization problem solution for multi-task based on ResNet-50.</bold> A ConvNet Layer <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>i</mml:mi></mml:math></inline-formula> can be defined as a function: <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the convolution operator, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the output tensor, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the input tensor. A ConvNet <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>N</mml:mi></mml:math></inline-formula> can be represented by a list of composed layers
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2026;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mrow><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:mrow></mml:math></inline-formula> represents the multiplication operation. There are five stages in ResNet-50, and every layer in the last four stages has the same convolutional type, except for the first layer which performs downsampling. Therefore, ConvNet can be defined as
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:munder><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2026;</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>N</mml:mi></mml:math></inline-formula> represents the abstracted whole network processing, <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>i</mml:mi></mml:math></inline-formula> is the module serial number. <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents a convolution operator, <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msup><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> means that the <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> operation is repeated <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> times in module <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>i</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>X</mml:mi></mml:math></inline-formula> denotes the input tensor in stage <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>i</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mo>&#x003E;</mml:mo></mml:math></inline-formula> indicates width, height and number of channels. Our purpose is to maximize the model accuracy for given resource constraints, which can be formulated as an optimization problem.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:munder><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The multi-task optimization problem described above is transformed into a single optimization problem, where <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mover><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula> indicates that ResNet-50 has undergone optimization for image classification tasks, and <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents ResNet-50 employed for the CMFD task, <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>d</mml:mi></mml:math></inline-formula> is a scaling factor, <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mover><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mover><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula>are predefined parameters in ResNet-50 and <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mrow><mml:mo>(</mml:mo><mml:mover><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mover><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>224</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mn>224</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, denoted as
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:munder><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:munder><mml:mo>&#x2061;</mml:mo><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mover><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mo>.</mml:mo><mml:mi>t</mml:mi><mml:mo>.</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2026;</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>.</mml:mo><mml:mover><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>y</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>F</mml:mi><mml:mi>L</mml:mi><mml:mi>O</mml:mi><mml:mi>P</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>f</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>s</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Initially, we assume that the model is the most optimal image classification solution. According to the steps above, this optimal solution can be adapted to the specifics of image copy-move tampering. In the deep model, the number of iterations between layers is scaled by <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>d</mml:mi></mml:math></inline-formula>. Following feature decoupling and model optimization, the implementation of a multi-task ResNet-50 exhibits significant advantages. It effectively circumvents catastrophic forgetting while also improving the detection accuracy for the CMFD tasks.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Tampered Image Positioning Algorithm</title>
<p>The MFPN incorporates top-down and bottom-up bidirectional fusion branches, which combine low-level features information and high-level semantic information to improve the accuracy of semantic representation [<xref ref-type="bibr" rid="ref-21">21</xref>]. Therefore, IMTNet leverages the inherent multi-scale and hierarchical characteristics of the network to construct MFPN, which greatly enhances its representational capabilities and localization quality. Moreover, it reduces the cost of building MFPN. The output results of the three matching maps and their combination are shown in <xref ref-type="fig" rid="fig-3">Fig. 3a,b</xref> show the forgery image and the ground truth mask, and <xref ref-type="fig" rid="fig-3">Fig. 3c</xref>&#x2013;<xref ref-type="fig" rid="fig-3">f</xref> shows the results of the three up-sampled matching maps I, II and III and the final suspicious area location results.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The output results of the three matching maps and their combination</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-3.tif"/>
</fig>
<p>The numerous untrained forgery classes or objects result in difficulty in applying a classic DNNs model to address those data. Therefore, an auxiliary image tampering localization model is proposed to learn the correlations between the rich hierarchical features. In the proposed scheme, <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>P</mml:mi></mml:math></inline-formula> features are extracted for each image, which are the different features of the image after applying different convolution operators to the image. Assuming that sets of the feature point is <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, the <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>M</mml:mi></mml:math></inline-formula>-dimensional description operator of <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>P</mml:mi></mml:math></inline-formula> can be expressed as
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="left left left left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where the parameter <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>M</mml:mi></mml:math></inline-formula> represents the depth of the feature, which is the number of channels, <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:math></inline-formula> represents the number of candidate pixels or the size of the feature maps in the candidate matrix. In the matching localization algorithm, the feature correlation coefficient between the defined feature points is denoted as
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mspace width="1em" /><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em" /><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em" /><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>M</mml:mi></mml:mfrac></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="left left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where the subscript <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>i</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>j</mml:mi></mml:math></inline-formula> represent the localizations of feature point <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the corresponding matching map. <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the matched measurement and represents the negative feature correlation coefficient between the feature point <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The definition of the correlation coefficient represents that the closer <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is to 0, the more similar <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are.</p>
<p>In the IMTNet, the 2NN matching algorithm [<xref ref-type="bibr" rid="ref-22">22</xref>] is used to reduce the matching errors. Assume <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the second minor characteristic correlation coefficient and <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the third minor characteristic correlation coefficient, <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> satisfy the following condition:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mfrac><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mfrac><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>The relevant features can be filtered according to <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref> while <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.65</mml:mn></mml:math></inline-formula>.</p>
<p>In <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>, set <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is transformed into a binary classification problem by setting it to an activation function that approximates the sigmoid function as</p>
<p><disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>2</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mfrac></mml:mstyle><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>2</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>. It is used to make the unmatched coefficient approach 0.</p>
<p>Finally, the feature matching coefficient <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is filled in the localization of the matching map. In this way, other feature points search for the best matching coefficients and fill in the matching localizations matrix. The hyperparameter <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>l</mml:mi></mml:math></inline-formula> are based on the input feature depths of the MFPN blocks <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mrow><mml:mtext>I&#xA0;</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>I I&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mrow><mml:mtext>I I I&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula>. <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> stands for three matching map combinations. The data in blocks <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mrow><mml:mtext>I&#xA0;</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>I I&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mrow><mml:mtext>I I I&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula> of MFPM is obtained through processing by the second, third, and fourth modules of the ResNet-50, respectively, denoted as</p>
<p><disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>I&#xA0;</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>v</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>I I&#xA0;</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>l</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>I I I&#xA0;</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>I&#xA0;</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>I I&#xA0;</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>I I I&#xA0;</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represent the numerical matrices of the first, second, and third modules, respectively, in the MFPM that is processed with the 2NN matching algorithm. The hyperparameter <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>l</mml:mi></mml:math></inline-formula> are obtained as</p>
<p><disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>32</mml:mn><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>32</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>I&#xA0;</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mn>48</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>I I&#xA0;</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mn>64</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>I I I&#xA0;</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>48</mml:mn><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>32</mml:mn><mml:mo>+</mml:mo><mml:mn>48</mml:mn><mml:mo>+</mml:mo><mml:mn>64</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>64</mml:mn><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>32</mml:mn><mml:mo>+</mml:mo><mml:mn>48</mml:mn><mml:mo>+</mml:mo><mml:mn>64</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mi>v</mml:mi><mml:mo>+</mml:mo><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where 32, 48, and 64 represent the number of channels in MFPM for locks <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mrow><mml:mtext>I&#xA0;</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>I I&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mrow><mml:mtext>I I I&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula>, respectively.</p>
<p>Task relevance can be trained based on the premise that similar tasks share the same model weights, and these tasks can be transformed or low-rank regularized to obtain richer representations. The feature decoupling approach used in this paper aims to capture multiple aspects of task relevance properties, such as sparsely and low-ranking of tasks. This is achieved by decomposing the model partial weights into the sum or product of different convolution operator components that capture information in addition to specific task information that is beneficial to each task. The flexibility of the feature decoupling technique provides a deeper understanding of the nature of multi-task, enabling feature sharing of model weights for both image classification and image tampering detection tasks [<xref ref-type="bibr" rid="ref-23">23</xref>].</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Experiments</title>
<sec id="s3_1">
<label>3.1</label>
<title>Experimental Sets</title>
<p>IMTNet is trained on two benchmark datasets: the CASIA2.0 [<xref ref-type="bibr" rid="ref-24">24</xref>] and CoMoFod_small [<xref ref-type="bibr" rid="ref-25">25</xref>] dataset. Meanwhile, the mentioned datasets are mixed as the target domain datasets. The blended dataset contains many manipulated images that have been attacked, which could enhance the &#x201C;quality&#x201D; of the dataset and improve its robustness. Moreover, MICC-F2000 [<xref ref-type="bibr" rid="ref-5">5</xref>], MICC-F600 [<xref ref-type="bibr" rid="ref-26">26</xref>], COVERAGE [<xref ref-type="bibr" rid="ref-27">27</xref>], and DEFACTO [<xref ref-type="bibr" rid="ref-28">28</xref>] are used to conduct generalization tests. Original and tampered images from the Ardizzone [<xref ref-type="bibr" rid="ref-29">29</xref>] dataset and the MICC-F2000 dataset are used. In addition, there are 140 images in the dataset for the attack resistance experiments, where 35 identical images come from the Ardizzone dataset that have undergone different tampering attacks, and in order to enlarge the size of the dataset, 35 images in the MICC-F2000 dataset are extracted that have also undergone different tampering attacks, totaling 70 tampered images. Thus, the interference due to different images is reduced and the model is realized for the anti-attack experiments.</p>
<p>Anchor points are used in the experiment to divide the semantic information of images, such as the visual layer, object layer, and concept layer. These points are chosen because each downsampling greatly enhances the representation of semantic information. Additionally, the shallow features of the images are generic, which is the reason that the shallow weights of the model are frozen.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Ablation Experiments</title>
<p>The ablation study is divided into two phases. The first phase assessed the similarity between the source and target domains to identify the appropriate number of layers to freeze. The second phase explored the correlation between the deep architecture of ResNet-50 and its performance on the CMFD task.</p>
<p><bold>Similarity between the image classification domain and CMFD domain.</bold> The deep architecture of ResNet-50 is analyzed at various scales. An experimental comparison is conducted to evaluate the representational capabilities of ResNet-50 when trained directly or with specific modules frozen. <xref ref-type="table" rid="table-1">Table 1</xref> and <xref ref-type="fig" rid="fig-4">Fig. 4</xref> present the experimental results from alternative perspectives, respectively. The data in <xref ref-type="table" rid="table-1">Table 1</xref> are the value of <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, accuracy (<inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula>), Precision (<inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>P</mml:mi></mml:math></inline-formula>) and Recall (<inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>R</mml:mi></mml:math></inline-formula>), which displays the average values that demonstrate the characterization ability of the models when different modules of ResNet-50 are frozen. The experiment images were obtained after data cleaning and mixing the MICC-2000 and MICC-600 datasets, the total number is 2600, and the valid data are 2560 after cleaning.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Image level ablation experiments on ResNet-50 using MICC-F2000 and MICC-F600 datasets</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Inter-layer numbers</th>
<th align="center" colspan="4">Only training</th>
<th align="center" colspan="4">Freeze all layers</th>
<th align="center" colspan="4">Freeze 1 layer</th>
<th align="center" colspan="4">Freeze 1, 2 layer</th>
<th align="center" colspan="4">Freeze 1, 2, 3 layer</th>
</tr>
<tr>
<th><italic>F</italic><sub>1</sub><break/>(%)</th>
<th><italic>ACC</italic><break/>(%)</th>
<th><italic>P</italic><break/>(%)</th>
<th><italic>R</italic><break/>(%)</th>
<th><italic>F</italic><sub>1</sub><break/>(%)</th>
<th><italic>ACC</italic><break/>(%)</th>
<th><italic>P</italic><break/>(%)</th>
<th><italic>R</italic><break/>(%)</th>
<th><italic>F</italic><sub>1</sub><break/>(%)</th>
<th><italic>ACC</italic><break/>(%)</th>
<th><italic>P</italic><break/>(%)</th>
<th><italic>R</italic><break/>(%)</th>
<th><italic>F</italic><sub>1</sub><break/>(%)</th>
<th><italic>ACC</italic><break/>(%)</th>
<th><italic>P</italic><break/>(%)</th>
<th><italic>R</italic><break/>(%)</th>
<th><italic>F</italic><sub>1</sub><break/>(%)</th>
<th><italic>ACC</italic><break/>(%)</th>
<th><italic>P</italic><break/>(%)</th>
<th><italic>R</italic><break/>(%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>3, 4, 2, 1</td>
<td>65.03</td>
<td>57.49</td>
<td>60.82</td>
<td>72.83</td>
<td>74.10</td>
<td>60.80</td>
<td>79.23</td>
<td>69.58</td>
<td>75.61</td>
<td>62.59</td>
<td>81.91</td>
<td>70.18</td>
<td>73.33</td>
<td>61.97</td>
<td>73.88</td>
<td>72.78</td>
<td><bold>76.35</bold></td>
<td><bold>62.59</bold></td>
<td><bold>85.34</bold></td>
<td><bold>71.08</bold></td>
</tr>
<tr>
<td>3, 4, 4, 2</td>
<td>62.69</td>
<td>53.79</td>
<td>56.82</td>
<td>71.18</td>
<td>74.29</td>
<td>61.52</td>
<td>78.58</td>
<td>70.43</td>
<td>75.60</td>
<td>62.55</td>
<td>81.99</td>
<td>70.13</td>
<td>71.67</td>
<td>60.20</td>
<td>71.18</td>
<td>72.18</td>
<td><bold>77.80</bold></td>
<td><bold>65.25</bold></td>
<td><bold>86.09</bold></td>
<td><bold>70.98</bold></td>
</tr>
<tr>
<td>3, 4, 6, 3</td>
<td>58.35</td>
<td>50.47</td>
<td>49.02</td>
<td>72.03</td>
<td>72.78</td>
<td>59.90</td>
<td>75.73</td>
<td>70.03</td>
<td>73.39</td>
<td>61.32</td>
<td>75.33</td>
<td>71.53</td>
<td>71.49</td>
<td>60.37</td>
<td>70.18</td>
<td>72.83</td>
<td><bold>75.60</bold></td>
<td><bold>63.20</bold></td>
<td><bold>80.54</bold></td>
<td><bold>71.23</bold></td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Interlayer relationships</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-4.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, the inter-layer relationship refers to freezing the pre-trained parameters of different modules within the same model, obtained after training on the image classification task. And then the weight parameters of the unfrozen modules are initialized and trained in the image copy-move tampering detection task, while the generalization ability of the model in image copy-move tampering detection is later measured to achieve the comparison of results in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>

<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows that the model trained directly is more stochastic and less stable compared to the pre-trained ResNet-50. It can be seen that the <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>P</mml:mi></mml:math></inline-formula> values are higher while the 1, 2, 3 layers are frozen and the inter-layers are repeated 3, 4, 4, 2 times.</p>

<p>The experimental results in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> indicate that the proposed IMTNet model achieves better performance in the CMFD task by applying the feature decoupling method, while freezing the first three module weights of ResNet-50.</p>

<p><bold>The relationship between the deep structure of ResNet-50 and CMFD.</bold> In <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, the experimental result of accuracy is weighted and summed in proportions. The experimental datasets are the mixture of generalized experimental data based on MICC-F2000 and MICC-F600 datasets. The comparison experiment is repeated 35 times for each group that is a series of values derived from the same model after 35 experiments under the same conditions. Removing the two highest and the three worst results, the first three modules are frozen to avoid disruptions of the parameter change. The relationship between model depth and the image tampering detection task is revealed by the change in the probability distribution of <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>L</mml:mi><mml:mi>O</mml:mi><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula>, conditioned on the number of repetitions of the residual block in the model&#x2019;s last layer.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Ablation experiment of model deep layers. (a) Last module comparison. (b) Penultimate module comparison. (c) Refined comparison. (d) Final detail comparison</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-5a.tif"/><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-5b.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-5">Fig. 5a</xref> characterizes the relationship between the model generalization ability and the model loss function. It can be seen from the distribution of model layers that model structure 3,4,6,6 does not perform as well as expected while consuming more hardware resources. In addition, the generalizability of the model is also reduced. The problem is mainly due to the model using concrete concepts to represent targets rather than abstract regions in tampered images. Furthermore, model structures 3,4,6,1 have insufficient characterization ability and the upper bound of the generalization error and the loss function error is large, which highlights the weakness of the model characterization ability and the instability of the characterization.</p>

<p><xref ref-type="fig" rid="fig-5">Fig. 5b</xref> illustrates the training of the last layer while keeping the first three modules frozen, and it also involves reducing the number of convolutional layers in the third module. Experimental evidence reveals that the abundance of highly specific semantic features extracted from the middle and deep layers hinders the detection of tampered images.</p>

<p>As demonstrated in <xref ref-type="fig" rid="fig-5">Fig. 5c</xref>, the comparison of the generalization performance of the model structures 3,4,2,3 and 3,4,2,2 reveals that they are comparable. <xref ref-type="fig" rid="fig-5">Fig. 5a</xref>,<xref ref-type="fig" rid="fig-5">c</xref> shows that the deep parameter reduction of the model has little effect on the model&#x2019;s ability in image tampering detection. This leads to the conclusion that lightweighting the model without compromising its representational capabilities can significantly reduce operational resource consumption and enhance computing speed. At present, the preference is given to the model with larger loss as it is believed to possess better generalization ability. Additionally, this model is more compact and consumes fewer resources.</p>

<p>Based on the results depicted in <xref ref-type="fig" rid="fig-5">Fig. 5d</xref>, the stable and well-generalized model structure 3432 has been selected as the optimal structure. By modifying the parameters of the last two layers of the model, the network can better characterize the problem of &#x201C;whether the image has been tampered with&#x201D; with an abstract entity description.</p>

<p><bold>An exploration of the relevant properties of IMTNet.</bold> The mean values from multiple experiments by using MICC-F2000 and MICC-F600 datasets are presented in <xref ref-type="table" rid="table-2">Table 2</xref>. The pre-training datasets we used are ImageNet-1K and ImageNet-21K. The ImageNet dataset is indeed a very common and important source of pre-trained models in the field of transfer learning. The ImageNet-1K is the most commonly used subset of the ImageNet dataset, containing 1000 classes and approximately 1.3 million images. The ImageNet-21K is the full ImageNet dataset, containing 21,841 classes and approximately 14 million images.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Image level generation experiments on IMTNet (1), (2)</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center" colspan="13">Image level generation experiments on IMTNet (1)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Pretained dataset</td>
<td align="center" colspan="12">Fine-tune</td>
</tr>
<tr>
<td/>
<td align="center" colspan="4">4 modules</td>
<td align="center" colspan="4">3, 4 modules</td>
<td align="center" colspan="4">2, 3, 4 modules</td>
</tr>
<tr>
<td></td>
<td><italic>F</italic><sub>1</sub> (%)</td>
<td><italic>ACC</italic> (%)</td>
<td><italic>P</italic> (%)</td>
<td><italic>R</italic> (%)</td>
<td><italic>F</italic><sub>1</sub> (%)</td>
<td><italic>ACC</italic> (%)</td>
<td><italic>P</italic> (%)</td>
<td><italic>R</italic> (%)</td>
<td><italic>F</italic><sub>1</sub> (%)</td>
<td><italic>ACC</italic> (%)</td>
<td><italic>P</italic> (%)</td>
<td><italic>R</italic> (%)</td>
</tr>
<tr>
<td>ImageNet-1K</td>
<td>75.11</td>
<td>62.26</td>
<td>77.21</td>
<td>73.25</td>
<td>78.28</td>
<td>65.29</td>
<td>82.64</td>
<td>74.52</td>
<td>77.56</td>
<td>64.92</td>
<td>82.87</td>
<td>73.22</td>
</tr>
<tr>
<td>ImageNet-21K</td>
<td>76.56</td>
<td>64.13</td>
<td>77.53</td>
<td>74.76</td>
<td>78.60</td>
<td>66.01</td>
<td>82.91</td>
<td>74.83</td>
<td>77.16</td>
<td>63.97</td>
<td>81.94</td>
<td>73.89</td>
</tr>
<tr>
<td align="center" colspan="13">Image level generation experiments on IMTNet (2)</td>
</tr>
<tr>
<td>Pretained dataset</td>
<td align="center" colspan="12">Fine-tune after feature decoupling</td>
</tr>
<tr>
<td/>
<td align="center" colspan="4">Only transfer</td>
<td align="center" colspan="4">2, 3 modules</td>
<td align="center" colspan="4">3 modules</td>
</tr>
<tr>
<td></td>
<td><italic>F</italic><sub>1</sub> (%)</td>
<td><italic>ACC</italic> (%)</td>
<td><italic>P</italic> (%)</td>
<td><italic>R</italic> (%)</td>
<td><italic>F</italic><sub>1</sub> (%)</td>
<td><italic>ACC</italic> (%)</td>
<td><italic>P</italic> (%)</td>
<td><italic>R</italic> (%)</td>
<td><italic>F</italic><sub>1</sub> (%)</td>
<td><italic>ACC</italic> (%)</td>
<td><italic>P</italic> (%)</td>
<td><italic>R</italic> (%)</td>
</tr>
<tr>
<td>ImageNet-1K</td>
<td>78.28</td>
<td>65.29</td>
<td>85.39</td>
<td>73.23</td>
<td>77.49</td>
<td>64.89</td>
<td>84.96</td>
<td>71.93</td>
<td><bold>78.39</bold></td>
<td><bold>66.93</bold></td>
<td><bold>84.79</bold></td>
<td><bold>72.88</bold></td>
</tr>
<tr>
<td>ImageNet-21K</td>
<td>78.60</td>
<td>66.01</td>
<td>85.33</td>
<td>74.96</td>
<td>78.52</td>
<td>65.74</td>
<td>85.44</td>
<td>73.58</td>
<td><bold>79.67</bold></td>
<td><bold>68.10</bold></td>
<td><bold>85.39</bold></td>
<td><bold>75.53</bold></td>
</tr>
</tbody>
</table>
</table-wrap>

<p>From the perspective of the generalization dataset, feature decoupling and fine-tuning are compared in the CMFD task quantitatively. The experimental data is utilized to measure the similarity between the image classification domain and the CMFD domain. Then, the comparison of the number of frozen layers <italic>vs.</italic> the number of fine-tuned layers is made. The experimental results have led to the determination that the most optimal approach involves fine-tuning the third module subsequent to the feature decoupling process.</p>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref> depicts the correlation between the number of training parameters in a model and its representation capabilities. The trainable parameters are derived by subtracting the model&#x2019;s frozen parameters from the model&#x2019;s full parameters, while the model detection accuracy is derived from the model&#x2019;s performance in the generation experiments. This reveals that IMTNet exhibits outstanding generalization performance by employing a trace of parameters.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Generalization accuracy-training parameter</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-6.tif"/>
</fig>
<p>Additionally, <xref ref-type="fig" rid="fig-7">Fig. 7</xref> presents the comparison between feature decoupling and fine-tuning across five distinct dataset sizes: 166, 766, 2046, 4171, and 5337. These datasets are part of the training dataset. After subtracting 5337 images from the training dataset, the test dataset is uniformly the remaining 15,328 images. The highest value is taken in 25 runs. The results of the experiment indicate that, for datasets with less than 1000 samples, fine-tuning is generally a more effective approach compared to feature decoupling followed by fine-tuning. <xref ref-type="fig" rid="fig-7">Fig. 7a</xref>,<xref ref-type="fig" rid="fig-7">b</xref> portrays the relationship from different perspectives.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Comparative analysis of fine-tuning and feature decoupling. (a) F<sub>1</sub>-images number. (b) Accuracy-images number</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-7.tif"/>
</fig>
<p>The following observations, derived from the comparison of the results obtained from these experiments, have been listed below: (1) The findings from the quantitative analysis experiments indicate that the CMFD task necessitates shallower layers for the ResNet-50 model as compared to the image classification task. (2) Fine-tuning achieves superior performance compared to feature decoupling when the dataset is limited in size. (3) For the ResNet-50 model, the CMFD task achieves better performance when reducing the number of convolutional layers in the third module, compared to reducing the number of convolutional layers in the last module.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Anti-Attack Experiments</title>
<p>The anti-attack experiments of various algorithms are described below to evaluate the robustness of the pre-trained network. 1: Rotation attack. 2: Gaussian noise attack. 3: JPEG image compression attack. 4: Blurring attack. 5: Scaling attack.</p>
<p>As shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the value is used as an indicator to evaluate the detection capabilities of various methods in different tampering environments. The curve depicted is obtained through the quadratic interpolation of the relevant data points. In which, <xref ref-type="fig" rid="fig-8">Fig. 8a</xref>&#x2013;<xref ref-type="fig" rid="fig-8">e</xref> show the result comparisons under various attacks. In anti-attack experiments, the following observations are presented: (1) under a variety of attacks, IMTNet exhibits excellent robustness compared to other methods. (2) IMTNet possesses the property conferred by feature decoupling as seen in the comparison of ResNet-50.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Resistance to attack performance. (a) Rotation attack. (b) Noise attack. (c) JPEG compression attack. (d) Blur attack. (e) Scaling attack</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-8.tif"/>
</fig>
<p>To test the performance of IMTNet, MICC-F2000, MICC-F600 and COVERAGE datasets are used to conduct generalization tests. <xref ref-type="table" rid="table-3">Table 3</xref> shows the value of <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, accuracy (<inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula>), Precision (<inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>P</mml:mi></mml:math></inline-formula>) and Recall (<inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>R</mml:mi></mml:math></inline-formula>).</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Image generation experiments</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Pretained<break/>dataset</th>
<th align="center" colspan="4">MICC-F2000</th>
<th align="center" colspan="4">MICC-F600</th>
<th align="center" colspan="4">COVERAGE</th>
<th align="center" colspan="4">DEFACTO</th>
</tr>
<tr>
<th><italic>F</italic><sub>1</sub>(%)</th>
<th><italic>ACC</italic>(%)</th>
<th><italic>P</italic>(%)</th>
<th><italic>R</italic>(%)</th>
<th><italic>F</italic><sub>1</sub>(%)</th>
<th><italic>ACC</italic>(%)</th>
<th><italic>P</italic>(%)</th>
<th><italic>R</italic>(%)</th>
<th><italic>F</italic><sub>1</sub>(%)</th>
<th><italic>ACC</italic>(%)</th>
<th><italic>P</italic>(%)</th>
<th><italic>R</italic>(%)</th>
<th><italic>F</italic><sub>1</sub>(%)</th>
<th><italic>ACC</italic>(%)</th>
<th><italic>P</italic>(%)</th>
<th><italic>R</italic>(%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>ResNet-50 [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>54.29</td>
<td>46.48</td>
<td>45.42</td>
<td>67.48</td>
<td>59.63</td>
<td>51.76</td>
<td>48.57</td>
<td>77.18</td>
<td>51.91</td>
<td>46.28</td>
<td>57.97</td>
<td>46.97</td>
<td>52.11</td>
<td>46.53</td>
<td>48.32</td>
<td>56.57</td>
</tr>
<tr>
<td>Zhong [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>48.47</td>
<td>46.33</td>
<td>50.52</td>
<td>48.53</td>
<td>46.67</td>
<td>47.55</td>
<td>47.68</td>
<td>41.83</td>
<td>41.93</td>
<td>46.36</td>
<td>49.78</td>
<td>55.05</td>
<td>47.98</td>
<td>50.44</td>
<td>52.57</td>
<td>52.03</td>
</tr>
<tr>
<td>Chen [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>62.16</td>
<td>54.49</td>
<td>53.43</td>
<td>74.33</td>
<td>67.91</td>
<td>60.15</td>
<td>57.47</td>
<td>82.94</td>
<td>56.35</td>
<td>52.63</td>
<td>61.18</td>
<td>52.22</td>
<td>54.92</td>
<td>50.17</td>
<td>50.37</td>
<td>60.38</td>
</tr>
<tr>
<td>Priyanka [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>47.34</td>
<td>46.67</td>
<td>48.59</td>
<td>46.78</td>
<td>50.09</td>
<td>49.56</td>
<td>47.46</td>
<td>48.29</td>
<td>49.01</td>
<td>49.92</td>
<td>50.03</td>
<td>51.49</td>
<td>49.98</td>
<td>41.29</td>
<td>45.69</td>
<td>46.46</td>
</tr>
<tr>
<td>IMTNet</td>
<td><bold>78.66</bold></td>
<td><bold>66.47</bold></td>
<td><bold>86.24</bold></td>
<td><bold>72.93</bold></td>
<td><bold>81.12</bold></td>
<td><bold>70.40</bold></td>
<td><bold>86.74</bold></td>
<td><bold>76.18</bold></td>
<td><bold>65.74</bold></td>
<td><bold>58.83</bold></td>
<td><bold>75.03</bold></td>
<td><bold>59.27</bold></td>
<td><bold>63.54</bold></td>
<td><bold>57.94</bold></td>
<td><bold>60.78</bold></td>
<td><bold>66.53</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The algorithm in [<xref ref-type="bibr" rid="ref-19">19</xref>] introduces a classification task based on ResNet-50, the first row in <xref ref-type="table" rid="table-3">Table 3</xref> is the result of CMFD task by training directly on ResNet-50, it can be seen that applying the classification task model directly to the CMFD task does not work well. However, the accuracy and <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> values of the proposed IMTNet combining feature decoupling and transfer learning are higher than other algorithms in different datasets. Therefore, it is believed that IMTNet exhibits better generalization than other proprietary methods. The outcomes presented in <xref ref-type="fig" rid="fig-9">Fig. 9</xref> showcase the robust localization proficiency of IMTNet in generalization experiments. The 1st and 2nd rows show the forgery images and corresponding detection results.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>The CMFD results of IMTNet. a(1)&#x007E;e(1) The forgery images. a(2)&#x007E;e(2) Corresponding detection results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_53740-fig-9.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Conclusions</title>
<p>In the proposed scheme, the relationship between ResNet-50 and CMFD is thoroughly demonstrated through quantitative experiments. IMTNet is proposed by leveraging the image classification feature domains and reducing the deep architecture of ResNet-50. Firstly, the relationship between CMFD and deep network architecture is formulated as an optimization problem. In the CMFD task, IMTNet exhibits outstanding performance compared to ResNet-50 and other CMFD algorithms by reducing the deep structure of ResNet-50 and utilizing the feature decoupling method. Secondly, experiments demonstrate that the IMTNet reduced the number of ResNet-50 parameters while enhancing the generalization capability of the model. Furthermore, the integration of MFPN improved the capability of the proposed method in detecting suspicious areas in tampered images.</p>
</sec>
</body>
<back>
<ack><p>None.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported and founded by the Guizhou Provincial Science and Technology Project under the Grant No. QKH-Basic-ZK[2021]YB311, the Youth Science and Technology Talent Growth Project of Guizhou Provincial Education Department under Grant No. QJH-KY-ZK[2021]132, the Guizhou Provincial Science and Technology Project under the Grant No. QKH-Basic-ZK[2021]YB319, the National Natural Science Foundation of China (NSFC) under Grant 61902085, and the Key Laboratory Program of Blockchain and Fintech of Department of Education of Guizhou Province (2023-014).</p>
</sec>
<sec><title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Huan Wang, Zhongyuan Jiang; data collection: Huan Wang, Qing Qian and Yong Long; analysis and interpretation of results: Huan Wang, Zhongyuan Jiang and Qing Qian; draft manuscript preparation: Huan Wang, Zhongyuan Jiang and Yong Long. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are available from the corresponding author, <email>jiangzydj@163.com</email>, upon reasonable request.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Fridrich</surname></string-name>, <string-name><given-names>B. D.</given-names> <surname>Soukal</surname></string-name>, and <string-name><given-names>A. J.</given-names> <surname>Lukas</surname></string-name></person-group>, &#x201C;<article-title>Detection of copy-move forgery in digital images</article-title>,&#x201D; in <conf-name>Proc. Digit. Forensic Res. Workshop</conf-name>, <publisher-loc>Cleveland, OH, USA</publisher-loc>, <year>2003</year>, pp. <fpage>19</fpage>&#x2013;<lpage>23</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. P.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wei</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Long</surname></string-name></person-group>, &#x201C;<article-title>Improved DCT-based detection of copy-move forgery in images</article-title>,&#x201D; <source>Forensic Sci. Int.</source>, vol. <volume>206</volume>, no. <issue>3</issue>, pp. <fpage>178</fpage>&#x2013;<lpage>184</lpage>, <year>2011</year>. doi: <pub-id pub-id-type="doi">10.1016/j.forsciint.2010.08.001</pub-id>; <pub-id pub-id-type="pmid">20832208</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Priyanka</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Singh</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Singh</surname></string-name></person-group>, &#x201C;<article-title>An improved block based copy-move forgery detection technique</article-title>,&#x201D; <source>Multimed. Tools Appl.</source>, vol. <volume>79</volume>, no. <issue>19</issue>, pp. <fpage>13011</fpage>&#x2013;<lpage>13035</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1007/s11042-019-08354-x</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Akram</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Rashid</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jaffar</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Hajjej</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Iqbal</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Sarwar</surname></string-name></person-group>, &#x201C;<article-title>Weber law based approach for multi-class image forgery detection</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>78</volume>, pp. <fpage>145</fpage>&#x2013;<lpage>166</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2023.041074</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Amerini</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ballan</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Caldelli</surname></string-name>, <string-name><given-names>A.</given-names> <surname>BDel Bimbo</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Serra</surname></string-name></person-group>, &#x201C;<article-title>A SIFT-based forensic method for copy-move attack detection and transformation recovery</article-title>,&#x201D; <source>IEEE Trans. Inf. Forensics Secur.</source>, vol. <volume>6</volume>, no. <issue>3</issue>, pp. <fpage>1099</fpage>&#x2013;<lpage>1110</lpage>, <year>2011</year>. doi: <pub-id pub-id-type="doi">10.1109/TIFS.2011.2129512</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Gan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhong</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Vong</surname></string-name></person-group>, &#x201C;<article-title>A novel copy-move forgery detection algorithm via feature label matching and hierarchical segmentation filtering</article-title>,&#x201D; <source>Inf. Process. Manage.</source>, vol. <volume>59</volume>, no. <issue>1</issue>, pp. <fpage>167</fpage>&#x2013;<lpage>178</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1016/j.ipm.2021.102783</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. L.</given-names> <surname>Zhong</surname></string-name>, <string-name><given-names>J. X.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zeng</surname></string-name>, and <string-name><given-names>Y. Q.</given-names> <surname>Zhao</surname></string-name></person-group>, &#x201C;<article-title>A novel image copy-move forgery detection algorithm using the characteristics of local descriptors</article-title>,&#x201D; <source>Int. J. Pattern Recognit. Artif. Intell.</source>, vol. <volume>36</volume>, no. <issue>15</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1142/S0218001422540192</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Ayd&#x0131;n</surname></string-name></person-group>, &#x201C;<article-title>Automated identification of copy-move forgery using Hessian and patch feature extraction techniques</article-title>,&#x201D; <source>J. Forensic Sci.</source>, vol. <volume>69</volume>, no. <issue>1</issue>, pp. <fpage>131</fpage>&#x2013;<lpage>138</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.1111/1556-4029.15415</pub-id>; <pub-id pub-id-type="pmid">37888436</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Yue</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Abd-Almageed</surname></string-name>, and <string-name><given-names>P.</given-names> <surname>Natarajan</surname></string-name></person-group>, &#x201C;<article-title>Busternet: Detecting copy-move image forgery with source/target localization</article-title>,&#x201D; in <conf-name>Proc. Eur. Conf. on Comput. Vis. (ECCV)</conf-name>, <publisher-loc>Munich, Germany</publisher-loc>, <year>2018</year>, pp. <fpage>170</fpage>&#x2013;<lpage>186</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-030-01231-1_11</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Coatrieux</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zheng</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Shafique</surname></string-name></person-group>, &#x201C;<article-title>A serial image copy-move forgery localization <italic>scheme</italic> with source/target distinguishment</article-title>,&#x201D; <source>IEEE Trans. Multimedia</source>, vol. <volume>23</volume>, pp. <fpage>3506</fpage>&#x2013;<lpage>3517</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TMM.2020.3026868</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. L.</given-names> <surname>Zhong</surname></string-name> and <string-name><given-names>C. M.</given-names> <surname>Pun</surname></string-name></person-group>, &#x201C;<article-title>An end-to-end dense-inceptionnet for image copy-move forgery detection</article-title>,&#x201D; <source>IEEE Trans. Inf. Forensics Secur.</source>, vol. <volume>15</volume>, pp. <fpage>2134</fpage>&#x2013;<lpage>2146</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1109/TIFS.2019.2957693</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Chaudhuri</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Nevatia</surname></string-name></person-group>, &#x201C;<article-title>SPAN: Spatial pyramid attention network for image manipulation localization</article-title>,&#x201D; <year>2020</year>, <comment><italic>arXiv:2009.00726</italic></comment>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F. Z.</given-names> <surname>Mehrjardi</surname></string-name>, <string-name><given-names>A. M.</given-names> <surname>Latif</surname></string-name>, and <string-name><given-names>M. S.</given-names> <surname>Zarchi</surname></string-name></person-group>, &#x201C;<article-title>Copy-move forgery detection and localization using deep learning</article-title>,&#x201D; <source>Int. J. Pattern Recognit. Artif. Intell.</source>, vol. <volume>37</volume>, no. <issue>9</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>21</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1142/S0218001423520122</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. K.</given-names> <surname>Thyagharajan</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Nirmala</surname></string-name></person-group>, &#x201C;<article-title>Image manipulation detection through laterally linked pixels and kernel algorithms</article-title>,&#x201D; <source>Comput. Syst. Sci. Eng.</source>, vol. <volume>41</volume>, no. <issue>1</issue>, pp. <fpage>357</fpage>&#x2013;<lpage>371</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.32604/csse.2022.020258</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F. Z.</given-names> <surname>Zhuang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A comprehensive survey on transfer learning</article-title>,&#x201D; <source>Proc. IEEE</source>, vol. <volume>109</volume>, no. <issue>1</issue>, pp. <fpage>43</fpage>&#x2013;<lpage>76</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/JPROC.2020.3004555</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Abbas</surname></string-name></person-group>, &#x201C;<article-title>A survey on deep learning and its applications</article-title>,&#x201D; <source>Comput. Sci. Rev.</source>, vol. <volume>40</volume>, no. <issue>5</issue>, <year>2021</year>, Art. no. <comment>100379</comment>. doi: <pub-id pub-id-type="doi">10.1016/j.cosrev.2021.100379</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Ran</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Xu</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Guo</surname></string-name></person-group>, &#x201C;<article-title>Decoupling makes weakly supervised local feature better</article-title>,&#x201D; <year>2022</year>, <comment><italic>arXiv:2201.02861</italic></comment>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Ran</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Lu</surname></string-name></person-group>, &#x201C;<article-title>Feature decoupling and reorganization network for single image deraining</article-title>,&#x201D; <source>Multimed. Syst.</source>, vol. <volume>30</volume>, <year>2024</year>, Art. no. <comment>154</comment>. doi: <pub-id pub-id-type="doi">10.1007/s00530-024-01348-2</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ren</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Deep residual learning for image recognition</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Comput. Vis. Pattern Recognit.</conf-name>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>2016</year>, pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J. H.</given-names> <surname>Liew</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zou</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Zhou</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Feng</surname></string-name></person-group>, &#x201C;<article-title>PANet: Few-shot image semantic segmentation with prototype alignment</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. on Comput. Vis. Pattern Recognit.</conf-name>, <publisher-loc>COEX Convention Center, Seoul, Republic of Korea</publisher-loc>, <year>2019</year>, pp. <fpage>9197</fpage>&#x2013;<lpage>9206</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICCV.2019.00929</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Y. X.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Qiao</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Xiang</surname></string-name></person-group>, &#x201C;<article-title>Domain adaptive ensemble learning</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>30</volume>, pp. <fpage>8008</fpage>&#x2013;<lpage>8018</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2021.3112012</pub-id>; <pub-id pub-id-type="pmid">34534081</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J. S.</given-names> <surname>Beis</surname></string-name> and <string-name><given-names>G. D.</given-names> <surname>Lowe</surname></string-name></person-group>, &#x201C;<article-title>Shape indexing using approximate nearest-neighbor search in high-dimensional spaces</article-title>,&#x201D; in <conf-name>Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit.</conf-name>, <publisher-loc>San Juan, PR, USA</publisher-loc>, <year>1997</year>, pp. <fpage>1000</fpage>&#x2013;<lpage>1006</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.1997.609451</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Yu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Unleashing the power of multi-task learning: A comprehensive survey spanning traditional, deep, and pretrained foundation model Eras</article-title>,&#x201D; <year>2024</year>, <comment><italic>arXiv:2404.18961</italic></comment>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>T. N.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<article-title>CASIA image tampering detection evaluation database</article-title>,&#x201D; in <conf-name>Proc. IEEE China Summit &#x0026; Int. Conf. on Signal and Inf. Process.</conf-name>, <publisher-loc>Beijing, China</publisher-loc>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1109/ChinaSIP.2013.6625374</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Tralic</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Zupancic</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Grgic</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Grgic</surname></string-name></person-group>, &#x201C;<article-title>CoMoFoD-new database for copy-move forgery detection</article-title>,&#x201D; in <conf-name>Proc. Int. Symp. on Electron. in Mar. (ELMAR)</conf-name>, <publisher-loc>Zadar, Croatia</publisher-loc>, <year>2013</year>, pp. <fpage>49</fpage>&#x2013;<lpage>54</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Amerini</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ballan</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Caldelli</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Del Bimbo</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Del Tongo</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Serra</surname></string-name></person-group>, &#x201C;<article-title>Copy-move forgery detection and localization by means of robust clustering with J-Linkage</article-title>,&#x201D; <source>Signal Process.: Image Commun.</source>, vol. <volume>28</volume>, no. <issue>6</issue>, pp. <fpage>659</fpage>&#x2013;<lpage>669</lpage>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1016/j.image.2013.03.006</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Wen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Subramanian</surname></string-name>, <string-name><given-names>T. T.</given-names> <surname>Ng</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Shen</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Winkler</surname></string-name></person-group>, &#x201C;<article-title>COVERAGE&#x2014;A novel database for copy-move forgery detection</article-title>,&#x201D; in <conf-name>Proc. Int. Conf. on Inf. Photonics(ICIP)</conf-name>, <publisher-loc>Phoenix, AZ, USA</publisher-loc>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1109/ICIP.2016.7532339</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Mahfoudi</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Tajini</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Retraint</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Morain-Nicolier</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Dugelay</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Pic</surname></string-name></person-group>, &#x201C;<article-title>DEFACTO: Image and face manipulation dataset</article-title>,&#x201D; in <conf-name>Proc. 27Th Eur. Sig. Process. Conf. (EUSIPCO)</conf-name>, <publisher-loc>Coruna, Spain</publisher-loc>, <year>2019</year>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi: <pub-id pub-id-type="doi">10.23919/EUSIPCO.2019.8903181</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Ardizzone</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Bruno</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Mazzola</surname></string-name></person-group>, &#x201C;<article-title>Copy-move forgery detection by matching triangles of keypoints</article-title>,&#x201D; <source>IEEE Trans. Inf. Forensics Secur.</source>, vol. <volume>10</volume>, no. <issue>10</issue>, pp. <fpage>2084</fpage>&#x2013; <lpage>2094</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TIFS.2015.2445742</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>