<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">81922</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.081922</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>HiFraud: Hierarchical Privacy-Preserving Federated Learning with Star-Chain Knowledge Transfer for Cross-Institutional Fraud Detection</article-title>
<alt-title alt-title-type="left-running-head">HiFraud: Hierarchical Privacy-Preserving Federated Learning with Star-Chain Knowledge Transfer for Cross-Institutional Fraud Detection</alt-title>
<alt-title alt-title-type="right-running-head">HiFraud: Hierarchical Privacy-Preserving Federated Learning with Star-Chain Knowledge Transfer for Cross-Institutional Fraud Detection</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0006-2759-4971</contrib-id>
<name name-style="western"><surname>Zhang</surname><given-names>Zhihao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="author-notes" rid="afn1">#</xref></contrib>
<contrib id="author-2" contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0009-9535-2941</contrib-id>
<name name-style="western"><surname>Liu</surname><given-names>Zhuodong</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="author-notes" rid="afn1">#</xref></contrib>
<contrib id="author-3" contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-4226-1255</contrib-id>
<name name-style="western"><surname>Li</surname><given-names>Xiangyu</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-7217-6919</contrib-id>
<name name-style="western"><surname>Zhang</surname><given-names>Lei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>zhlei@bjtu.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>School of Economics and Management, Beijing Jiaotong University</institution>, <addr-line>Beijing</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Electronic Engineering, Shanghai Jiao Tong University</institution>, <addr-line>Shanghai</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Lei Zhang. Email: <email>zhlei@bjtu.edu.cn</email></corresp>
<fn id="afn1">
<p><sup>#</sup>These authors contributed equally to this work</p>
</fn>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>34</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_81922.pdf"></self-uri>
<abstract>
<p>Financial fraud detection across institutions faces a fundamental tension between the need for diverse training data and regulatory prohibitions on sharing sensitive records. Existing federated learning approaches suffer from performance degradation under non-IID distributions and substantial utility losses when uniform differential privacy is applied to inherently sparse fraud signals. To this end, this paper proposes HiFraud, a hierarchical federated framework featuring three key components: fraud-aware dynamic clustering with complementarity regularization to group institutions by fraud pattern similarity while preserving rare-type representation; star-chain knowledge transfer augmented by not-true-class distillation to propagate novel fraud patterns rapidly within clusters while mitigating catastrophic forgetting; and privacy-adaptive aggregation via R&#x00E9;nyi differential privacy composition, calibrating noise intensity to distributional divergence and fraud rarity. Experiments on IEEE-CIS, PaySim, and Worldline datasets show that HiFraud achieves an area under the receiver operating characteristic curve (AUC-ROC) of 0.935 under <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2.3</mml:mn></mml:math></inline-formula>, outperforming DP-FedAvg by 10.5% while reducing convergence from 49 to 30 rounds. The framework also suppresses membership inference attack success to 10.2%, detects emerging fraud patterns within 3 h inside clusters, and improves rare fraud type detection by 23.0% over uniform privacy baselines. These results demonstrate that hierarchical architectures can effectively reconcile detection performance, formal privacy guarantees, and rapid threat response in collaborative fraud detection.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Hierarchical federated learning</kwd>
<kwd>differential privacy</kwd>
<kwd>fraud detection</kwd>
<kwd>star-chain transfer</kwd>
<kwd>knowledge distillation</kwd>
<kwd>adaptive clustering</kwd>
<kwd>non-IID data</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Financial fraud poses a persistent and escalating threat to banking, e-commerce, and digital payment ecosystems, where rapid anomaly identification is essential to preventing substantial economic losses [<xref ref-type="bibr" rid="ref-1">1</xref>]. In 2024, the U.S. Federal Trade Commission reported &#x00024;12.5 billion in consumer fraud losses, representing a 25% increase over the prior year, with investment scams alone accounting for &#x00024;5.7 billion [<xref ref-type="bibr" rid="ref-2">2</xref>]. This growth has intensified the demand for detection systems capable of identifying diverse and evolving fraud patterns across institutional boundaries. However, privacy regulations such as the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA) strictly prohibit the sharing of raw transaction data between organizations [<xref ref-type="bibr" rid="ref-3">3</xref>]. This creates a fundamental tension: effective fraud detection requires large-scale, heterogeneous datasets, yet the very data needed to build robust models cannot leave institutional boundaries.</p>
<p>Federated learning (FL) has emerged as a promising paradigm for enabling collaborative model training without centralizing sensitive data [<xref ref-type="bibr" rid="ref-4">4</xref>]. By allowing institutions to jointly optimize a shared model through the exchange of model parameters rather than raw data, FL offers a potential resolution to the privacy&#x2013;utility dilemma in fraud detection. Nevertheless, the direct application of standard FL algorithms to cross-institutional fraud detection introduces two compounding challenges. First, different institutions face fundamentally distinct fraud types&#x2014;credit card skimming at banks, account takeovers at e-commerce platforms, and money laundering at payment processors&#x2014;resulting in extreme non-IID (non-independent and identically distributed) data distributions that violate the convergence assumptions underlying standard federated algorithms [<xref ref-type="bibr" rid="ref-5">5</xref>]. Second, fraudulent transactions are inherently rare events, with some institutions reporting fraud rates below 0.1%, rendering isolated local training insufficient to learn discriminative representations for minority-class patterns [<xref ref-type="bibr" rid="ref-6">6</xref>]. The simultaneous presence of distributional heterogeneity and extreme class imbalance demands architectural innovations beyond incremental adaptations of existing methods.</p>
<sec id="s1_1">
<label>1.1</label>
<title>Federated Fraud Detection: Progress and Limitations</title>
<p>Early federated fraud detection work established the viability of collaborative training: Yang et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] achieved an AUC of 95.5% with the first federated credit card fraud framework, while Abdul Salam et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] introduced hybrid resampling to address class imbalance. However, the former assumes homogeneous fraud distributions and the latter lacks formal privacy guarantees. Subsequent studies have tackled class imbalance more systematically: Shah et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] and Hilou et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] applied client-side synthetic minority over-sampling technique (SMOTE) variants before federated aggregation, Farooq et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] integrated adaptive aggregation with privacy-preserving mechanisms, Sarkar et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] proposed Fed-Focal Loss to focus gradient updates on difficult fraud instances, and Wang et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] developed Ratio Loss to dynamically counteract local&#x2013;global imbalance mismatch without direct data access.</p>
<p>Despite these advances, three structural limitations persist across existing federated fraud detection systems. First, all aforementioned frameworks adopt flat federated architectures in which every participant is treated identically during aggregation, failing to exploit the natural similarities in fraud patterns among subsets of institutions. Second, class imbalance mitigation remains confined to the client level, without federated-layer coordination that could leverage cross-institutional knowledge about rare fraud types. Third, the integration of differential privacy&#x2014;essential for regulatory compliance&#x2014;typically imposes uniform noise that disproportionately degrades the detection of already-sparse fraud signals, creating a privacy&#x2013;utility trade-off that existing approaches have not satisfactorily resolved. Recent work has begun to address these limitations from different angles. For instance, Aljunaid et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed an explainable AI-driven federated learning model for financial fraud detection that integrates secure aggregation with interpretable decision mechanisms, demonstrating the growing recognition that collaborative architectures must balance privacy, transparency, and detection accuracy in banking environments. However, their approach does not address hierarchical aggregation, dynamic clustering based on fraud pattern semantics, or the interplay between adaptive differential privacy and knowledge transfer that is central to our work.</p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Hierarchical Architectures and Knowledge Transfer in Federated Learning</title>
<p>Hierarchical federated learning introduces intermediate aggregation layers to address the communication overhead and data heterogeneity of large-scale federated systems. Liu et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] proposed a cloud-edge-client architecture reducing communication costs by 60%, and comprehensive reviews have further examined cloud&#x2013;edge&#x2013;end collaboration for privacy-preserving AI [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. Clustered federated learning refines this paradigm by grouping clients according to distributional similarity: Sattler et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] demonstrated that model-agnostic clustering substantially improves convergence on non-IID data, while Gong et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] achieved a 9.2% accuracy gain with adaptive cluster scheduling. On the clustering criteria side, Duan et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] and Ali et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] independently developed dynamic clustering frameworks supporting client migration in response to distribution drift, and Islam et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a weight-based one-shot method achieving up to 45% accuracy gains. However, Yang et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] identified a critical limitation: purely similarity-driven clustering may marginalize clients with rare patterns into low-influence groups, suggesting that clustering objectives should balance similarity with complementarity&#x2014;an insight directly relevant to fraud detection.</p>
<p>In the knowledge transfer dimension, Wang et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] provided theoretical and empirical evidence that sequential model passing achieves faster convergence under extreme non-IID conditions than parallel averaging, a finding further validated by Yan et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] within hierarchical architectures. Xie et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] proposed StarCPFL, combining centralized model distribution with chain-style sequential refinement. Despite these advances, sequential transfer inherits a well-known vulnerability: catastrophic forgetting, where training on subsequent clients overwrites previously accumulated knowledge [<xref ref-type="bibr" rid="ref-27">27</xref>]. Knowledge distillation has emerged as the predominant countermeasure. Lee et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed FedNTD, distilling knowledge on non-true classes to preserve discriminative capacity for absent categories, while He et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] introduced selective self-distillation conditioned on teacher confidence. Arafeh et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] further designed a warmup-based protocol to reduce forgetting during sequential initialization. These techniques provide a mature foundation for integrating distillation into sequential transfer, yet this combination has not been explored in cross-institutional fraud detection.</p>
<p>It is important to note how HiFraud&#x2019;s star-chain mechanism fundamentally differs from existing approaches. Unlike StarCPFL [<xref ref-type="bibr" rid="ref-26">26</xref>], which applies star-chain communication in a general personalized federated learning setting with layer-wise clustering, HiFraud introduces three domain-specific innovations: (i) the star institution is selected based on a fraud-rate-adjusted performance metric (<xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>) rather than simple accuracy, ensuring that institutions capable of detecting fraud under scarcity lead knowledge propagation; (ii) the chain ordering is determined by distributional similarity to the star in fraud pattern space, creating a curriculum-like transfer path that progressively adapts to increasingly diverse fraud distributions; and (iii) not-true-class distillation is integrated at each chain step specifically to preserve knowledge of fraud types absent from the current institution, which is critical in fraud detection where each institution may observe only a subset of fraud categories. These design choices transform the general-purpose star-chain topology into a fraud-aware knowledge propagation mechanism with theoretical and empirical advantages over flat aggregation in non-IID fraud settings.</p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>Privacy Preservation and Adversarial Robustness</title>
<p>Differential privacy (DP) remains the dominant formal privacy framework in federated learning. The moments accountant [<xref ref-type="bibr" rid="ref-31">31</xref>] enabled tight privacy loss tracking, while R&#x00E9;nyi differential privacy (RDP) [<xref ref-type="bibr" rid="ref-32">32</xref>] further tightened multi-round composition bounds. However, uniform noise injection poses a particular challenge for fraud detection: because fraud signals are inherently sparse, flat noise mechanisms disproportionately obscure the patterns the model needs to learn, as demonstrated by Truex et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] for local DP on imbalanced datasets. Adaptive mechanisms have been proposed to address this limitation: Xue et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] dynamically adjusted clipping thresholds based on gradient norms, Yuan et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] introduced amplitude-varying perturbation that reduces noise in later training stages, and Lin et al. [<xref ref-type="bibr" rid="ref-36">36</xref>] formalized the M<sup>2</sup>FDP framework decomposing privacy contributions across multi-tier networks. Nevertheless, a unified composition theorem tailored to hierarchical adaptive mechanisms remains an open challenge.</p>
<p>Beyond formal guarantees, federated systems must withstand practical attacks. Bai et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] surveyed membership inference attacks (MIA) in FL, concluding that standard configurations are vulnerable to success rates well above random chance, while Deng and Yang [<xref ref-type="bibr" rid="ref-38">38</xref>] showed that composite defenses combining gradient compression, selective sharing, and regularization can suppress MIA accuracy to below 38% with minimal task degradation. On the robustness front, Li et al. [<xref ref-type="bibr" rid="ref-39">39</xref>] demonstrated that classical Byzantine-robust aggregation methods degrade under non-IID distributions, motivating hierarchical solutions such as the two-tier Byzantine-resilient scheme of Nordlund et al. [<xref ref-type="bibr" rid="ref-40">40</xref>] and the cross-device protocol of Liu et al. [<xref ref-type="bibr" rid="ref-41">41</xref>]. However, neither MIA defenses nor Byzantine robustness mechanisms have been systematically integrated with hierarchical federated architectures for fraud detection, leaving a significant design gap.</p>
<p>The privacy-preserving mechanism in HiFraud is designed to defend against three specific categories of privacy threats that are particularly relevant to cross-institutional fraud detection. First, <italic>model inversion attacks</italic>, in which an adversary attempts to reconstruct sensitive transaction features from shared model parameters, are mitigated through the Gaussian noise mechanism applied during both star distribution (<xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>) and chain transfer (<xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>), which ensures that the transmitted parameters do not reveal individual transaction characteristics. Second, <italic>gradient leakage attacks</italic>, where an attacker infers training data from observed gradient updates, are countered by gradient clipping to sensitivity bound <italic>S</italic> combined with calibrated noise injection at each local adaptation step, following the Gaussian mechanism framework of Abadi et al. [<xref ref-type="bibr" rid="ref-31">31</xref>]. Third, <italic>membership inference attacks</italic>, in which an adversary determines whether a specific transaction was used in training, are addressed through the hierarchical aggregation structure that limits external visibility to cluster-level models rather than individual institutional updates, compounded by the adaptive noise calibration that provides stronger protection for institutions with distinctive fraud patterns. The formal privacy guarantee (Theorem 1) establishes that these layered defenses collectively satisfy <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>-differential privacy under R&#x00E9;nyi composition.</p>
</sec>
<sec id="s1_4">
<label>1.4</label>
<title>Our Approach and Contributions</title>
<p>The preceding analysis reveals a fragmented research landscape: clustered federated learning benefits non-IID data but has not been designed around fraud pattern semantics; sequential knowledge transfer offers convergence advantages but lacks privacy guarantees and forgetting mitigation; and adaptive differential privacy improves the privacy&#x2013;utility trade-off but has not been coupled with hierarchical architectures. No existing framework unifies these individually mature components into a coherent system tailored to cross-institutional fraud detection.</p>
<p>This paper proposes HiFraud, a hierarchical privacy-preserving federated learning framework that addresses these challenges through a three-layer architecture integrating fraud-aware clustering, star-chain knowledge transfer with distillation-based forgetting mitigation, and privacy-adaptive aggregation grounded in R&#x00E9;nyi differential privacy composition. The main contributions of this work are as follows:<list list-type="bullet">
<list-item>
<p>We propose a three-layer hierarchical architecture that combines fraud-aware dynamic clustering, intra-cluster star-chain transfer learning, and privacy-adaptive global aggregation. The framework achieves an AUC-ROC of 0.935 under <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>&#x03B5;</mml:mi><mml:mo>=</mml:mo><mml:mn>2.3</mml:mn></mml:math></inline-formula> differential privacy, outperforming standard DP-FedAvg by 10.5% while reducing convergence rounds from 49 to 30.</p></list-item>
<list-item>
<p>We develop a star-chain knowledge transfer mechanism augmented with not-true-class distillation to mitigate catastrophic forgetting during sequential model passing. This mechanism enables detection of novel fraud patterns within 3 h inside clusters, compared to 24 h for flat federated architectures.</p></list-item>
<list-item>
<p>We introduce a fraud-pattern-specific dynamic clustering strategy that balances distributional similarity with complementarity, preventing the marginalization of institutions with rare fraud types. This design improves detection performance for rare fraud categories by 18% compared to static geographic clustering.</p></list-item>
<list-item>
<p>We design a hierarchical adaptive privacy allocation scheme based on R&#x00E9;nyi differential privacy composition, calibrating noise intensity according to both distributional divergence and fraud pattern rarity. This approach reduces overall privacy budget consumption by 35% compared to uniform allocation while maintaining equivalent formal guarantees.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> presents the HiFraud framework, detailing the fraud-aware dynamic clustering mechanism, the star-chain knowledge transfer with not-true-class distillation, and the privacy-adaptive aggregation scheme with formal privacy and convergence guarantees. <xref ref-type="sec" rid="s3">Section 3</xref> provides comprehensive experimental evaluations on three benchmark datasets, including comparisons with state-of-the-art baselines, ablation studies, privacy&#x2013;utility analysis, and scalability assessments. <xref ref-type="sec" rid="s4">Section 4</xref> discusses the practical implications of the results and identifies limitations alongside directions for future research. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> concludes the paper.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Methodology</title>
<p>This section presents the design of HiFraud, a hierarchical federated learning framework for cross-institutional fraud detection. The framework adopts a three-layer architecture in which participating institutions are first grouped into clusters based on fraud pattern similarity, then engage in intra-cluster knowledge transfer through a star-chain mechanism augmented with distillation-based forgetting mitigation, and finally contribute to global model refinement via privacy-adaptive aggregation grounded in R&#x00E9;nyi differential privacy. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> provides an overview of the complete architecture. The global coordination layer manages cross-cluster aggregation and privacy budget allocation. Each cluster coordination layer facilitates star-chain knowledge transfer among its member institutions. At the institutional layer, local models are trained on private transaction data with adaptive differential privacy protection. The interplay among these three layers enables efficient knowledge sharing between similar institutions while maintaining formal privacy guarantees across the entire system. As illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the communication flow proceeds as follows: (1) at Layer 1, each institution trains locally on its private dataset <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> with adaptive DP noise <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> calibrated to its distributional divergence; (2) at Layer 2, the star institution (marked with <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mo>&#x22C6;</mml:mo></mml:math></inline-formula>) in each cluster broadcasts its model to all cluster members via solid arrows (star distribution), after which models are sequentially refined along dashed arrows (chain enhancement) with not-true-class distillation at each step; and (3) at Layer 3, the global coordinator collects cluster-level aggregated models via dotted arrows every <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> rounds and redistributes the updated global model <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> to all clusters. This layered communication design ensures that intra-cluster knowledge transfer occurs at high frequency (every round) while cross-cluster aggregation occurs at lower frequency (every <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> rounds), reducing both privacy cost and communication overhead.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of the HiFraud framework. The three-layer architecture comprises a global coordination layer (Layer 3) for cross-cluster aggregation and privacy budget management, cluster coordination layers (Layer 2) for intra-cluster star-chain knowledge transfer, and institutional layers (Layer 1) for local training with adaptive differential privacy. Within each cluster, the star institution (<inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mo>&#x22C6;</mml:mo></mml:math></inline-formula>) distributes its model to all members (solid arrows), which then sequentially refine models along a similarity-ordered chain (dashed arrows). Cluster-level models are aggregated globally every <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> rounds (dotted arrows). Adaptive DP noise <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is applied at each transfer step, with intensity calibrated to institutional distributional divergence.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-1.tif"/>
</fig>
<p>Let <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:mi>&#x02110;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> denote the set of <italic>N</italic> participating financial institutions, where each institution <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> maintains a private local dataset <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> with fraud rate <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0.001</mml:mn><mml:mo>,</mml:mo><mml:mn>0.05</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. The objective of HiFraud is to collaboratively learn a global fraud detection model <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> and a set of cluster-specialized models <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:msubsup></mml:math></inline-formula> that collectively maximize detection performance across all institutions while satisfying a total differential privacy budget <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula>. The hierarchical structure partitions <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mi>&#x02110;</mml:mi></mml:mrow></mml:math></inline-formula> into <italic>K</italic> clusters <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>K</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where each cluster <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2286;</mml:mo><mml:mrow><mml:mi>&#x02110;</mml:mi></mml:mrow></mml:math></inline-formula> contains institutions with similar fraud characteristics, and the number of clusters <italic>K</italic> is determined dynamically through the clustering procedure described in <xref ref-type="sec" rid="s2_1">Section 2.1</xref>.</p>
<p>Algorithm 1 presents the complete training procedure of HiFraud.</p>
<fig id="fig-14">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-14.tif"/>
</fig>
<sec id="s2_1">
<label>2.1</label>
<title>Fraud-Aware Dynamic Clustering</title>
<p>Conventional federated clustering strategies group clients based on geographic proximity, organizational hierarchy, or generic model-weight similarity. In the context of fraud detection, however, institutions that are geographically distant may face nearly identical fraud schemes, while co-located institutions may encounter entirely different threat profiles. To capture this domain-specific structure, HiFraud introduces a fraud-aware clustering mechanism that groups institutions according to the distributional characteristics of their observed fraud patterns.</p>
<p>For each institution <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, a fraud pattern feature vector <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula> is constructed by concatenating four component encodings:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mtext>type</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mtext>temporal</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mtext>amount</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mtext>merchant</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mtext>type</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> encodes the distribution over fraud categories observed in the local dataset, <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mtext>temporal</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> captures temporal periodicity patterns extracted via a lightweight Transformer encoder, <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mtext>amount</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the statistical moments of transaction amount distributions stratified by fraud label, and <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mtext>merchant</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> encodes merchant category frequencies weighted by fraud incidence. Each component is computed locally and normalized to unit variance before transmission.</p>
<p>To prevent the clustering process from leaking sensitive institutional information, each feature vector is perturbed with calibrated Laplace noise prior to transmission to the global coordinator:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mtext>Lap</mml:mtext></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>f</mml:mi></mml:mrow><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>f</mml:mi></mml:math></inline-formula> is the <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>-sensitivity of the feature extraction function and <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> is the privacy budget allocated to the clustering phase. The sensitivity <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>f</mml:mi></mml:math></inline-formula> is bounded by the normalization applied to each component, ensuring that the noise magnitude remains controlled.</p>
<p>Given the set of perturbed feature vectors <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup></mml:math></inline-formula>, the global coordinator solves the following clustering optimization problem:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:munder><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mtext>fraud</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>bal</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>Var</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>K</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>fraud</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> denotes a fraud-specific distance metric defined as the weighted combination of Jensen&#x2013;Shannon divergence on fraud type distributions and Euclidean distance on temporal and amount features, <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> is the centroid of cluster <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mtext>Var</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> penalizes imbalanced cluster sizes to prevent degenerate partitions. The third term <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is a complementarity regularizer that penalizes configurations in which institutions holding rare fraud types are concentrated into a single small cluster. Formally, this regularizer is defined as:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mtext>type</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>type</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the cross-entropy between the local fraud type distribution and the cluster-average distribution <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mtext>type</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Minimizing this term encourages each cluster to maintain internal diversity of fraud types, thereby ensuring that rare fraud patterns receive sufficient representation within their assigned cluster rather than being marginalized. The hyperparameters <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>bal</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> control the relative importance of balance and complementarity, respectively.</p>
<p>The clustering is re-executed every <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> communication rounds to adapt to evolving fraud landscapes. Between re-clustering events, the cluster assignments remain fixed to preserve learning stability within each cluster. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates how clusters evolve over the course of training, transitioning from initial geographic groupings toward fraud-pattern-based configurations.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Evolution of fraud-aware clusters over 50 communication rounds. Early rounds exhibit geographic groupings inherited from initialization, which progressively transition to fraud-pattern-based clusters as the feature vectors capture increasingly discriminative fraud characteristics. By round 30, institutions handling similar fraud types are co-located in the same cluster regardless of geographic origin.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-2.tif"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Star-Chain Knowledge Transfer with Distillation</title>
<p>Within each cluster, HiFraud employs a star-chain transfer mechanism to efficiently propagate fraud detection knowledge among member institutions. This mechanism operates in two sequential phases: a star distribution phase that disseminates the best-performing model to all cluster members, followed by a chain enhancement phase that refines models through sequential passing with distillation-based forgetting mitigation. The design is motivated by two complementary findings from the recent literature: sequential model transfer achieves superior convergence under extreme non-IID conditions compared to parallel aggregation [<xref ref-type="bibr" rid="ref-24">24</xref>], while knowledge distillation on non-true classes effectively preserves global discriminative capacity during local adaptation [<xref ref-type="bibr" rid="ref-28">28</xref>].</p>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Star Distribution Phase</title>
<p>In each cluster <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>, the star institution <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msup><mml:mi>i</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:math></inline-formula> is selected as the member whose local model achieves the highest detection performance, with an adjustment for fraud rate scarcity:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msup><mml:mi>i</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>AUC</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mtext>AUC</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the area under the receiver operating characteristic curve evaluated on a held-out validation set at institution <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>i</mml:mi></mml:math></inline-formula>, and the term <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> upweights institutions with lower fraud rates to prioritize models that have learned to detect fraud under scarcity. This selection criterion ensures that the star model reflects strong detection capability rather than merely access to abundant fraud samples.</p>
<p>The star model <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msup><mml:mi>i</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> is then distributed to all other cluster members with Gaussian noise calibrated to the distributional distance between the star and each recipient:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msubsup><mml:mrow><mml:mover><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mi>i</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msup><mml:mi>i</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where the noise scale <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> for recipient institution <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>i</mml:mi></mml:math></inline-formula> is determined by the adaptive privacy mechanism described in <xref ref-type="sec" rid="s2_3">Section 2.3</xref>. Institutions that are distributionally closer to the star receive lower noise, preserving more of the transferred knowledge, while distributionally distant institutions receive stronger perturbation to protect against information leakage about the star&#x2019;s private data.</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Chain Enhancement Phase</title>
<p>Following star distribution, models are refined through a chain of sequential local adaptations. Institutions within each cluster are ordered by their distributional similarity to the star, forming a transfer chain <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>&#x03C0;</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> where <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mi>i</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:math></inline-formula>. Each institution in the chain receives the model from its predecessor, adapts it on local data, and passes the updated model to the next institution.</p>
<p>To mitigate the catastrophic forgetting that arises from sequential training on heterogeneous data, each local adaptation step incorporates a not-true-class distillation loss. Specifically, for institution <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> receiving model <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> from its predecessor, the local objective combines the standard supervised loss with a distillation term:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>NTD</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>CE</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the cross-entropy loss on local labeled data, <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>NTD</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the not-true-class distillation loss that penalizes divergence between the current model&#x2019;s predictions and the predecessor&#x2019;s predictions on classes other than the ground-truth label, and <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> controls the strength of distillation. The not-true-class formulation is particularly well-suited to fraud detection because it explicitly preserves the model&#x2019;s capacity to distinguish fraud types that may be absent from the current institution&#x2019;s data, thereby counteracting the forgetting of previously learned fraud patterns.</p>
<p>The model update at each chain step is then given by:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>LocalAdapt</mml:mtext></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the transfer coefficient that balances between retaining the institution&#x2019;s existing knowledge and incorporating transferred knowledge, and <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mtext>LocalAdapt</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> performs a fixed number of local gradient descent steps on the combined loss function in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>. After local adaptation, the model parameters are clipped to sensitivity bound <italic>S</italic> and perturbed with Gaussian noise before being passed to the next institution in the chain:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Clip</mml:mtext></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mrow><mml:mtext>chain</mml:mtext></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The chain enhancement is executed for <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>d</mml:mi></mml:math></inline-formula> sequential rounds within each cluster. After completion, the cluster model is obtained by weighted aggregation of all member models:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:munder><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> weights each institution proportionally to its dataset size. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> illustrates the two-phase star-chain transfer process within a single cluster.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Star-chain knowledge transfer within a cluster. In the star distribution phase (left), the best-performing institution (marked with <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mo>&#x22C6;</mml:mo></mml:math></inline-formula>, selected via <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>) broadcasts its model to all cluster members with adaptive noise. In the chain enhancement phase (right), models are sequentially refined along a similarity-ordered chain, with not-true-class distillation applied at each step to preserve knowledge of previously learned fraud patterns.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-3.tif"/>
</fig>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Privacy-Adaptive Mechanism</title>
<p>A uniform differential privacy mechanism applies identical noise to all institutions regardless of their data characteristics, which can disproportionately degrade detection performance for institutions with unique or rare fraud patterns. HiFraud addresses this through an adaptive noise calibration scheme that allocates stronger privacy protection to institutions whose data distributions diverge significantly from the cluster norm, while preserving model utility for institutions with representative distributions.</p>
<p>The noise scale <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> for each institution is determined by a calibration function that jointly considers distributional divergence and fraud pattern rarity:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>sigmoid</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mrow><mml:mtext>dp</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext>dp</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mtext>KL</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>P</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> measures the Kullback&#x2013;Leibler divergence between institution <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>i</mml:mi></mml:math></inline-formula>&#x2019;s fraud distribution <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and the cluster-average distribution <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mrow><mml:mover><mml:mi>P</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mi>r</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the local fraud rate. The parameters <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mtext>dp</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mtext>dp</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> control the sensitivity of noise calibration to distributional divergence and fraud rarity, respectively, while <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> bound the noise range. Institutions with high KL divergence (atypical fraud patterns) or low fraud rates (rare fraud observations) receive proportionally stronger noise to protect their distinctive and potentially sensitive patterns, while institutions with representative distributions receive less noise to maximize detection utility.</p>
<sec id="s2_3_1">
<label>2.3.1</label>
<title>Hierarchical Privacy Budget Allocation</title>
<p>The total privacy budget <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow></mml:math></inline-formula> is distributed across the three layers of the framework. Let <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>3</mml:mn></mml:msub></mml:math></inline-formula> denote the budgets allocated to the clustering phase, the star-chain transfer phase, and the global aggregation phase, respectively. Under the sequential composition property of R&#x00E9;nyi differential privacy, the total privacy cost across all phases and communication rounds is bounded by:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munder><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>&#x2113;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mrow><mml:mtext>clust</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>star</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>chain</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>global</mml:mtext></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:munder><mml:msubsup><mml:mi>R</mml:mi><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x2113;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>log</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>&#x03B4;</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msubsup><mml:mi>R</mml:mi><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x2113;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the R&#x00E9;nyi divergence of order <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> accumulated at layer <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>&#x2113;</mml:mi></mml:math></inline-formula> over all communication rounds, and <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is the relaxation parameter. This formulation provides tighter privacy accounting than basic sequential composition because R&#x00E9;nyi divergences compose linearly across independent mechanisms, enabling the framework to allocate privacy resources more efficiently.</p>
<p>Within the star-chain transfer phase, the budget <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> is further decomposed into a star distribution component and a chain enhancement component. The star distribution consumes privacy budget proportional to the number of recipient institutions, while each chain step consumes a budget proportional to the clipping bound and inverse noise scale. By reducing the frequency of global aggregation from every round to every <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> rounds, the hierarchical structure significantly reduces the privacy cost of the global aggregation phase relative to flat architectures, as the global coordinator receives only cluster-level models rather than individual institutional updates.</p>
</sec>
<sec id="s2_3_2">
<label>2.3.2</label>
<title>Formal Privacy Guarantee</title>
<p>The following theorem establishes the end-to-end privacy guarantee of HiFraud.</p>

<p><bold>Theorem 1:</bold> <italic>(Privacy Guarantee). Under the Gaussian mechanism with adaptive noise calibration defined in <xref ref-type="disp-formula" rid="eqn-11">Eqs. (11)</xref> and <xref ref-type="disp-formula" rid="eqn-12">(12)</xref>, gradient clipping bound S, and R&#x00E9;nyi DP composition in <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>, the HiFraud framework satisfies <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>-differential privacy for each participating institution over T communication rounds, where</italic>
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munder><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">chain</mml:mtext></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mi>T</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:mfrac><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">global</mml:mtext></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mi>log</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>&#x03B4;</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>

<p><bold>Proof:</bold> The proof follows from three observations. First, the clustering phase satisfies <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>-differential privacy through the Laplace mechanism applied to bounded-sensitivity feature vectors (<xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>). Specifically, each feature vector <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is normalized to unit variance before transmission, bounding the <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>-sensitivity <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>f</mml:mi></mml:math></inline-formula> to at most <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mn>2</mml:mn><mml:msqrt><mml:mi>d</mml:mi></mml:msqrt></mml:math></inline-formula>, where <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>d</mml:mi></mml:math></inline-formula> is the dimensionality of <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">f</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. By the Laplace mechanism guarantee [<xref ref-type="bibr" rid="ref-31">31</xref>], perturbing each coordinate with <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mtext>Lap</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>f</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> ensures <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>-differential privacy for the clustering input. Second, within the star-chain phase, each local adaptation step applies the Gaussian mechanism with clipping bound <italic>S</italic>, yielding a R&#x00E9;nyi divergence of <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> per step, and the <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mi>d</mml:mi></mml:math></inline-formula>-step chain composes linearly in R&#x00E9;nyi divergence. This follows directly from the composition property of R&#x00E9;nyi DP [<xref ref-type="bibr" rid="ref-32">32</xref>]: for <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mi>d</mml:mi></mml:math></inline-formula> sequential applications of the Gaussian mechanism, each with noise variance <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext>chain</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>-sensitivity <italic>S</italic>, the total R&#x00E9;nyi divergence of order <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext>chain</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. For the star distribution phase, each broadcast to a single recipient constitutes one application of the Gaussian mechanism with institution-specific noise <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula>, contributing <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to the R&#x00E9;nyi divergence. Since <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> by construction (<xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>), the worst-case per-round R&#x00E9;nyi divergence is bounded by <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, yielding the first term <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mi>T</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref> over <italic>T</italic> rounds. Third, the global aggregation occurs every <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> rounds, contributing <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mi>T</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> composition terms. Each global aggregation step applies the Gaussian mechanism to cluster-level model parameters with noise <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext>global</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula>, contributing <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext>global</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> per aggregation event. Over <italic>T</italic> rounds with aggregation every <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> rounds, this yields <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mo fence="false" stretchy="false">&#x230A;</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo fence="false" stretchy="false">&#x230B;</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mtext>global</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> total R&#x00E9;nyi divergence. The minimum over <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is computed numerically to obtain the tightest bound for any given configuration of noise parameters. Converting from R&#x00E9;nyi DP to <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>-DP follows from the standard conversion theorem [<xref ref-type="bibr" rid="ref-32">32</xref>]: for any <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mi>&#x03B4;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, an <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>-R&#x00E9;nyi divergence of <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:msub><mml:mi>R</mml:mi><mml:mi>&#x03B1;</mml:mi></mml:msub></mml:math></inline-formula> implies <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>-DP with <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>&#x03B1;</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Taking the minimum over <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> yields the tightest possible <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>-DP guarantee, completing the proof. <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mi>&#x25FB;</mml:mi></mml:math></inline-formula></p>
<p><italic>Discussion of Assumptions</italic>. The privacy guarantee in Theorem 1 relies on three key assumptions that merit explicit examination in the context of fraud detection. First, the gradient clipping bound <italic>S</italic> assumes that all per-sample gradient norms can be bounded by <italic>S</italic> without significant information loss. In practice, fraud detection models may exhibit larger gradient norms for rare fraud samples; we mitigate this by setting <italic>S</italic> based on the 95th percentile of observed gradient norms during a non-private warmup phase of 5 rounds, following the adaptive clipping strategy of Xue et al. [<xref ref-type="bibr" rid="ref-34">34</xref>]. Second, the composition assumes that the noise mechanisms at different layers are applied independently, which holds in our architecture because each layer operates on distinct parameter spaces (feature vectors for clustering, model parameters for star-chain, aggregated models for global). Third, the use of <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> as a worst-case bound in the first term of <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref> is conservative; in practice, most institutions receive noise levels above <inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula>, and the actual privacy cost is tighter than the stated bound. We verify empirically in <xref ref-type="sec" rid="s3_5">Section 3.5</xref> that the operational privacy budget consumption is approximately 35% lower than the theoretical worst case.</p>
</sec>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Global Aggregation and Convergence</title>
<p>At each global communication round, the cluster models <inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:msubsup></mml:math></inline-formula> are aggregated at the global coordinator to produce the updated global model:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The global model is then redistributed to all clusters as the initialization for the next round of local training and star-chain transfer.</p>

<p><bold>Theorem 2:</bold> <italic>(Convergence Rate). Assume that the global loss function <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is L-smooth, each local stochastic gradient has bounded variance <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:msup><mml:mi>&#x03C2;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>, and the data heterogeneity across clusters is bounded such that <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>&#x2264;</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula> for all <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>. Then the iterates of HiFraud satisfy:</italic>
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>T</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>[</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:mi>m</mml:mi><mml:mi>T</mml:mi></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mrow><mml:mi>T</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>
 where <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mi>m</mml:mi></mml:math></inline-formula> is the average cluster size, <italic>T</italic> is the total number of communication rounds, <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub></mml:math></inline-formula> is the model dimensionality, and <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula> is the maximum privacy noise variance.</p>

<p><bold>Proof:</bold> The proof proceeds by bounding the expected gradient norm through a standard one-step descent analysis adapted to the hierarchical structure. By <italic>L</italic>-smoothness:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mo fence="false" stretchy="false">&#x27E8;</mml:mo><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">&#x27E9;</mml:mo><mml:mo>+</mml:mo><mml:mfrac><mml:mi>L</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Substituting the global update rule (<xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>) and decomposing the update into three sources of error&#x2014;stochastic gradient variance within clusters, privacy noise, and inter-cluster heterogeneity&#x2014;yields:<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>&#x03B7;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:msup><mml:mi>&#x03C2;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mi>m</mml:mi></mml:mfrac><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x03B7;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x03B7;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula> is the effective learning rate. The first term arises from the variance of stochastic gradients averaged over <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:mi>m</mml:mi></mml:math></inline-formula> institutions per cluster. The second term captures the variance introduced by the Gaussian noise mechanism with maximum variance <inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula> applied to <inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub></mml:math></inline-formula>-dimensional model parameters. The third term bounds the bias due to inter-cluster distribution shift, quantified by the heterogeneity measure <inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>.</p>
<p>Telescoping across <italic>T</italic> rounds with learning rate <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mi>&#x03B7;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msqrt><mml:mi>m</mml:mi><mml:mi>T</mml:mi></mml:msqrt><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and rearranging:<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>T</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>[</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mfrac><mml:mrow><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03B8;</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>F</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow><mml:mrow><mml:mi>&#x03B7;</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mi>L</mml:mi><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msup><mml:mi>&#x03C2;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mi>m</mml:mi></mml:mfrac><mml:mo>+</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Substituting <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mi>&#x03B7;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msqrt><mml:mi>m</mml:mi><mml:mi>T</mml:mi></mml:msqrt><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> yields the bound in <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref>. The star-chain mechanism contributes to reducing <inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula> by improving intra-cluster alignment through sequential knowledge transfer: since institutions within each cluster are trained on increasingly similar models through the chain process, the effective inter-cluster heterogeneity is smaller than in flat architectures where each institution trains independently on its local data. <inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mi>&#x25FB;</mml:mi></mml:math></inline-formula></p>
<p>The convergence bound in <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref> decomposes into three interpretable terms. The first term <inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msqrt><mml:mi>m</mml:mi><mml:mi>T</mml:mi></mml:msqrt><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> reflects the convergence rate of distributed stochastic gradient descent (SGD) within clusters, where the effective parallelism <inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mi>m</mml:mi></mml:math></inline-formula> is determined by cluster size rather than the total number of institutions. The second term <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> captures the impact of privacy noise, which diminishes with additional communication rounds and is controlled by the adaptive noise calibration. The third term <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents an irreducible bias due to inter-cluster heterogeneity, bounded by the quality of the fraud-aware clustering. The star-chain mechanism contributes to reducing all three terms: it improves intra-cluster alignment (reducing the effective variance in the first term), enables more targeted noise allocation (reducing the second term through adaptive calibration), and produces cluster models that better capture local fraud patterns (reducing the residual heterogeneity in the third term).</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Computational Complexity and Communication Overhead</title>
<p>The per-round computational cost comprises: clustering (<inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mi>K</mml:mi><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, amortized over <inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> rounds, adding &#x003C;2% overhead), star-chain transfer (<inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>E</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> per cluster, where <inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> is model size and <italic>E</italic> is local epochs), and global aggregation (<inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>). For communication, HiFraud transmits <inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mi>K</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mi>m</mml:mi><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> parameters intra-cluster and <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mi>K</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> globally per round, where <inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>K</mml:mi></mml:math></inline-formula>. With <inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>=</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mn>20</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mi>K</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>=</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mn>5</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mi>d</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>=</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mn>3</mml:mn></mml:math></inline-formula>, total per-round cost is <inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mn>95</mml:mn><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula>; however, global aggregation occurs every <inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mi>&#x03C4;</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>=</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mn>10</mml:mn></mml:math></inline-formula> rounds, yielding amortized global cost of <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mn>0.5</mml:mn><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> per round vs. <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mn>20</mml:mn><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> for FedAvg. A detailed timing breakdown is provided in <xref ref-type="sec" rid="s3_6">Section 3.6</xref>.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Experimental Results</title>
<p>This section presents a comprehensive evaluation of HiFraud across multiple dimensions: overall detection performance, component-wise ablation, fraud pattern propagation speed, privacy&#x2013;utility trade-offs, adversarial robustness, scalability, per-type detection, and sensitivity to key hyperparameters.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Datasets and Experimental Setup</title>
<p>We evaluate HiFraud on three benchmark fraud detection datasets with distinct characteristics. <xref ref-type="table" rid="table-1">Table 1</xref> summarizes the key statistics of each dataset.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of benchmark datasets used in the experiments. Fraud rate denotes the proportion of fraudulent transactions in each dataset. Feature types include numerical (N), categorical (C), and temporal (T).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Transactions</th>
<th>Fraud Rate</th>
<th>Features</th>
<th>Feature Types</th>
<th>Domain</th>
</tr>
</thead>
<tbody>
<tr>
<td>IEEE-CIS</td>
<td>590,540</td>
<td>3.50%</td>
<td>433</td>
<td>N, C, T</td>
<td>E-commerce</td>
</tr>
<tr>
<td>PaySim</td>
<td>6,362,620</td>
<td>0.13%</td>
<td>11</td>
<td>N, C</td>
<td>Mobile payment</td>
</tr>
<tr>
<td>Worldline</td>
<td>284,807</td>
<td>0.172%</td>
<td>30</td>
<td>N (PCA)</td>
<td>Credit card</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The IEEE-CIS Fraud Detection dataset contains 590,540 e-commerce transactions with a 3.5% fraud rate and 433 heterogeneous features spanning transaction metadata (amount, product category, device information), identity features (email domain, device type, address), and V-features derived from principal component analysis of anonymized variables. PaySim provides 6.36 million synthetic mobile money transactions simulating real-world transfer, cash-out, and payment operations with 11 features including transaction type, amount, origin and destination account balances, and a binary fraud indicator; despite being synthetically generated, PaySim preserves the statistical properties of a real mobile money dataset from a developing country, including realistic class imbalance (0.13% fraud rate). The Worldline dataset comprises 284,807 credit card transactions with an extreme fraud rate of 0.172%, representing the most challenging class imbalance scenario among the three benchmarks ; all 28 numerical features are transformed via principal component analysis (PCA) for anonymization, with only the &#x201C;Time&#x201D; and &#x201C;Amount&#x201D; features retaining their original semantics.</p>
<p>To simulate realistic cross-institutional settings, we partition each dataset across <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>20</mml:mn></mml:math></inline-formula> institutions using three complementary strategies: fraud-type specialization, in which different institutions observe different fraud categories; temporal splitting, where institutions join the federation at staggered intervals; and geographic distribution, which introduces regional variations in fraud patterns. For IEEE-CIS, fraud-type specialization assigns each institution a primary fraud category based on the &#x201C;ProductCD&#x201D; and &#x201C;card6&#x201D; features, creating heterogeneous fraud distributions with Dirichlet parameter <inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula> to control the degree of non-IID-ness. For PaySim, institutions are specialized by transaction type (TRANSFER, CASH_OUT, PAYMENT, DEBIT), with each institution receiving 60%&#x2013;80% of transactions from its primary type. For Worldline, geographic distribution is simulated by partitioning transactions chronologically into 20 segments, with each segment assigned to one institution, reflecting the temporal evolution of fraud patterns across different &#x201C;regions&#x201D; of the transaction timeline.</p>
<p>The base fraud detection model at each institution is a 4-layer fully connected neural network with hidden dimensions [256, 128, 64, 32], ReLU activations, batch normalization after each hidden layer, and a sigmoid output layer. The model contains approximately 168 K trainable parameters. We use the Adam optimizer with an initial learning rate of <inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and cosine annealing decay over the total training rounds. Each local adaptation step consists of <inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> local epochs with batch size 256. The gradient clipping bound is set to <inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula>, determined by the 95th percentile of gradient norms during a 5-round non-private warmup phase. For the distillation loss <inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>NTD</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, we use a temperature of <inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>3.0</mml:mn></mml:math></inline-formula> applied to the softmax outputs of both the current and predecessor models.</p>
<p>All experiments are implemented in PyTorch 1.13.0 with Opacus 1.4.0 for differential privacy accounting. Unless otherwise stated, we use the following default configuration: the number of clusters <italic>K</italic> is determined dynamically within <inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>7</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, the star-chain enhancement runs for <inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> sequential rounds, the transfer coefficient is set to <inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula>, the distillation weight to <inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula>, the re-clustering interval to <inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> rounds, the privacy budget to <inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2.3</mml:mn></mml:math></inline-formula> with <inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and the noise bounds to <inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2.5</mml:mn></mml:math></inline-formula>. The adaptive noise calibration parameters are set to <inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mtext>dp</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2.0</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mtext>dp</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>, and the clustering hyperparameters to <inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>bal</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>comp</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula>. All results are averaged over five independent runs with different random seeds. All baseline methods are trained with the same base model architecture, privacy budget (<inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2.3</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), gradient clipping bound (<inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula>), and model capacity to ensure fair comparison. Hyperparameters specific to each baseline (e.g., the proximal coefficient <inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:mi>&#x03BC;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula> for FedProx) are individually tuned via grid search on a held-out validation set.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Main Results</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> compares HiFraud against seven baselines spanning centralized training, standard federated methods, and recent hierarchical and clustered approaches. All federated methods are evaluated under the same data partition and, where applicable, the same total privacy budget <inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2.3</mml:mn></mml:math></inline-formula>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Overall performance comparison on the IEEE-CIS dataset. MIA denotes membership inference attack success rate (%); lower is better. Precision and FPR (false positive rate) are additionally reported. Best federated results are in bold.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>AUC-ROC</th>
<th>F1</th>
<th>Recall</th>
<th>Precision</th>
<th>FPR (%)</th>
<th>Rounds</th>
<th>MIA (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Centralized<break/> (no privacy)</td>
<td>0.912</td>
<td>0.824</td>
<td>0.810</td>
<td>0.839</td>
<td>4.8</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>FedAvg [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>0.883</td>
<td>0.772</td>
<td>0.748</td>
<td>0.798</td>
<td>7.2</td>
<td>45</td>
<td>28.7</td>
</tr>
<tr>
<td>DP-FedAvg</td>
<td>0.845</td>
<td>0.721</td>
<td>0.695</td>
<td>0.749</td>
<td>9.1</td>
<td>49</td>
<td>15.2</td>
</tr>
<tr>
<td>FedProx</td>
<td>0.900</td>
<td>0.798</td>
<td>0.780</td>
<td>0.817</td>
<td>5.9</td>
<td>42</td>
<td>25.3</td>
</tr>
<tr>
<td>HierFL [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>0.908</td>
<td>0.810</td>
<td>0.793</td>
<td>0.828</td>
<td>5.3</td>
<td>35</td>
<td>18.6</td>
</tr>
<tr>
<td>ClusteredFL [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>0.914</td>
<td>0.818</td>
<td>0.801</td>
<td>0.836</td>
<td>5.0</td>
<td>33</td>
<td>14.8</td>
</tr>
<tr>
<td>FedFraud</td>
<td>0.917</td>
<td>0.825</td>
<td>0.808</td>
<td>0.843</td>
<td>4.7</td>
<td>38</td>
<td>12.3</td>
</tr>
<tr>
<td><bold>HiFraud (Ours)</bold></td>
<td><bold>0.935</bold></td>
<td><bold>0.852</bold></td>
<td><bold>0.838</bold></td>
<td><bold>0.867</bold></td>
<td><bold>3.8</bold></td>
<td><bold>30</bold></td>
<td><bold>10.2</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>HiFraud achieves the highest AUC-ROC of 0.935, surpassing the centralized baseline by 2.3% and the strongest federated competitor (FedFraud) by 1.8%. The improvement over DP-FedAvg is 10.5%, demonstrating that the hierarchical architecture substantially recovers the performance typically lost to differential privacy. In terms of convergence speed, HiFraud reaches its plateau in 30 communication rounds, representing a 37% reduction compared to DP-FedAvg (49 rounds) and a 21% reduction compared to FedFraud (38 rounds). The membership inference attack success rate of 10.2% approaches the random-guess baseline of 10%, indicating that the combination of hierarchical aggregation and adaptive noise provides strong empirical privacy protection.</p>
<p>We note that HiFraud&#x2019;s AUC-ROC (0.935) exceeds the centralized baseline (0.912) by 2.3%. This seemingly counterintuitive result can be attributed to two factors. First, the hierarchical structure introduces an implicit regularization effect: by training specialized cluster models before global aggregation, the framework prevents overfitting to the dominant non-fraud class that occurs in centralized training on imbalanced datasets. Second, the complementarity-aware clustering ensures that rare fraud patterns, which may be underrepresented in a single centralized training pass, receive focused attention within their assigned clusters through the star-chain mechanism. A similar phenomenon has been observed in clustered federated learning settings where local specialization outperforms global averaging on heterogeneous data [<xref ref-type="bibr" rid="ref-18">18</xref>]. However, we emphasize that this advantage is dataset- and partition-dependent: the centralized baseline represents a single-model upper bound under our specific non-IID partition, and centralized training with ensemble methods or specialized imbalance handling could potentially match or exceed HiFraud&#x2019;s performance.</p>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> presents the convergence trajectories of all methods across 50 communication rounds. HiFraud exhibits the steepest initial ascent and reaches 90% of its final performance by round 12, while DP-FedAvg requires approximately 30 rounds to reach the same relative milestone. The acceleration is attributable to the star-chain transfer mechanism, which enables efficient knowledge propagation within clusters of similar institutions, reducing the number of global rounds needed to disseminate useful fraud patterns.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Convergence trajectories of all methods on the IEEE-CIS dataset. HiFraud achieves the fastest convergence and highest final AUC-ROC, reaching 90% of its plateau by round 12.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-4.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> extends the comparison across all three datasets. HiFraud consistently outperforms all federated baselines on every dataset. The performance advantage is most pronounced on the Worldline dataset (0.948 vs. 0.858 for DP-FedAvg), where the extreme class imbalance (0.172% fraud rate) amplifies the benefit of the complementarity-aware clustering and adaptive privacy allocation. On PaySim, HiFraud achieves 0.962, exceeding even the centralized baseline (0.955), which we attribute to the regularization effect of the hierarchical structure preventing overfitting to the dominant non-fraud class.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>AUC-ROC comparison across three datasets. HiFraud consistently achieves the highest performance among federated methods and surpasses centralized training on PaySim.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Ablation Study</title>
<p>To quantify the contribution of each architectural component, we conduct an ablation study in which individual modules are removed while keeping all other components unchanged. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> summarizes the results in terms of AUC-ROC and convergence rounds.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Ablation study on the IEEE-CIS dataset. Each bar pair shows the AUC-ROC (left axis, blue) and convergence rounds (right axis, orange) when one component is removed from the full HiFraud framework.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-6.tif"/>
</fig>
<p>Removing the star-chain transfer mechanism produces the largest performance degradation, reducing AUC-ROC from 0.935 to 0.867 (a 6.8 percentage point drop) and increasing convergence rounds from 30 to 42. This result confirms that sequential knowledge transfer within clusters is the primary driver of both accuracy and convergence speed improvements. Without NTD distillation, AUC-ROC decreases to 0.901, demonstrating that forgetting mitigation accounts for approximately 3.4 percentage points of the total gain. Notably, even without distillation, the star-chain mechanism alone still outperforms all baselines, indicating that the transfer topology provides value independent of the forgetting mitigation strategy.</p>
<p>Removing fraud-aware clustering and reverting to random cluster assignment reduces AUC-ROC to 0.895 and increases convergence to 42 rounds, highlighting the importance of grouping institutions by fraud pattern similarity rather than arbitrary criteria. The adaptive differential privacy mechanism contributes 2.5 percentage points over uniform noise allocation, with its removal reducing AUC-ROC to 0.910 while only modestly affecting convergence. The complementarity regularizer provides a smaller but meaningful improvement of 1.7 percentage points, with its primary benefit concentrated on rare fraud types as discussed in <xref ref-type="sec" rid="s3_7">Section 3.7</xref>.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Fraud Pattern Propagation</title>
<p>A critical operational requirement for fraud detection systems is the ability to rapidly disseminate knowledge of newly emerging fraud patterns across institutions. To evaluate this capability, we simulate the injection of a novel fraud type at round 25 in a single institution and track the detection rate of this new pattern across the cluster hierarchy over subsequent rounds. <xref ref-type="fig" rid="fig-7">Fig. 7</xref> presents the results.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Detection rate of a novel fraud pattern injected at round 25. Within the same cluster, 80% of institutions detect the new pattern within 2 rounds (<inline-formula id="ieqn-215"><mml:math id="mml-ieqn-215"><mml:mo>&#x223C;</mml:mo></mml:math></inline-formula>3 h), compared to 12 rounds for adjacent clusters and over 20 rounds for flat FL.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-7.tif"/>
</fig>
<p>Within the same cluster, the star-chain mechanism enables 80% of institutions to detect the novel fraud pattern within 2 communication rounds, corresponding to approximately 3 h in our experimental setup. This &#x201C;3-h&#x201D; figure is derived from our experimental configuration in which each communication round takes approximately 90 min, comprising local training (<inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> epochs <inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> approximately 12 min per epoch &#x003D; 60 min), star-chain transfer (sequential model passing among <inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:mi>m</mml:mi><mml:mo>&#x2248;</mml:mo><mml:mn>4</mml:mn></mml:math></inline-formula> institutions per cluster <inline-formula id="ieqn-212"><mml:math id="mml-ieqn-212"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> approximately 5 min per pass &#x003D; 20 min), and global aggregation plus communication overhead (approximately 10 min). Thus, 2 rounds <inline-formula id="ieqn-213"><mml:math id="mml-ieqn-213"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 90 min <inline-formula id="ieqn-214"><mml:math id="mml-ieqn-214"><mml:mo>&#x2248;</mml:mo></mml:math></inline-formula> 3 h. We note that this latency is configuration-dependent: the actual propagation time in production deployments would scale with dataset size, model complexity, network bandwidth, and the number of institutions per cluster. In a setting with faster hardware or fewer local epochs, propagation could be significantly faster; conversely, larger models or slower networks would increase the latency. Propagation to adjacent clusters, which occurs through the global aggregation pathway, requires approximately 7 additional rounds. In contrast, flat federated learning (FedAvg) requires over 20 rounds to achieve comparable detection rates, as the new pattern must influence the global model before being redistributed to all participants. This result demonstrates that the hierarchical architecture provides a substantial operational advantage for responding to emerging threats, enabling institutions within a cluster to benefit from each other&#x2019;s observations with minimal latency.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Privacy Analysis</title>
<sec id="s3_5_1">
<label>3.5.1</label>
<title>Privacy&#x2013;Utility Trade-Off</title>
<p><xref ref-type="fig" rid="fig-8">Fig. 8</xref> illustrates the relationship between the total privacy budget <inline-formula id="ieqn-216"><mml:math id="mml-ieqn-216"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow></mml:math></inline-formula> and detection performance for HiFraud and three baselines. Across the entire range of privacy budgets evaluated (<inline-formula id="ieqn-217"><mml:math id="mml-ieqn-217"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>1.0</mml:mn><mml:mo>,</mml:mo><mml:mn>5.0</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>), HiFraud consistently achieves the highest AUC-ROC for any given privacy level.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Privacy&#x2013;utility trade-off. HiFraud achieves superior AUC-ROC at every privacy budget level. At the operating point <inline-formula id="ieqn-218"><mml:math id="mml-ieqn-218"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2.3</mml:mn></mml:math></inline-formula>, HiFraud attains 0.935, compared to 0.845 for DP-FedAvg under the same budget.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-8.tif"/>
</fig>
<p>The advantage of HiFraud is most pronounced at tighter privacy budgets. At <inline-formula id="ieqn-219"><mml:math id="mml-ieqn-219"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula>, HiFraud achieves an AUC-ROC of 0.890, compared to 0.780 for DP-FedAvg, representing a 14% relative improvement. This gap narrows to 7% at <inline-formula id="ieqn-220"><mml:math id="mml-ieqn-220"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>5.0</mml:mn></mml:math></inline-formula>, indicating that the hierarchical privacy allocation provides the greatest benefit precisely when privacy constraints are most stringent. The efficiency gain stems from two sources: the reduced frequency of global aggregation lowers the cumulative privacy cost of inter-cluster communication, and the adaptive noise calibration directs privacy budget toward institutions that need it most while minimizing unnecessary noise for representative institutions.</p>
<p><xref ref-type="table" rid="table-3">Table 3</xref> provides a detailed breakdown of how the privacy budget is allocated across the three layers of HiFraud, compared to uniform and non-hierarchical adaptive allocation.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Privacy budget allocation across framework layers and corresponding AUC-ROC. The hierarchical approach achieves higher performance with the same total budget by reducing redundant noise in global aggregation.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Component</th>
<th>Uniform DP</th>
<th>Adaptive DP</th>
<th>Hierarchical (Ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clustering</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>0.3</td>
</tr>
<tr>
<td>Local Training</td>
<td>1.2</td>
<td>0.8&#x2013;1.8</td>
<td>0.6&#x2013;1.2</td>
</tr>
<tr>
<td>Star Distribution</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>0.4</td>
</tr>
<tr>
<td>Chain Transfer</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>0.3</td>
</tr>
<tr>
<td>Global Aggregation</td>
<td>1.1</td>
<td>0.5</td>
<td>0.1</td>
</tr>
<tr>
<td>Total <inline-formula id="ieqn-221"><mml:math id="mml-ieqn-221"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>2.3</td>
<td>2.3</td>
<td>2.3</td>
</tr>
<tr>
<td><bold>AUC-ROC</bold></td>
<td>0.845</td>
<td>0.917</td>
<td><bold>0.935</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_5_2">
<label>3.5.2</label>
<title>Resistance to Membership Inference Attacks</title>
<p>We evaluate the empirical privacy protection of HiFraud against membership inference attacks using the gradient-based attack framework described in Bai et al. [<xref ref-type="bibr" rid="ref-37">37</xref>]. The attacker is assumed to have passive access to the aggregated model updates at the cluster level and employs a binary classifier trained to distinguish between member and non-member samples. <xref ref-type="fig" rid="fig-9">Fig. 9</xref> reports the attack success rates across all methods.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Membership inference attack success rate. Lower values indicate stronger privacy. HiFraud achieves 10.2%, approaching the 10% random-guess baseline.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-9.tif"/>
</fig>
<p>HiFraud achieves the lowest MIA success rate of 10.2%, approaching the 10% random-guess baseline for our 10-class attack formulation. This represents a 64% reduction relative to FedAvg (28.7%) and a 17% reduction relative to FedFraud (12.3%). The strong privacy protection results from the compounding effect of three mechanisms: the adaptive noise injection obscures individual gradient contributions, the hierarchical aggregation limits the attacker&#x2019;s visibility to cluster-level updates rather than institutional-level parameters, and the star-chain transfer introduces additional noise through sequential model passing. Notably, HiFraud provides stronger empirical privacy than DP-FedAvg (15.2%) despite achieving substantially higher detection performance, demonstrating that the hierarchical architecture enables a more favorable privacy&#x2013;utility operating point.</p>
</sec>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Scalability Analysis</title>
<p><xref ref-type="fig" rid="fig-10">Fig. 10</xref> evaluates the performance and communication efficiency of HiFraud as the number of participating institutions increases from 10 to 100.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Scalability analysis. HiFraud maintains stable AUC-ROC as the federation grows to 100 institutions, while FedAvg degrades by 4.1 percentage points. Communication cost scales as <inline-formula id="ieqn-222"><mml:math id="mml-ieqn-222"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for HiFraud vs. <inline-formula id="ieqn-223"><mml:math id="mml-ieqn-223"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for FedAvg.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-10.tif"/>
</fig>
<p>HiFraud maintains stable detection performance across all federation sizes, with AUC-ROC decreasing only marginally from 0.930 (10 institutions) to 0.927 (100 institutions), a degradation of 0.3 percentage points. In contrast, FedAvg suffers a 4.1 percentage point decline over the same range, from 0.882 to 0.842, as increasing data heterogeneity overwhelms the flat aggregation mechanism. The communication cost of HiFraud grows sublinearly with the number of institutions due to the hierarchical structure: at 100 institutions, HiFraud requires 4.8<inline-formula id="ieqn-224"><mml:math id="mml-ieqn-224"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> the communication cost of the 10-institution baseline, compared to 28.0<inline-formula id="ieqn-225"><mml:math id="mml-ieqn-225"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> for FedAvg. This efficiency arises because only <italic>K</italic> cluster models (rather than <italic>N</italic> individual models) are transmitted during global aggregation, and the star-chain transfer within each cluster is sequential, requiring only <inline-formula id="ieqn-226"><mml:math id="mml-ieqn-226"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> transmissions per cluster. <xref ref-type="table" rid="table-4">Table 4</xref> provides a detailed breakdown of the computational overhead per communication round.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Computational overhead breakdown per communication round on the IEEE-CIS dataset with <inline-formula id="ieqn-227"><mml:math id="mml-ieqn-227"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>20</mml:mn></mml:math></inline-formula> institutions and <inline-formula id="ieqn-228"><mml:math id="mml-ieqn-228"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> clusters. Re-clustering is amortized over <inline-formula id="ieqn-229"><mml:math id="mml-ieqn-229"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> rounds. All times are measured on a single NVIDIA A100 GPU.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Component</th>
<th>Time (min)</th>
<th>Proportion (%)</th>
<th>Comm. Cost (<inline-formula id="ieqn-230"><mml:math id="mml-ieqn-230"><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula>)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Local Training (<inline-formula id="ieqn-231"><mml:math id="mml-ieqn-231"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> epochs)</td>
<td>60.0</td>
<td>66.7</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Star Distribution</td>
<td>5.0</td>
<td>5.6</td>
<td><inline-formula id="ieqn-232"><mml:math id="mml-ieqn-232"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> per cluster</td>
</tr>
<tr>
<td>Chain Transfer (<inline-formula id="ieqn-233"><mml:math id="mml-ieqn-233"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula>)</td>
<td>15.0</td>
<td>16.7</td>
<td><inline-formula id="ieqn-234"><mml:math id="mml-ieqn-234"><mml:mn>9</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> per cluster</td>
</tr>
<tr>
<td>Global Aggregation</td>
<td>8.0</td>
<td>8.9</td>
<td><inline-formula id="ieqn-235"><mml:math id="mml-ieqn-235"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>Re-clustering (amortized)</td>
<td>2.0</td>
<td>2.2</td>
<td><inline-formula id="ieqn-236"><mml:math id="mml-ieqn-236"><mml:mn>20</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>Total per round</bold></td>
<td><bold>90.0</bold></td>
<td><bold>100.0</bold></td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Per-Type Detection and Clustering Analysis</title>
<p>To evaluate whether the hierarchical architecture improves detection uniformly or preferentially benefits specific fraud categories, we report per-type AUC-ROC in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Per-fraud-type AUC-ROC. HiFraud achieves the most uniform performance across fraud types, with the greatest improvement on the rarest category (Synthetic ID: &#x002B;23.0% over DP-FedAvg).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-11.tif"/>
</fig>
<p>HiFraud provides substantial and consistent improvements across all fraud types, with the most pronounced gains on rare categories. For Synthetic ID fraud, the rarest type in our experimental setup, HiFraud achieves an AUC-ROC of 0.910, compared to 0.782 for FedFraud and 0.680 for DP-FedAvg. This 23.0 percentage point improvement over DP-FedAvg demonstrates the effectiveness of the complementarity regularizer in preventing the marginalization of institutions holding rare fraud types. The performance variance across fraud types is also notably reduced: the standard deviation of per-type AUC-ROC is 0.015 for HiFraud, compared to 0.058 for FedFraud and 0.078 for DP-FedAvg, indicating more equitable detection across all fraud categories.</p>
<p><xref ref-type="fig" rid="fig-12">Fig. 12</xref> visualizes the evolution of cluster assignments over the course of training, using a t-distributed stochastic neighbor embedding (t-SNE) projection of institutional fraud feature vectors with markers indicating the dominant fraud type at each institution.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Evolution of cluster assignments visualized via t-SNE projection. Colors indicate cluster assignment; marker shapes indicate dominant fraud type. By round 30, clusters align closely with fraud type rather than initial geographic grouping.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-12.tif"/>
</fig>
<p>At round 1, cluster assignments reflect the initial geographic grouping, with each cluster containing a heterogeneous mix of fraud types (indicated by diverse marker shapes within each color group). By round 15, the clustering begins to reorganize around fraud pattern similarity, as the fraud-aware feature vectors become more discriminative through iterative refinement. By round 30, clusters are strongly aligned with fraud type: institutions facing similar fraud categories are co-located in the same cluster regardless of their geographic origin. This transition demonstrates that the dynamic re-clustering mechanism successfully adapts to the underlying fraud structure of the data, enabling increasingly specialized intra-cluster knowledge transfer as training progresses.</p>
</sec>
<sec id="s3_8">
<label>3.8</label>
<title>Sensitivity Analysis</title>
<p>We investigate the sensitivity of HiFraud to three key hyperparameters: the transfer coefficient <inline-formula id="ieqn-237"><mml:math id="mml-ieqn-237"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, the distillation weight <inline-formula id="ieqn-238"><mml:math id="mml-ieqn-238"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, and the re-clustering interval <inline-formula id="ieqn-239"><mml:math id="mml-ieqn-239"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-13">Fig. 13</xref> presents the results.</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Sensitivity to key hyperparameters. (<bold>a</bold>) Transfer coefficient <inline-formula id="ieqn-240"><mml:math id="mml-ieqn-240"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>: optimal at 0.3, balancing knowledge absorption with local retention. (<bold>b</bold>) Distillation weight <inline-formula id="ieqn-241"><mml:math id="mml-ieqn-241"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>: optimal at 0.3, with excessive distillation (&#x003E;0.5) constraining model plasticity. (<bold>c</bold>) Re-clustering interval <inline-formula id="ieqn-242"><mml:math id="mml-ieqn-242"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>: optimal at 10, trading off adaptation responsiveness against learning stability.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81922-fig-13.tif"/>
</fig>
<p>The transfer coefficient <inline-formula id="ieqn-243"><mml:math id="mml-ieqn-243"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> achieves optimal performance at <inline-formula id="ieqn-244"><mml:math id="mml-ieqn-244"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula> (<xref ref-type="fig" rid="fig-13">Fig. 13a</xref>). Values below 0.2 underutilize transferred knowledge, while values above 0.5 cause excessive reliance on the predecessor model at the expense of local adaptation, reducing AUC-ROC by up to 2.7 percentage points. The distillation weight <inline-formula id="ieqn-245"><mml:math id="mml-ieqn-245"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> exhibits a similar concave profile with an optimum at <inline-formula id="ieqn-246"><mml:math id="mml-ieqn-246"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula> (<xref ref-type="fig" rid="fig-13">Fig. 13b</xref>). When distillation is disabled (<inline-formula id="ieqn-247"><mml:math id="mml-ieqn-247"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>), AUC-ROC drops to 0.901, confirming the value of forgetting mitigation. Excessive distillation (<inline-formula id="ieqn-248"><mml:math id="mml-ieqn-248"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KD</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>) constrains the model&#x2019;s ability to learn new local patterns, reducing performance by up to 2.3 percentage points.</p>
<p>The re-clustering interval <inline-formula id="ieqn-249"><mml:math id="mml-ieqn-249"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> presents a trade-off between adaptation responsiveness and learning stability (<xref ref-type="fig" rid="fig-13">Fig. 13c</xref>). Frequent re-clustering (<inline-formula id="ieqn-250"><mml:math id="mml-ieqn-250"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>) disrupts ongoing intra-cluster learning, as newly formed clusters must re-establish star-chain transfer from scratch, resulting in 4 additional convergence rounds compared to the optimal <inline-formula id="ieqn-251"><mml:math id="mml-ieqn-251"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula>. Infrequent re-clustering (<inline-formula id="ieqn-252"><mml:math id="mml-ieqn-252"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>50</mml:mn></mml:math></inline-formula>) fails to track evolving fraud patterns, reducing AUC-ROC by 2.5 percentage points as institutions with shifting fraud profiles remain trapped in suboptimal clusters. The optimal interval of <inline-formula id="ieqn-253"><mml:math id="mml-ieqn-253"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> rounds corresponds to approximately 15 h in our experimental setup, providing a practical balance between responsiveness and stability for real-world deployment.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Discussion</title>
<sec id="s4_1">
<label>4.1</label>
<title>Implications</title>
<p>The experimental results carry several practical implications. The finding that HiFraud surpasses centralized training on PaySim (0.962 vs. 0.955) challenges the assumption that federated approaches necessarily sacrifice detection quality for privacy, suggesting that the hierarchical structure introduces a beneficial inductive bias by increasing the diversity of training signals without exposing raw data. In production settings, financial institutions can thus achieve detection performance meeting or exceeding centralized alternatives while fully complying with GDPR and CCPA, and the 3-h propagation latency for novel fraud patterns enables rapid collective response to emerging threats. Equally important, the hierarchical architecture fundamentally alters the privacy&#x2013;utility trade-off: at the stringent budget of <inline-formula id="ieqn-254"><mml:math id="mml-ieqn-254"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula>, HiFraud achieves an AUC-ROC of 0.890&#x2014;higher than DP-FedAvg at the much looser <inline-formula id="ieqn-255"><mml:math id="mml-ieqn-255"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>3.0</mml:mn></mml:math></inline-formula> (0.860)&#x2014;demonstrating that institutions under strict data protection regimes need not accept severe performance penalties. The adaptive noise calibration directs privacy budget toward institutions with unique fraud patterns while avoiding unnecessary noise for representative ones, and the near-random MIA success rate of 10.2% confirms that the hierarchical structure inherently limits information leakage by exposing only cluster-level aggregates to potential attackers. For deployment, the global coordinator role can be assumed by a regulatory body or implemented via secure multi-party computation, institutional onboarding requires only local SQL-based feature computation (<xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>), and the framework scales to 100&#x002B; institutions through multi-level hierarchy with <inline-formula id="ieqn-256"><mml:math id="mml-ieqn-256"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> communication cost, as validated in <xref ref-type="sec" rid="s3_6">Section 3.6</xref>.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Limitations and Future Work</title>
<p>Despite the strong performance across multiple benchmarks, several limitations merit acknowledgment. The dynamic re-clustering mechanism introduces periodic disruptions to intra-cluster learning; our sensitivity analysis shows that the interval <inline-formula id="ieqn-257"><mml:math id="mml-ieqn-257"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> must be carefully tuned, and future work should investigate smooth cluster transition mechanisms leveraging soft clustering [<xref ref-type="bibr" rid="ref-23">23</xref>] and continual learning [<xref ref-type="bibr" rid="ref-27">27</xref>] to enable gradual migration without discarding accumulated knowledge. The current framework also assumes homogeneous computational capabilities, whereas real-world participants range from multinational banks to small credit unions, motivating extensions with adaptive computation allocation and asynchronous protocols. From a security perspective, the star institution occupies a privileged position that creates a potential single point of vulnerability: a compromised star could propagate poisoned updates to all cluster members. While clipping and noise injection in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref> provide partial mitigation, integrating dedicated Byzantine-robust star selection based on recent hierarchical robust aggregation [<xref ref-type="bibr" rid="ref-40">40</xref>] and trust-score filtering [<xref ref-type="bibr" rid="ref-39">39</xref>] would substantially strengthen resilience. The theoretical composition of formal privacy guarantees with Byzantine robustness in hierarchical adaptive settings also remains an open problem warranting further investigation.</p>
<p>Additionally, the current evaluation is conducted on benchmark datasets that, while widely used in the fraud detection literature, may not fully capture the complexity of production fraud systems. Real-world deployments involve continuously evolving fraud tactics, adversarial adaptation, and regulatory constraints that vary across jurisdictions. Future work should evaluate HiFraud on proprietary institutional datasets in controlled pilot studies to validate the framework&#x2019;s effectiveness under genuine operational conditions.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>This paper proposed HiFraud, a hierarchical federated learning framework that addresses the fundamental challenges of cross-institutional fraud detection through a three-layer architecture integrating fraud-aware dynamic clustering with complementarity regularization, star-chain knowledge transfer augmented by not-true-class distillation for forgetting mitigation, and privacy-adaptive aggregation grounded in R&#x00E9;nyi differential privacy composition. The key technical contributions include: (i) a fraud-aware dynamic clustering mechanism with complementarity regularization that groups institutions by fraud pattern similarity while preserving rare-type representation; (ii) a star-chain knowledge transfer mechanism with domain-specific innovations including fraud-rate-adjusted star selection, similarity-ordered chain traversal, and not-true-class distillation for forgetting mitigation; (iii) a hierarchical adaptive privacy allocation scheme based on R&#x00E9;nyi DP composition that calibrates noise to distributional divergence and fraud rarity; and (iv) formal privacy and convergence guarantees with detailed proofs under explicit assumptions. Experiments on three benchmark datasets demonstrated that HiFraud achieves an AUC-ROC of 0.935 under <inline-formula id="ieqn-258"><mml:math id="mml-ieqn-258"><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2.3</mml:mn></mml:math></inline-formula> differential privacy, outperforming standard DP-FedAvg by 10.5% while reducing convergence rounds by 39% and suppressing membership inference attack success to near-random levels (10.2%). The star-chain mechanism enables detection of emerging fraud patterns within approximately 3 h inside clusters under our experimental configuration, and the complementarity-aware clustering improves rare fraud type detection by 23.0% over uniform privacy baselines. These results establish that hierarchical architectures can effectively reconcile the competing demands of detection performance, formal privacy guarantees, and rapid threat response in collaborative financial fraud detection, providing a practical blueprint for privacy-preserving multi-institutional learning in regulated environments. However, the framework currently assumes homogeneous computational capabilities and a trusted global coordinator, and the dynamic re-clustering interval requires careful tuning. Future work should address these limitations through asynchronous protocols, Byzantine-robust star selection, and validation on proprietary institutional datasets.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Zhihao Zhang and Zhuodong Liu; methodology, Zhihao Zhang and Zhuodong Liu; software, Zhihao Zhang and Zhuodong Liu; validation, Zhihao Zhang, Zhuodong Liu and Xiangyu Li; formal analysis, Zhihao Zhang and Zhuodong Liu; investigation, Zhihao Zhang, Zhuodong Liu and Xiangyu Li; resources, Lei Zhang; data curation, Xiangyu Li; writing&#x2014;original draft preparation, Zhihao Zhang and Zhuodong Liu; writing&#x2014;review and editing, Xiangyu Li and Lei Zhang; visualization, Zhihao Zhang and Zhuodong Liu; supervision, Lei Zhang; project administration, Lei Zhang. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The three benchmark datasets used in this study are publicly available: IEEE-CIS Fraud Detection (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/c/ieee-fraud-detection">https://www.kaggle.com/c/ieee-fraud-detection</ext-link>), PaySim (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/ealaxi/paysim1">https://www.kaggle.com/datasets/ealaxi/paysim1</ext-link>), and Worldline (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud">https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud</ext-link>). The federated data partitioning configurations, which simulate cross-institutional settings as described in <xref ref-type="sec" rid="s3_1">Section 3.1</xref>, are not derived from real institutional records and do not contain sensitive information. The experimental code, including data partitioning scripts and all baseline implementations, will be released upon acceptance of this paper.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chatterjee</surname> <given-names>P</given-names></string-name>, <string-name><surname>Das</surname> <given-names>D</given-names></string-name>, <string-name><surname>Rawat</surname> <given-names>DB</given-names></string-name></person-group>. <article-title>Digital twin for credit card fraud detection: opportunities, challenges, and fraud detection advancements</article-title>. <source>Future Gener Comput Syst</source>. <year>2024</year>;<volume>158</volume>:<fpage>410</fpage>&#x2013;<lpage>26</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Federal Trade Commission</collab></person-group>. <article-title>New FTC data show consumers reported losing more than $12.5 billion to fraud in 2024 [Internet]. Washington, DC, USA: FTC</article-title>; <comment>2025 [cited 2025 Mar 15]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.ftc.gov/news-events/news/press-releases/2025/03/new-ftc-data-show-big-jump-reported-losses-fraud-125-billion-2024">https://www.ftc.gov/news-events/news/press-releases/2025/03/new-ftc-data-show-big-jump-reported-losses-fraud-125-billion-2024</ext-link>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Privacy-aware financial risk control: a federated learning approach with differential privacy optimization</article-title>. <source>J Comput Technol Softw</source>. <year>2025</year>;<volume>4</volume>:<fpage>37</fpage>&#x2013;<lpage>52</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>McMahan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Moore</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ramage</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hampson</surname> <given-names>S</given-names></string-name>, <string-name><surname>Arcas</surname> <given-names>BAY</given-names></string-name></person-group>. <article-title>Communication-efficient learning of deep networks from decentralized data</article-title>. In: <conf-name>Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS); 2017 Apr 20&#x2013;22</conf-name>; <publisher-loc>Fort Lauderdale, FL, USA</publisher-loc>. p. <fpage>1273</fpage>&#x2013;<lpage>82</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Federated learning on non-IID data: a survey</article-title>. <source>Neurocomputing</source>. <year>2021</year>;<volume>465</volume>:<fpage>371</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2021.07.098</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A survey on class imbalance in federated learning</article-title>. <comment>arXiv:2303.11673. 2023</comment>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>CZ</given-names></string-name></person-group>. <article-title>FFD:a federated learning based method for credit card fraud detection</article-title>. In: <conf-name>Proceedings of the International Conference on Big Data; 2019 Dec 10&#x2013;13</conf-name>; <publisher-loc>Los Angeles, CA, USA</publisher-loc>. p. <fpage>18</fpage>&#x2013;<lpage>32</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdul Salam</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fouad</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Elbably</surname> <given-names>DL</given-names></string-name>, <string-name><surname>Elsayed</surname> <given-names>SM</given-names></string-name></person-group>. <article-title>Federated learning model for credit card fraud detection with data balancing techniques</article-title>. <source>Neural Comput Appl</source>. <year>2024</year>;<volume>36</volume>(<issue>11</issue>):<fpage>7359</fpage>&#x2013;<lpage>78</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-023-09410-2</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shah</surname> <given-names>P</given-names></string-name>, <string-name><surname>Patil</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Secure and efficient fraud detection using federated learning and distributed search databases</article-title>. In: <conf-name>Proceedings of the IEEE 4th International Conference on AI in Cybersecurity (ICAIC)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hilou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dheeb</surname> <given-names>S</given-names></string-name>, <string-name><surname>Radhi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khadim</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Majeed</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Federated learning for credit card fraud detection: a privacy-preserving approach with SMOTE optimization</article-title>. <source>J Al-Qadisiyah Comput Sci Math</source>. <year>2025</year>;<volume>17</volume>(<issue>3</issue>):<fpage>Comp 44</fpage>&#x2013;<lpage>57</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Farooq</surname> <given-names>M</given-names></string-name>, <string-name><surname>Munir</surname> <given-names>S</given-names></string-name>, <string-name><surname>Manzoor</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shaheen</surname> <given-names>M</given-names></string-name></person-group>. <article-title>AI-driven adaptive federated learning with privacy preservation and imbalance adjustment for financial credit card fraud detection</article-title>. <source>Appl Comput Intell Soft Comput</source>. <year>2025</year>;<volume>2025</volume>:<fpage>7116768</fpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sarkar</surname> <given-names>D</given-names></string-name>, <string-name><surname>Narang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rai</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Fed-Focal Loss for imbalanced data classification in federated learning</article-title>. <comment>arXiv:2011.06283. 2020</comment>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Addressing class imbalance in federated learning</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence</conf-name>. <publisher-loc>Menlo Park, CA, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>; <year>2021</year>. Vol. <volume>35</volume>, p. <fpage>10165</fpage>&#x2013;<lpage>73</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aljunaid</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Almheiri</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>Dawood</surname> <given-names>H</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Secure and transparent banking: explainable AI-driven federated learning model for financial fraud detection</article-title>. <source>J Risk Financ Manag</source>. <year>2025</year>;<volume>18</volume>(<issue>4</issue>):<fpage>179</fpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Song</surname> <given-names>S</given-names></string-name>, <string-name><surname>Letaief</surname> <given-names>KB</given-names></string-name></person-group>. <article-title>Client-edge-cloud hierarchical federated learning</article-title>. In: <conf-name>Proceedings of the IEEE International Conference on Communications (ICC); 2020 Jun 7&#x2013;11</conf-name>; <publisher-loc>Dublin, Ireland</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chao</surname> <given-names>HC</given-names></string-name></person-group>. <article-title>A review on federated learning architectures for privacy-preserving AI: lightweight and secure cloud-edge&#x2013;end collaboration</article-title>. <source>Electronics</source>. <year>2025</year>;<volume>14</volume>(<issue>13</issue>):<fpage>2512</fpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Albshaier</surname> <given-names>L</given-names></string-name>, <string-name><surname>Almarri</surname> <given-names>S</given-names></string-name>, <string-name><surname>Albuali</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Federated learning for cloud and edge security: a systematic review of challenges and AI opportunities</article-title>. <source>Electronics</source>. <year>2025</year>;<volume>14</volume>(<issue>5</issue>):<fpage>1019</fpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sattler</surname> <given-names>F</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>KR</given-names></string-name>, <string-name><surname>Samek</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Clustered federated learning: model-agnostic distributed multitask learning for non-IID data</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2021</year>;<volume>32</volume>:<fpage>3710</fpage>&#x2013;<lpage>22</lpage>; <pub-id pub-id-type="pmid">32833654</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gong</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xing</surname> <given-names>T</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xi</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Adaptive client clustering for efficient federated learning over non-IID and imbalanced data</article-title>. <source>IEEE Trans Big Data</source>. <year>2024</year>;<volume>10</volume>(<issue>6</issue>):<fpage>1051</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tbdata.2022.3167994</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Duan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Flexible clustered federated learning for client-level data distribution shift</article-title>. <source>IEEE Trans Parallel Distrib Syst</source>. <year>2022</year>;<volume>33</volume>:<fpage>2661</fpage>&#x2013;<lpage>74</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpds.2021.3134263</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bhatti</surname> <given-names>DMS</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>BJ</given-names></string-name></person-group>. <article-title>dy-TACFL: dynamic temporal adaptive clustered federated learning for heterogeneous clients</article-title>. <source>Electronics</source>. <year>2025</year>;<volume>14</volume>(<issue>1</issue>):<fpage>152</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics14010152</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Islam</surname> <given-names>M</given-names></string-name>, <string-name><surname>Javaherian</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tzeng</surname> <given-names>N</given-names></string-name></person-group>. <article-title>FedClust: optimizing federated learning on non-IID data through weight-driven client clustering</article-title>. In: <conf-name>2024 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW); 2024 May 27&#x2013;31</conf-name>; <publisher-loc>San Francisco, CA, USA</publisher-loc>. p. <fpage>1184</fpage>&#x2013;<lpage>1186</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DA-PFL: dynamic affinity aggregation in personalized federated learning under class imbalance</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2025</year>;<volume>36</volume>:<fpage>20184</fpage>&#x2013;<lpage>98</lpage>; <pub-id pub-id-type="pmid">40902044</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ng</surname> <given-names>S</given-names></string-name></person-group>. <article-title>One-shot sequential federated learning for non-IID data by enhancing local model diversity</article-title>. In: <conf-name>MM &#x2019;24: Proceedings of the 32nd ACM International Conference on Multimedia</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2024</year>. p. <fpage>5201</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Sequential federated learning in hierarchical architecture on non-IID datasets</article-title>. <source>IEEE Trans Mob Comput</source>. <year>2024</year>;<volume>24</volume>(<issue>10</issue>):<fpage>11110</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmc.2025.3573928</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>StarCPFL: star-centric personalized federated learning with layer-wised clustering</article-title>. <source>Future Gener Comput Syst</source>. <year>2025</year>;<volume>175</volume>:<fpage>108037</fpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Criado</surname> <given-names>M</given-names></string-name>, <string-name><surname>Casado</surname> <given-names>F</given-names></string-name>, <string-name><surname>Iglesias</surname> <given-names>R</given-names></string-name>, <string-name><surname>Regueiro</surname> <given-names>C</given-names></string-name>, <string-name><surname>Barro</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Non-IID data and continual learning processes in federated learning: a long road ahead</article-title>. <source>Inf Fusion</source>. <year>2022</year>;<volume>88</volume>(<issue>3</issue>):<fpage>263</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2022.07.024</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>G</given-names></string-name>, <string-name><surname>Jeong</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bae</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yun</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Preservation of the global knowledge by not-true distillation in federated learning</article-title>. <comment>arXiv:2106.03097. 2022</comment>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Learning critically: selective self-distillation in federated learning on non-IID data</article-title>. <source>IEEE Trans Big Data</source>. <year>2024</year>;<volume>10</volume>:<fpage>789</fpage>&#x2013;<lpage>800</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Arafeh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hammoud</surname> <given-names>A</given-names></string-name>, <string-name><surname>Guizani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mourad</surname> <given-names>A</given-names></string-name>, <string-name><surname>Otrok</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ould-Slimane</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>WFSL: warmup-based federated sequential learning</article-title>. <source>IEEE Internet Things J</source>. <year>2025</year>;<volume>12</volume>:<fpage>1974</fpage>&#x2013;<lpage>89</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Abadi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goodfellow</surname> <given-names>I</given-names></string-name>, <string-name><surname>McMahan</surname> <given-names>HB</given-names></string-name>, <string-name><surname>Mironov</surname> <given-names>I</given-names></string-name>, <string-name><surname>Talwar</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep learning with differential privacy</article-title>. In: <conf-name>Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS); 2016 Oct 24&#x2013;28</conf-name>; <publisher-loc>Vienna, Austria</publisher-loc>. p. <fpage>308</fpage>&#x2013;<lpage>18</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mironov</surname> <given-names>I</given-names></string-name></person-group>. <article-title>R&#x00E9;nyi differential privacy</article-title>. In: <conf-name>Proceedings of the IEEE 30th Computer Security Foundations Symposium (CSF); 2017 Aug 21&#x2013;25</conf-name>; <publisher-loc>Santa Barbara, CA, USA</publisher-loc>. p. <fpage>263</fpage>&#x2013;<lpage>75</lpage>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Truex</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chow</surname> <given-names>KH</given-names></string-name>, <string-name><surname>Gursoy</surname> <given-names>ME</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>W</given-names></string-name></person-group>. <article-title>LDP-Fed: federated learning with local differential privacy</article-title>. In: <conf-name>Proceedings of the EdgeSys Workshop; 2020 Apr 27</conf-name>; <publisher-loc>Heraklion, Greece</publisher-loc>. p. <fpage>61</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Differentially private federated learning with an adaptive noise mechanism</article-title>. <source>IEEE Trans Inf Forensics Secur</source>. <year>2024</year>;<volume>19</volume>:<fpage>74</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tifs.2023.3318944</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ni</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Poor</surname> <given-names>HV</given-names></string-name></person-group>. <article-title>Amplitude-varying perturbation for balancing privacy and utility in federated learning</article-title>. <source>IEEE Trans Inf Forensics Secur</source>. <year>2023</year>;<volume>18</volume>:<fpage>1884</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tifs.2023.3258255</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>F</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>E</given-names></string-name>, <string-name><surname>Han</surname> <given-names>D</given-names></string-name>, <string-name><surname>Brinton</surname> <given-names>CG</given-names></string-name></person-group>. <article-title>Differentially-private multi-tier federated learning: a formal analysis and evaluation</article-title>. <source>IEEE/ACM Trans Netw</source>. <year>2025</year>;<volume>34</volume>:<fpage>2226</fpage>&#x2013;<lpage>41</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bai</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Membership inference attacks and defenses in federated learning: a survey</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>57</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3704633</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Multi-layer defense strategies and privacy preserving enhancements for membership reasoning attacks in a federated learning framework</article-title>. In: <conf-name>Proceedings of the 5th International Conference on Computer Science and Blockchain (CCSB); 2025 Aug 1&#x2013;3</conf-name>; <publisher-loc>Shenzhen, China</publisher-loc>. p. <fpage>278</fpage>&#x2013;<lpage>82</lpage>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ngai</surname> <given-names>ECH</given-names></string-name>, <string-name><surname>Voigt</surname> <given-names>T</given-names></string-name></person-group>. <article-title>An experimental study of Byzantine-robust aggregation schemes in federated learning</article-title>. <source>IEEE Trans Big Data</source>. <year>2024</year>;<volume>10</volume>(<issue>6</issue>):<fpage>975</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tbdata.2023.3237397</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nordlund</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Byzantine-resilient hierarchical federated learning with clustered over-the-air aggregation</article-title>. In: <conf-name>2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) Workshops; 2024 Apr 14&#x2013;19</conf-name>; <publisher-loc>Seoul, Republic of Korea</publisher-loc>. p. <fpage>715</fpage>&#x2013;<lpage>19</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Du</surname> <given-names>W</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Byzantine-robust hierarchical aggregation for cross-device federated learning in consumer IoT</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2025</year>;<volume>71</volume>(<issue>2</issue>):<fpage>6359</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tce.2024.3450649</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>