<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">73647</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.073647</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>KMFC-GWO: A Hybrid Fuzzy-Metaheuristic Algorithm for Privacy-Preservation in Graph-Based Social Networks</article-title>
<alt-title alt-title-type="left-running-head">KMFC-GWO: A Hybrid Fuzzy-Metaheuristic Algorithm for Privacy-Preservation in Graph-Based Social Networks</alt-title>
<alt-title alt-title-type="right-running-head">KMFC-GWO: A Hybrid Fuzzy-Metaheuristic Algorithm for Privacy-Preservation in Graph-Based Social Networks</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Memarian</surname><given-names>Saeideh</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Oprescu</surname><given-names>Andreea M.</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Moreno-Naranjo</surname><given-names>Natalia</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Mir&#x00F3;-Amarante</surname><given-names>Gloria</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Romero-Ternero</surname><given-names>M. Carmen</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref ref-type="aff" rid="aff-3">3</xref><email>mcromerot@us.es</email></contrib>
<aff id="aff-1"><label>1</label><institution>Doctoral Program in Computer Science Engineering, Universidad de Sevilla</institution>, <addr-line>Sevilla</addr-line>, <country>Spain</country></aff>
<aff id="aff-2"><label>2</label><institution>Departamento Tecnolog&#x00ED;a Electr&#x00F3;nica, Universidad de Sevilla</institution>, <addr-line>Sevilla</addr-line>, <country>Spain</country></aff>
<aff id="aff-3"><label>3</label><institution>Instituto de Ingenier&#x00ED;a Inform&#x00E1;tica, Universidad de Sevilla</institution>, <addr-line>Sevilla</addr-line>, <country>Spain</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: M. Carmen Romero-Ternero. Email: <email>mcromerot@us.es</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>4</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>1</issue>
<elocation-id>25</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>09</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>01</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_73647.pdf"></self-uri>
<abstract>
<p>In recent years, the proliferation of social networks has been remarkable, providing a rich source for data mining endeavors. However, a significant challenge lies in safeguarding the privacy of individuals while sharing these databases publicly. Current approaches, such as K-anonymity, L-diversity, and T-closeness, are commonly employed for data anonymization in social networks. However, these techniques entail considerable information loss due to random alterations in the graph-based datasets. To address these limitations, this paper introduces a new anonymization technique called KMFC-GWO, which combines K-Member Fuzzy Clustering with Grey Wolf Optimizer. This integrated method is designed to strengthen the anonymized graph against a range of threats, including identity, attribute, link disclosure, and similarity attacks, while significantly reducing information loss. Within the KMFC-GWO framework, K-member fuzzy c-means clustering is utilized to create well-balanced clusters, each meeting the K-anonymity requirement. Subsequently, the Grey Wolf Optimizer is applied to optimize cluster formation and effectively anonymize the social network graph. The objective function is carefully crafted to minimize both clustering error and information loss, while ensuring adherence to predefined anonymity criteria. Experimentation on three major graph-based social networks extracted from Facebook, Twitter, and YouTube validates the effectiveness of the KMFC-GWO approach. Results demonstrate its ability to significantly reduce information loss in published graph data, while concurrently satisfying requirements for K-anonymity, L-diversity, and T-closeness.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Gaph-based social networks</kwd>
<kwd>privacy preserving</kwd>
<kwd>K-anonymity</kwd>
<kwd>L-diversity</kwd>
<kwd>T-closeness</kwd>
<kwd>fuzzy clustering</kwd>
<kwd>grey wolf optimizar (GWO)</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>MICIU/AEI/10.13039/501100011033</funding-source>
</award-group>
<award-group id="awg2">
<funding-source>ERDF/EU</funding-source>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With advances in technology, social networks have emerged as pervasive platforms for worldwide social interaction, information sharing, and self-expression [<xref ref-type="bibr" rid="ref-1">1</xref>]. While these networks harbor vast amounts of user data that can enhance service quality, they also pose risks to individual privacy due to the presence of sensitive information. Users in social media platforms do not only post textual content but also frequently share personal photographs, videos, and interactions with friends and family members. These data, while voluntarily disclosed, can still pose severe privacy threats. For example, attackers may infer a user&#x2019;s identity by cross-linking their facial images with public databases, or they can reconstruct sensitive information such as a user&#x2019;s political views or medical conditions by analyzing patterns of likes, comments, and social connections. Even when users willingly share content, they usually do not anticipate large-scale data mining or de-anonymization attacks that exploit structural graph information combined with multimedia. This highlights the importance of designing anonymization methods that preserve privacy beyond simple text-based disclosures. Consequently, both users and data owners seek to safeguard data privacy while leveraging its insights for various purposes [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<p>Sharing social network data with data miners necessitates a careful balance between privacy protection and knowledge retention. Anonymizing techniques, such as K-anonymity (KA) and its extensions like L-diversity (LD) and T-closeness (TC), are commonly employed to mitigate privacy risks while preserving data utility [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>]. KA aims to group users into clusters with at least K members to prevent identity disclosure, though it doesn&#x2019;t safeguard against attribute or link disclosure. LD addresses attribute disclosure by ensuring each cluster contains diverse attribute values, while TC focuses on maintaining the global attribute distribution within clusters to mitigate similarity attacks.</p>
<p>Privacy threats in social network data publishing encompass different attacks such as identity, attribute, and link disclosure [<xref ref-type="bibr" rid="ref-5">5</xref>]. Identity disclosure exposes a user&#x2019;s identity, while attribute disclosure reveals sensitive user attributes and link disclosure unveils sensitive relationships between users. LD and TC complement KA by addressing attribute and similarity attacks, respectively, enhancing overall privacy protection in anonymized datasets [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>Despite the significant advancements in privacy-preserving techniques for social network data publishing, several research gaps remain unaddressed. Existing methods like KA, LD, and TC primarily focus on mitigating specific privacy threats, often at the cost of high information loss or computational complexity. While KA is effective in preventing identity disclosure, it falls short in safeguarding against attribute and link disclosures. LD and TC extend KA&#x2019;s capabilities but often struggle with scalability and preserving data utility when applied to large-scale or complex social networks. Moreover, most existing approaches treat privacy threats in isolation, lacking a unified framework that comprehensively addresses identity, attribute, and link disclosures simultaneously. This creates a pressing need for innovative solutions that balance robust privacy protection with minimal information loss, especially in the context of graph-based social networks.</p>
<p>This research concentrates on protecting social network data publication from a range of threats, including revealing identities, disclosing attributes/links, and potential similarity attacks. To address these challenges, we introduce a hybrid anonymization approach called KMFC-GWO, which combines K-member Fuzzy Clustering (KMFC) with Grey Wolf Optimizer (GWO). The main objective of KMFC-GWO is to fortify the privacy of graph-based social networks while reducing information loss. Our approach employs a modified variant of fuzzy c-means (FCM), referred to as KMFC, to create well-balanced clusters containing a minimum of K members in each cluster, thereby fulfilling the KA criterion. Furthermore, an optimization problem is formulated and solved using GWO, to satisfy the LD and TC conditions. The main contributions of this paper can be summarized as follows:
<list list-type="bullet">
<list-item>
<p>Hybrid Anonymization Framework: We propose a hybrid approach (KMFC-GWO) that combines K-member Fuzzy Clustering (KMFC) with the Grey Wolf Optimizer (GWO) to address multiple privacy threats in graph-based social networks.</p></list-item>
<list-item>
<p>Comprehensive Privacy Protection: Our framework ensures robust privacy by simultaneously addressing identity, attribute, and link disclosure risks, going beyond the limitations of traditional anonymization techniques.</p></list-item>
<list-item>
<p>Enhanced Clustering with KMFC: We introduce a modified variant of FCM clustering, i.e., KMFC, to form well-balanced clusters that meet the KA criterion, ensuring a minimum of K members per cluster.</p></list-item>
<list-item>
<p>Optimized Privacy and Utility Balance: Through the integration of the GWO algorithm, we optimize cluster properties to satisfy LD and TC conditions, minimizing information loss while preserving data utility.</p></list-item>
<list-item>
<p>Scalability and Effectiveness: Our approach demonstrates scalability and effectiveness on graph-based social networks, showcasing its potential for practical applications in large-scale data publishing scenarios.</p></list-item>
</list></p>
<p>In the rest of the paper, <xref ref-type="sec" rid="s2">Section 2</xref> reviews the literature on anonymization methods. <xref ref-type="sec" rid="s3">Section 3</xref> introduces the KMFC-GWO methodology. <xref ref-type="sec" rid="s4">Section 4</xref> outlines the simulation of the KMFC-GWO algorithm on three graph-based social networks. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> concludes the paper with future works.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Review</title>
<sec id="s2_1">
<label>2.1</label>
<title>Comparison with LocJury</title>
<p>LocJury is a privacy-preserving framework proposed for protecting location data in Internet of Connected Vehicles (IoCV) environments. It employs identity-based networking and clustering mechanisms to reduce the risk of location disclosure while maintaining service utility [<xref ref-type="bibr" rid="ref-7">7</xref>]. Although LocJury is designed for vehicular location privacy rather than social network anonymization, it is considered here for comparison due to its use of clustering-based privacy enforcement and constraint-driven data protection, which conceptually relate to the objectives of the proposed KMFC-GWO framework. LocJury, an Identity-Based Networking (IBN) location-privacy preservation framework for Internet of Connected Vehicles (IoCV), focuses primarily on protecting spatiotemporal trajectories, trust-based decisions, and contextual location disclosure. Unlike LocJury, which operates in a vehicular mobility environment, KMFC-GWO focuses on anonymization of graph-structured social networks while jointly satisfying K-anonymity, L-diversity, and T-closeness. KMFC-GWO performs optimization on both node attributes and structural edges, providing a multi-constraint anonymization method that addresses risks not covered in LocJury.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Definitions</title>
<p>The characteristics of users within social networks commonly involve both graph properties, which pertain to the connections between users, and matrix properties, relating to personal attributes. Graph-based binary features indicate whether connections exist between users, while matrix-based features are typically classified into three categories [<xref ref-type="bibr" rid="ref-6">6</xref>]:
<list list-type="bullet">
<list-item>
<p>Primary Identifying Attributes: This category encompasses direct identifiers such as full name, surname, and national identification numbers, which can directly expose a user&#x2019;s identity. Consequently, these attributes must be eliminated prior to any additional processing.</p></list-item>
<list-item>
<p>Sensitive Attributes: These attributes hold considerable significance and necessitate enhanced protection, examples include postal codes, specific medical conditions, or salary details. Even the disclosure of a single value or a combination of these attributes could potentially jeopardize user privacy or divulge sensitive information.</p></list-item>
<list-item>
<p>Standard Attributes: Attributes falling into this category exhibit lower sensitivity compared to the aforementioned types and therefore require less rigorous privacy safeguards. Examples include age, gender, and educational background. Prior to applying any anonymization algorithm, it is generally assumed that explicit identifiers have been removed from the dataset. Consequently, anonymization is performed on quasi-identifiers, sensitive attributes, and graph-based features. The concept of k-anonymity and its extensions form the foundation of many privacy-preserving models in data publishing [<xref ref-type="bibr" rid="ref-4">4</xref>]</p></list-item>
<list-item>
<p>KA (K-anonymity): A modified table adheres to the KA condition if no user can be uniquely identified from fewer than <italic>K</italic> &#x2212; 1 other users based on the features. In cluster-based methods, this equates to grouping all users into distinct clusters, ensuring that each cluster comprises a minimum of K samples.</p></list-item>
<list-item>
<p>LD (L-diversity): A cluster or group of users demonstrates LD when there are at least <italic>L</italic> unique values for every sensitive attribute within that cluster. Thus, a revised table achieves LD when each cluster within it satisfies this condition.</p></list-item>
<list-item>
<p>TC (T-closeness): A cluster is considered to adhere to TC conditions when the variance between the distribution of each sensitive attribute within the cluster and the distribution of that attribute across the entire dataset remains below a threshold <italic>T</italic>. If all clusters meet the TC requirements, the revised table is regarded as fulfilling TC criteria.</p></list-item>
</list></p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Existing Methods</title>
<p>The majority of anonymization techniques rely on clustering-based KA, wherein users are clustered into separate groups, each containing at least K members [<xref ref-type="bibr" rid="ref-8">8</xref>]. Graph-based anonymization methods typically involve initial modifications to the network data, followed by propagation of the altered graph. Edge randomization techniques introduce random changes to the graph-based network structure by randomly adding/deleting edges or swapping pairs of edges. This randomization aids in safeguarding users against re-identification through probabilistic means. Graph structure modifications can occur random perturbation or with the goal of satisfying some predetermined constraints such as KA, LD, and TC [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Random unconstrained perturbation involves fundamental techniques that alter graph structures by randomly adding or removing edges. Casas-Roma [<xref ref-type="bibr" rid="ref-10">10</xref>] introduced a spectral approach to safeguard critical network edges while optimizing utility within specified privacy constraints. Although spectral methods enhance utility while minimizing information loss, they entail higher computational demands. Nguyen et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] employed quadratic programming to create uncertain graphs, focusing on edge manipulation for anonymization. However, these techniques, despite their simplicity, overlook the diversity and distribution of sensitive attributes, making them vulnerable to attribute and similarity attacks. Moreover, they may not ensure the K-anonymity condition, as edge modification does not typically consider it.</p>
<p>Kumar and Kumar [<xref ref-type="bibr" rid="ref-12">12</xref>] proposed a K-degree technique based on the upper approximation method of rough sets, utilizing split-join operations to evaluate privacy preservation. They employed various metrics including variation of information, mutual information, and adjusted Rand index to assess utility. Kiabod et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] introduced a time-efficient K-degree method for anonymizing network graphs without the need for rescanning the structure for different anonymity levels. This method constructs a connection tree by computing the degree sequence of the anonymized graph, ensuring it exceeds a specified threshold. While KA and the K-degree approach protect against identity attacks, they do not address threats such as attribute/link and similarity attacks.</p>
<p>Yazdanjue et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed a KA technique for social networks through integrating clustering with evolutionary algorithms. This technique initially clusters nodes in the social network based on similarities. Fitness evaluation follows, assessing cluster fitness considering intra-cluster similarity and inter-cluster dissimilarity. Employing an evolutionary process, clusters with low fitness undergo modification through genetic operators like mutation and crossover. Throughout this process, the algorithm ensures KA, maintaining that each cluster contains at least K indistinguishable individuals. Termination occurs upon finding a satisfactory k-anonymous solution or reaching a predefined stopping criterion.</p>
<p>Langari et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] introduced a K-member Fuzzy Clustering and Firefly Algorithm (KFCFA) approach for privacy preservation in social networks by combining fuzzy clustering with the firefly algorithm. The method begins with fuzzy clustering to group nodes based on similarity, followed by the application of the firefly algorithm to optimize privacy parameters. Through an iterative optimization procedure, this algorithm aims to minimize information disclosure while maintaining network utility. Termination occurs upon reaching a satisfactory privacy-preserving solution.</p>
<p>Rajabzadeh et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] proposed a graph-based modification technique for achieving KA by employing a genetic algorithm. The approach involves iteratively modifying the network&#x2019;s structure to ensure that each generated subgraph satisfies the KA property. By employing a genetic algorithm, this method optimizes the required anonymity conditions and network utility. The algorithm terminates when a satisfactory level of KA is generated.</p>
<p>Clustering-based anonymization strategies commonly rely on K-anonymity as their fundamental privacy requirement, upon which additional diversity and distributional constraints can be enforced to mitigate attribute disclosure risks [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Research Gaps and Our Contributions against the Reviewed Techniques</title>
<p>The existing methods for anonymizing social network data, while effective in addressing specific privacy concerns, exhibit notable limitations. Many clustering-based KA techniques, such as those employing random edge perturbations or K-degree strategies, fail to adequately address threats like attribute and similarity attacks. Methods like spectral approaches or quadratic programming-based uncertain graphs, though utility-focused, often neglect the diversity and distribution of sensitive attributes. Moreover, the high computational cost of these methods makes them impractical for large-scale networks. Evolutionary and metaheuristic-based approaches, such as genetic algorithms or firefly algorithms, offer improved optimization but tend to prioritize specific conditions, such as KA, at the expense of balanced utility and privacy across multiple dimensions like LD and TC. These limitations highlight a critical gap in developing comprehensive methods that achieve robust privacy protection while maintaining low information loss and ensuring balanced clustering.</p>
<p>Our proposed KMFC-GWO anonymization technique addresses the mentioned gaps by combining KMFC with GWO to form a hybrid method capable of meeting KA, LD, and TC conditions simultaneously. Unlike existing methods, KMFC-GWO incorporates a modified fuzzy c-means clustering algorithm to ensure well-balanced clusters, where each cluster contains at least K members, thereby guaranteeing the KA condition. Additionally, our method formulates privacy protection as an optimization problem, leveraging GWO to optimize LD and TC constraints. This dual-layer approach effectively mitigates identity, attribute, and similarity attacks while preserving the structural and attribute-based integrity of the data. Furthermore, KMFC-GWO demonstrates superior computational efficiency and scalability compared to existing metaheuristic approaches, making it a practical solution for large-scale social network datasets. By achieving a balanced trade-off between privacy and utility, KMFC-GWO represents a significant advancement over existing anonymization techniques.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed KMFC-GWO Algorithm</title>
<p>In this section, we present a hybrid fuzzy-metaheuristic graph-based privacy-preserving approach for social networks, termed KMFC-GWO, which combines KMFC with GWO, offering a highly effective solution for both balanced clustering and anonymization in social networks. Our method addresses the optimization problem of privacy preservation within graph-based social networks by integrating KMFC, considering KA, LD, and TC criteria. Subsequently, GWO is leveraged to tackle this optimization challenge, aiming to maximize anonymity while minimizing information loss in the published graph. Specifically, the objective function of GWO is formulated to minimize information loss and uphold anonymity conditions (KA, LD, and TC). The proposed KMFC-GWO approach effectively safeguards the anonymized social network graph against identity, attribute disclosure, and similarity attacks by enforcing these three constraints.</p>
<p>Consider a social network represented as a graph, where the edges (referred to as G) have dimensions N &#x00D7; N, and the attribute data matrix (referred to as A) has dimensions N &#x00D7; M. Here, N denotes the number of users, and M denotes the number of personal characteristics associated with each user. Each user, denoted as i (where i ranges from 1 to N), possesses a vector of edges indicating connections with other users and a vector of features represented by M. The graph matrix, G, is binary: Gij &#x003D; 1 signifies the existence of an edge between users i and j, while Gij &#x003D; 0 denotes the absence of such an edge. Furthermore, Aim represents the m-th personal characteristic of user i. The aim of the anonymization procedure is to alter the original data to generate a modified edge graph (denoted as Gnew) and a modified data matrix (denoted as Anew). Before introducing the proposed KMFC-GWO algorithm, <xref ref-type="fig" rid="fig-1">Fig. 1</xref> provides a simple example of a graph-based social network, where users are modeled as nodes and their interactions as edges.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>An illustrative example of a graph-based social network, used to demonstrate how users and their interactions are modeled as a graph in real-world social platforms</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_73647-fig-1.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> is not intended as a purely hypothetical illustration, but rather as an intuitive abstraction of the types of real graph-based social networks analyzed in this study. The social platforms used in our experiments&#x2014;Twitter, Facebook, and YouTube&#x2014;can each be rigorously modeled as graphs, where users constitute the nodes and various forms of interactions define the edges. In Twitter, user relationships form a directed graph through the follower&#x2013;following mechanism, while mentions, replies, and retweets generate additional interaction edges. Facebook predominantly exhibits an undirected friendship graph, yet its structure is further enriched through comments, likes, and group interactions. Although primarily a content-sharing platform, YouTube also functions as a social network: users can subscribe to other users, interact through comments, share content within communities, and therefore form user&#x2013;user interaction graphs as well as user&#x2013;content relational structures. Consequently, the anonymization problem addressed in this work inherently belongs to the domain of graph-based social network modeling. The heterogeneous connectivity, neighborhood structures, and community-like patterns reflected in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> are directly representative of the structural properties present in these real-world platforms. This establishes a clear rationale for adopting a graph-based anonymization framework such as KMFC-GWO and justifies the applicability of our proposed method to Twitter, Facebook, and YouTube datasets.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Example of a graph-based social network with users U01&#x2013;U10 and their interaction links</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_73647-fig-2.tif"/>
</fig>
<p>In the KMFC-GWO approach, the K-anonymity requirement is met through KMFC forming C clusters, each comprising a minimum of K users. Then, the GWO algorithm is utilized to address the anonymity criteria of LD and TC, while simultaneously reducing information loss in the published social network graph.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Satisfying KA Using KMFC</title>
<p>As mentioned above, to satisfy the KA condition, we apply a KMFC technique on the FCM. Our KMFC method utilizes a customized version of the FCM algorithm to construct balanced clusters with at least K members, satisfying the KA condition. To achieve this purpose, first, we generate an initial clustering using FCM, and then, the generated clusters are balanced to satisfy the KA condition.</p>
<p>The FCM algorithm was initially introduced by Bezdek in 1981 [<xref ref-type="bibr" rid="ref-16">16</xref>]. Unlike traditional clustering techniques such as c-means and k-means, FCM assigns a degree of membership for each data point to every cluster. The objective of FCM is to minimize the total distance between the data points and the cluster centroids. Its primary aim is to partition N data points into C distinct clusters. Each data point can be characterized by two feature vectors: binary edges and personal attributes. More specifically, <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is the input matrix of dimension <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow></mml:math></inline-formula> is the number of clustering attributes, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow></mml:math></inline-formula> is the number of all users, and M is the number of personal features. Also, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>C</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is the centroid of clusters (<inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mtext>C</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:math></inline-formula>). The objective of FCM can be formulated as follows [<xref ref-type="bibr" rid="ref-17">17</xref>]:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>O</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>C</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>iF</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>] represents sample <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>i</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>kf</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>kF</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>] is the center of cluster <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>k</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>q</mml:mi></mml:math></inline-formula> (<italic>q</italic> &#x003E; 1) is the exponential weight of FCM that adjusts the fuzziness degree, and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is the membership of node <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>i</mml:mi></mml:math></inline-formula> to cluster <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>k</mml:mi></mml:math></inline-formula> which should fulfil the condition of <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. Also, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> measures the Euclidian distance between <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, as formulated in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>. To optimize the objective function of the FCM algorithm described in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>, iterative partial derivations of <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>O</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>C</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> with respect to <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> need to be computed as per <xref ref-type="disp-formula" rid="eqn-4">Eqs. (4)</xref> and <xref ref-type="disp-formula" rid="eqn-5">(5)</xref>.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mtext>&#xA0;&#xA0;</mml:mtext><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mfrac><mml:mn>2</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>The termination criterion for the FCM algorithm is determined by the condition <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow></mml:math></inline-formula>, where <italic>t</italic> represents the current iteration count, and &#x03B5; is a small positive value (we set &#x03B5; &#x003D; 0.001). Upon termination of the FCM algorithm, each data point (user) is allocated to the cluster with the highest fuzzy membership. Generally, the traditional KA-based clustering methods suffer from two drawbacks:
<list list-type="bullet">
<list-item>
<p>Number of clusters: Selecting the best value for the number of clusters <italic>C</italic> is a challenging issue.</p></list-item>
<list-item>
<p>Unbalanced clusters: In certain clusters, the sample count may be lower than K, while in others, there could be significantly more members than K.</p></list-item>
</list></p>
<p>Our proposed revision phase of the KMFC algorithm is provided in Algorithm 1. To handle the first problem, we consider the number of clusters in a such a way that minimizes the CAVG criterion [<xref ref-type="bibr" rid="ref-6">6</xref>], as formulated in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>. The CAVG metric (where CAVG &#x2265; 1) quantifies the degree of cluster balance, with values closer to 1 indicating higher balance among clusters. In the case of K-member clustering, CAVG tends to approximate 1, i.e., CAVG &#x2248; 1.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>V</mml:mi><mml:mi>G</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mi>N</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Since <italic>N</italic> and <italic>K</italic> are fixed parameters, we consider <italic>C</italic> <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo>&#x2248;</mml:mo></mml:math></inline-formula> <italic>N</italic>/<italic>K</italic>. To have somewhat a free level for clustering, we have set <italic>C</italic> &#x003D; 0.9 &#x00D7; <italic>N</italic>/<italic>K</italic> in our simulations to have around 10% more samples on average in each cluster. Furthermore, to handle the second problem to satisfy the KA condition, we present a revision phase on the initial clustering results of the FCM algorithm.</p>
<fig id="fig-7">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_73647-fig-7.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Impact of KMFC on Cluster Formation</title>
<p>The KMFC component ensures that clusters satisfy the minimum K-member requirement while preserving the inherent structural relationships within the graph. By incorporating both edge-based similarities and attribute-based distances, KMFC groups nodes with similar structural and semantic characteristics. The revision phase redistributes borderline nodes from overloaded clusters to underloaded ones, maintaining cluster balance and minimizing distortion. This ensures that the resulting clusters remain coherent and suitable for subsequent LD and TC optimization via the GWO module.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Satisfying LD and TC Using GWO</title>
<p>Once users are clustered into C balanced clusters (with at least K users in each cluster) to meet the KA requirement, GWO is used to further anonymize the social network, considering the requirements of LD and TC. This process is conducted simultaneously on both the graph matrix G and the attribute matrix A. At the graph level, anonymization involves randomly adding or removing edges between users, while at the attribute level, it entails randomly altering the values of user features.</p>
<p>Grey Wolf Optimizer (GWO) is a swarm intelligence&#x2013;based metaheuristic algorithm designed to balance exploration and exploitation in continuous optimization problems. In this study, GWO is adopted as the optimization engine within the proposed framework due to its simple structure and effective search mechanism. The algorithm was originally introduced by Mirjalili et al. in 2014, inspired by the social hierarchy and hunting behavior of grey wolves [<xref ref-type="bibr" rid="ref-18">18</xref>]. Beyond its original formulation, GWO has been successfully applied to various real-world optimization problems, including adaptive fuzzy systems and healthcare-related applications, demonstrating its robustness and practical effectiveness [<xref ref-type="bibr" rid="ref-19">19</xref>]. Since its introduction, GWO has been widely employed in various optimization applications due to its balanced search mechanism and straightforward implementation [<xref ref-type="bibr" rid="ref-18">18</xref>]. As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the algorithm commences by randomly initializing a population of grey wolves. In each iteration, the fitness of the current population is evaluated using a customized fitness function tailored to the particular application. Subsequently, the entire population undergoes adjustments through two phases: attacking prey (exploitation) and searching for prey (exploration), as described in the standard GWO procedure [<xref ref-type="bibr" rid="ref-18">18</xref>]. These main steps of GWO including random generation of initial population, objective function evaluation, and population updating, are described in the following:</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Flowchart of GWO-based optimization procedure for LD and TC anonymization</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_73647-fig-3.tif"/>
</fig>
<p>Generation of initial population: As seen in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, a feasible solution (SOL) can be represented as a matrix of N &#x00D7; (N &#x002B; M) &#x003D; N &#x00D7; F, where each row i represents the edge and attribute modifications in G and A matrices. In the case of edge modification, SOLi,j &#x003D; 1 if the edge between nodes i and j is modified (adding a new edge between nodes i and j or removing a previously connection link between nodes i and j). Furthermore, for the attribute modification, SOLi, N &#x002B; m &#x003D; 1 if the value of the personal feature m is randomly changed for the data of user i.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Representation of a feasible solution in GWO</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_73647-fig-4.tif"/>
</fig>
<p>&#x02022;&#x02003;Objective function evaluation: There is a trade-off between the information loss (due to the modifications in G and A) and the anonymity level. The more changes in the G and A, the higher anonymity level, however by accepting more information loss. To minimize the information loss within the published social network data, we present an objective function to minimize the summation of the information loss of the modified G and A, denoted by <italic>IL</italic><sub><italic>G</italic></sub> and <italic>IL</italic><sub><italic>A</italic></sub>, while satisfying the constraints of KA, LD and TC. More specifically, the objective function of Grey Wolf Optimizer (OFGWO) is defined as follows:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>O</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>W</mml:mi><mml:mi>O</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>minimize</mml:mtext></mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>I</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>I</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>P</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula>subject to:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>I</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mi>M</mml:mi><mml:mi>D</mml:mi><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>I</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mi>M</mml:mi><mml:mi>D</mml:mi><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>M</mml:mi><mml:mi>D</mml:mi><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mspace width="1em" /><mml:mspace width="1em" /><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>edge</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>is modified</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mspace width="1em" /><mml:mspace width="1em" /><mml:mrow><mml:mtext>otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>M</mml:mi><mml:mi>D</mml:mi><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mspace width="1em" /><mml:mspace width="1em" /><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>attribute</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>m</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>of</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>user</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>i</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>is modified</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mspace width="1em" /><mml:mspace width="1em" /><mml:mrow><mml:mtext>otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>P</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>D</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>D</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> &#x003D; 1 and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> &#x003D; 1 if the LD and TC conditions are not satisfied within cluster <italic>k</italic>, respectively. Furthermore, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the relative weights of the information losses within the published graph and attributes, respectively. According to <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>, <italic>Penalty</italic> is measured as the total number of unsatisfied <italic>LD</italic> and <italic>TC</italic> constraints. Considering the penalty function as 2<sup><italic>Penalty</italic></sup> ensures that satisfying the constraints is of higher importance than the main objective function. It ensures obtaining the final solutions with no penalties.
<list list-type="bullet">
<list-item>
<p>Population updating: GWO models a hierarchical social structure consisting of four main levels, namely alpha, beta, delta, and omega wolves. The alpha wolf represents the best solution found so far, followed by beta and delta wolves, which guide the search process, while the remaining woves update their positions accordingly [<xref ref-type="bibr" rid="ref-17">17</xref>]. The alpha assumes the leadership role, with commands to be followed by all other wolves. The beta acts as an advisor to the alpha, supporting the alpha&#x2019;s directives and offering feedback. The delta complies with the alpha and beta but holds dominance over the rest of the pack (omegas). During each iteration of the algorithm, the fitness function assesses the quality of all gray wolves, which are then sorted based on their fitness values, ranging from best to worst. The top three solutions are assigned as alpha (<italic>X</italic><sub><italic>&#x03B1;</italic></sub>), beta (<italic>X</italic><sub><italic>&#x03B2;</italic></sub>), and delta (<italic>X</italic><sub><italic>&#x03B4;</italic></sub>), respectively. In each iteration <italic>t</italic>, the position of each gray wolf <italic>s</italic> is updated according to the following process:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mn>3</mml:mn></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are respectively the factors of the encircling prey of the grey wolf <italic>s</italic> according to <italic>X</italic><sub><italic>&#x03B1;</italic></sub>, <italic>X</italic><sub><italic>&#x03B2;</italic></sub>, and <italic>X</italic><sub><italic>&#x03B4;</italic></sub>, respectively, which can be calculated as follows:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>&#x03B4;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where A &#x003D; 2a.r1-a and C &#x003D; 2.r2 are random vectors of the same dimension as <italic>X</italic><sub><italic>s</italic></sub>, where r<sub>1</sub> and r<sub>2</sub> are uniformly generated random vectors with elements within [0, 1], and &#x03B1; is decreased from 2 to 0 during the execution of the GWO.</p></list-item>
</list></p>
<p>After reaching the maximum iterations of GWO, the global best solution is considered as the final anonymized graph-based social network, which is ready to be sent to the data miners for further processing.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Computational Complexity Analysis</title>
<p>The computational complexity of the proposed KMFC-GWO framework is derived as follows:
KMFC Phase:
O(N &#x00D7; C &#x00D7; I_fcm)
GWO Phase
O(Pop &#x00D7; Iter &#x00D7; F)
where F &#x003D; N &#x002B; M represents the dimensionality of the solution space.
Total Complexity:
O(NCI_fcm &#x002B; Pop &#x00D7; Iter &#x00D7; (N &#x002B; M))
This combined complexity shows that the framework scales linearly with the size of the network and the attribute dimensions.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Simulation Results</title>
<p>This section presents the experimental evaluation of the proposed KMFC-GWO framework. We begin by reporting structural utility preservation results, followed by information-loss evaluation, cluster-level anonymization metrics, and runtime analysis. Subsequently, we present attack-resilience experiments, where three widely-studied re-identification and inference attacks are simulated to assess the robustness of the anonymized graphs. Each subsection includes detailed methodological explanations, quantitative results, and discussion of the scientific implications of KMFC-GWO.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Settings</title>
<p>The KMFC-GWO algorithm was effectively implemented using MATLAB R2022b on a Windows 10 PC equipped with a 2.6 GHz Core i7 CPU and 16 GB of RAM. Since GWO is a self-adjustable algorithm, we should only set the population size and iteration count of the algorithm to ensure enough function evaluation leading to a proper convergence in terms of OFGWO. In our simulations, we set the population size of GWO as 50 and maximum iterations as 200. The both weights <italic>W</italic><sub><italic>G</italic></sub> and <italic>W</italic><sub><italic>A</italic></sub> were set to 1 (<italic>W</italic><sub><italic>G</italic></sub> &#x003D; <italic>W</italic><sub><italic>A</italic></sub> &#x003D; 1), ensuring that the same weights are specified for the edge and attribute modifications. However, these weights can have any other relative values, based on the requirements specified by the system manager.</p>
<p>The datasets considered in this study are subsets of three large-scale social networks (Twitter, Facebook, and YouTube). Each dataset is represented in the form of a graph structure where nodes denote users and edges represent their relationships (e.g., friendships, subscriptions, or follow links). Alongside the graph matrices, we used an attribute matrix that contains numerical and categorical features such as age, gender, or activity-related metadata. While many social media users post images and videos, the publicly available benchmark datasets primarily provide graph and attribute information, not raw multimedia content. We therefore focus on graph-based anonymization, which remains highly relevant because the structural and attribute-based information alone is sufficient to re-identify users or disclose sensitive relations, as shown in prior privacy attacks. Due to the extensive size of these networks in terms of user count and edges, a subset of each network was selected for the simulations, as outlined in <xref ref-type="table" rid="table-1">Table 1</xref>. Each dataset includes an attribute matrix A, where each row represents a feature vector containing the personal attributes of a user. Additionally, there is an edge graph matrix G, depicting the connection edges between users. While many social media users also share images and videos, the benchmark datasets used in this study primarily provide graph structures and attribute information, not raw multimedia content. Therefore, our evaluation focuses on graph-based anonymization, which remains highly relevant because structural and attribute-based data alone have been shown to enable user re-identification, attribute disclosure, or sensitive relationship inference in prior works. The KMFC-GWO algorithm was then employed to anonymize each social network, considering anonymity parameters set to K &#x003D; 6, L &#x003D; 4, and T &#x003D; 0.5.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Social networks used in this paper, based on [<xref ref-type="bibr" rid="ref-6">6</xref>]</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>No. users</th>
<th>No. features</th>
<th>No. edges</th>
</tr>
</thead>
<tbody>
<tr>
<td>Twitter</td>
<td>244</td>
<td>1364</td>
<td>3621</td>
</tr>
<tr>
<td>Facebook</td>
<td>347</td>
<td>224</td>
<td>5038</td>
</tr>
<tr>
<td>YouTube</td>
<td>450</td>
<td>47</td>
<td>3704</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Results of KMFC</title>
<p>The results of the KMFC algorithm for different datasets are reported in <xref ref-type="table" rid="table-2">Table 2</xref>. As mentioned above, we considered the number of clusters in such a way that minimizes the CAVG criterion as much as possible, which means generating more balanced clusters. By performing the KMFC algorithm on Twitter, Facebook, and YouTube datasets, CAVG has been obtained as 1.099, 1.112, and 1.103, respectively. It shows a proper generation of balanced clusters using the KMFC algorithm. Another point is that as expected, the KA condition has been satisfied in all datasets, however, the conditions of LD and TC are still not fulfilled, which need further anonymization using the GWO algorithm.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Results of clustering using KMFC algorithm</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th><italic>K</italic></th>
<th><italic>C</italic></th>
<th><italic>CAVG</italic></th>
<th>KA</th>
<th>LD</th>
<th>TC</th>
</tr>
</thead>
<tbody>
<tr>
<td>Twitter</td>
<td>6</td>
<td>37</td>
<td>1.099</td>
<td>&#x2713;</td>
<td><bold>&#x00D7;</bold></td>
<td><bold>&#x00D7;</bold></td>
</tr>
<tr>
<td>Facebook</td>
<td>6</td>
<td>52</td>
<td>1.112</td>
<td>&#x2713;</td>
<td><bold>&#x00D7;</bold></td>
<td><bold>&#x00D7;</bold></td>
</tr>
<tr>
<td>YouTube</td>
<td>6</td>
<td>68</td>
<td>1.103</td>
<td>&#x2713;</td>
<td><bold>&#x00D7;</bold></td>
<td><bold>&#x00D7;</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Results of GWO</title>
<p>After satisfying the KA condition using the proposed KMFC algorithm, GWO is applied for further anonymization to fulfill the LD and TC conditions. The convergence of the GWO algorithm for the Twitter dataset is provided in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The figure shows that the algorithm started from an initial population with the average and best objective function values of 0.0491 and 0.0458, and after an iterative-based optimization procedure with 200 iterations, the algorithm converged finally to the global best solution with OFGWO &#x003D; 0.0374.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Convergence of GWO for Twitter dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_73647-fig-5.tif"/>
</fig>
<p>To justify the performance of the GWO algorithm against other metaheuristic algorithms, we have considered the same optimization problem to be solved using Genetic Algorithm (GA) [<xref ref-type="bibr" rid="ref-20">20</xref>], and Greylag Goose Optimization (GGO) [<xref ref-type="bibr" rid="ref-21">21</xref>]. A comparison of the objective function value of the GWO algorithm with GA, AO, and GGO, on the different datasets is provided in <xref ref-type="table" rid="table-3">Table 3</xref>. The results demonstrate that the AO and GGO have a slightly better performance than the GA. However, the GWO outperforms all compared metaheuristics in all datasets. To better understand the superiority of the GWO compared to other metaheuristics, the gain (improvement rate) of the GWO against GA, AO, and GGO is illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref> for different datasets.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of different metaheuristic algoruithms</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>GA</th>
<th>AO</th>
<th>GGO</th>
<th>GWO</th>
</tr>
</thead>
<tbody>
<tr>
<td>Twitter</td>
<td>0.0440</td>
<td>0.0407</td>
<td>0.0412</td>
<td>0.0374</td>
</tr>
<tr>
<td>Facebook</td>
<td>0.0306</td>
<td>0.0288</td>
<td>0.0285</td>
<td>0.0245</td>
</tr>
<tr>
<td>YouTube</td>
<td>0.0397</td>
<td>0.0369</td>
<td>0.0376</td>
<td>0.0348</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Improvement % of the GWO algorithm against GA, AO, and GGO in terms of the objective function value on different datasets</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_73647-fig-6.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Results of KMFC-GWO</title>
<p>To justify the proposed KMFC-GWO algorithm in terms of the anonymity results and clustering performance, the results of the proposed method are compared with KA, 3-Layer T-closeness L-diversity K-anonymity (3L-TLK), and KFCFA. Comparison of the obtained results in terms of the information loss of the attribute matrix A (ILA), information loss of the graph G (ILG), and the clustering balance metric (CAVG) are reported in <xref ref-type="table" rid="table-4">Tables 4</xref>&#x2013;<xref ref-type="table" rid="table-6">6</xref>, respectively. The results presented in <xref ref-type="table" rid="table-4">Tables 4</xref>&#x2013;<xref ref-type="table" rid="table-6">6</xref> illustrate the effectiveness of the proposed KMFC-GWO technique in comparison to K-anonymity (KA), 3L-TLK, and KFCFA across three datasets: Twitter, Facebook, and YouTube. These tables evaluate the methods based on information loss in the attribute and graph matrices, as well as clustering balance.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparison of the information loss of the attribute matrix (ILA) for different techniques</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>KA</th>
<th>3L-TLK</th>
<th>KFCFA</th>
<th>KMFC-GWO</th>
</tr>
</thead>
<tbody>
<tr>
<td>Twitter</td>
<td>0</td>
<td>0.0302</td>
<td>0.0321</td>
<td>0.0260</td>
</tr>
<tr>
<td>Facebook</td>
<td>0</td>
<td>0.0324</td>
<td>0.0239</td>
<td>0.0151</td>
</tr>
<tr>
<td>YouTube</td>
<td>0</td>
<td>0.0274</td>
<td>0.0325</td>
<td>0.0243</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comparison of the information loss of the graph matrix (ILG) for different techniques</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>KA</th>
<th>3L-TLK</th>
<th>KFCFA</th>
<th>KMFC-GWO</th>
</tr>
</thead>
<tbody>
<tr>
<td>Twitter</td>
<td>0</td>
<td>0.0135</td>
<td>0.0191</td>
<td>0.0114</td>
</tr>
<tr>
<td>Facebook</td>
<td>0</td>
<td>0.0164</td>
<td>0.0126</td>
<td>0.0094</td>
</tr>
<tr>
<td>YouTube</td>
<td>0</td>
<td>0.0175</td>
<td>0.0142</td>
<td>0.0105</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparison of the clustering balance metric (CAVG) for different techniques</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>KA</th>
<th>3L-TLK</th>
<th>KFCFA</th>
<th>KMFC-GWO</th>
</tr>
</thead>
<tbody>
<tr>
<td>Twitter</td>
<td>3.89</td>
<td>3.82</td>
<td>1.28</td>
<td>1.099</td>
</tr>
<tr>
<td>Facebook</td>
<td>4.47</td>
<td>5.03</td>
<td>1.32</td>
<td>1.112</td>
</tr>
<tr>
<td>YouTube</td>
<td>4.09</td>
<td>3.40</td>
<td>1.44</td>
<td>1.103</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>When it comes to minimizing information loss in the attribute matrix (<xref ref-type="table" rid="table-4">Table 4</xref>), KMFC-GWO consistently outperforms the alternative techniques. For instance, in the Facebook dataset, KMFC-GWO achieves an information loss value (ILA) of 0.0151, which is significantly lower than those reported for 3L-TLK (0.0324) and KFCFA (0.0239). Similar trends are observed in the Twitter and YouTube datasets, where KMFC-GWO consistently delivers the best performance, reflecting its ability to better preserve data utility.</p>

<p>A similar advantage is observed in minimizing information loss in the graph matrix (<xref ref-type="table" rid="table-5">Table 5</xref>). KMFC-GWO exhibits superior results by achieving the lowest ILG values across all datasets. For example, in the Twitter dataset, KMFC-GWO reports an ILG of 0.0114, outperforming 3L-TLK (0.0135) and KFCFA (0.0191). This pattern of reduced information loss highlights the robustness of KMFC-GWO in preserving the structural integrity of graph-based data.</p>

<p>In terms of clustering balance (<xref ref-type="table" rid="table-6">Table 6</xref>), KMFC-GWO demonstrates exceptional performance, achieving significantly lower clustering balance metric (CAVG) values compared to its counterparts. For example, in the YouTube dataset, KMFC-GWO achieves a CAVG of 1.103, which is notably better than KFCFA (1.44) and 3L-TLK (3.40). This indicates that KMFC-GWO not only protects privacy effectively but also ensures well-balanced clusters, which is critical for maintaining data utility.</p>

<p>The results show the superiority of the KMFC-GWO algorithm against the compared methods. Since KA considers just the clustering of users into distinguished groups, it does not generate any information loss neither within the attribute matrix nor graph matrix. However, as it does not consider the LD and TC conditions, it cannot protect the published data against attribute/link and similarity attacks. However, the other techniques have some distortions in the attribute and graph matrices, among them, the KMFC-GWO algorithm obtained the least information loss. Another remark is that the KMFC-GWO outperforms by far all techniques in terms of CAVG, by generating balanced clusters using the KMFC algorithm in the first phase of the proposed algorithm. These results underscore the advantages of KMFC-GWO as a powerful privacy-preserving approach that excels in reducing information loss in both attribute and graph matrices while achieving superior clustering balance, making it a reliable and efficient solution for publishing anonymized social network data.</p>
<p>Although KMFC-GWO is compared primarily with GA, AO, and GGO in the experimental results, further justification is provided for not adopting Particle Swarm Optimization (PSO) and Ant Colony Optimization (ACO) (<xref ref-type="table" rid="table-7">Table 7</xref>). These methods either lack the exploitation intensity required to satisfy LD/TC constraints or impose computational costs unsuitable for high-dimensional graph anonymization.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparison of optimization heuristics</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Algorithm</th>
<th>Strength</th>
<th>Weakness</th>
<th>Reason not used</th>
</tr>
</thead>
<tbody>
<tr>
<td>GA</td>
<td>Strong exploration</td>
<td>Slow convergence</td>
<td>Not suitable for LD/TC constraints</td>
</tr>
<tr>
<td>PSO</td>
<td>Fast convergence</td>
<td>Prone to local optima</td>
<td>LD/TC require strong exploitation</td>
</tr>
<tr>
<td>ACO</td>
<td>Efficient in path search</td>
<td>High computational overhead</td>
<td>Not suitable for matrix-level optimization
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To further illustrate the motivation and effectiveness of KMFC-GWO, we provide two examples. First, consider a Facebook subgraph where a small community of users share sensitive attributes (e.g., medical conditions). Without anonymization, even if the users voluntarily posted these attributes, attackers could easily re-identify them through structural overlaps with other public networks. Second, when applying our algorithm, the anonymized graph retains its utility for data mining tasks while substantially lowering the risk of identity and attribute disclosure compared to 3L-TLK and KFCFA. This demonstrates that KMFC-GWO not only improves clustering balance and reduces information loss but also delivers practical protection against real-world privacy threats in social media environments.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Attack Resilience Evaluation</title>
<p>To evaluate the robustness of the anonymized graphs, we simulate three widely-used attack models that represent realistic adversarial behavior in social networks: degree-based re-identification, neighborhood-based structural attacks, and attribute inference attacks. These attacks were selected because they are among the most commonly used adversarial strategies in the social network anonymization literature and target different aspects of user privacy.
<list list-type="bullet">
<list-item>
<p>Degree-based re-identification exploits the uniqueness of node degrees. Attackers match a node in the anonymized graph to its unique or rare degree signature in the original graph.</p></list-item>
<list-item>
<p>Neighborhood attacks make use of local structural similarity by comparing ego-network patterns such as clustering coefficient, neighbor sets, or motif frequencies.</p></list-item>
<list-item>
<p>Attribute inference attacks attempt to guess sensitive user attributes by leveraging structural correlation and cluster homogeneity.</p></list-item>
</list></p>
<p>For each attack, we implement a baseline adversary that has partial background knowledge of the original graph. The attacker attempts to map anonymized nodes back to their true identities using structural or attribute-based fingerprints. <xref ref-type="table" rid="table-8">Table 8</xref> reports the attack success rate before and after anonymization. As shown, KMFC-GWO reduces adversarial success by 63%&#x2013;85% across all attack categories. This improvement results from both the KMFC clustering process, which ensures attribute diversity, and the GWO optimization mechanism, which disrupts structural uniqueness while preserving overall utility.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Attack success rate before and after KMFC-GWO</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Attack type</th>
<th>Original success (%)</th>
<th>After KMFC-GWO (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Degree-based Re-ID</td>
<td>72</td>
<td>9</td>
</tr>
<tr>
<td>Neighborhood Attack</td>
<td>63</td>
<td>6</td>
</tr>
<tr>
<td>Attribute Inference Attack</td>
<td>59</td>
<td>11</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Runtime Analysis</title>
<p>The information-loss evaluation examines how the anonymization process affects both structural and attribute components of the graph. Instead of relying solely on global IL scores, we separately analyze edge-modification rates, attribute-distortion rates, and their contribution to the final IL measure. A detailed breakdown demonstrates that KMFC-GWO achieves lower IL compared to baseline anonymization techniques due to its joint optimization of LD and TC constraints. The results confirm that KMFC-GWO introduces the minimum structural perturbation necessary for satisfying privacy requirements. <xref ref-type="table" rid="table-9">Table 9</xref> reports the execution time of KMFC and GWO for the three evaluated datasets.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Runtime of KMFC-GWO</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>KMFC time (s)</th>
<th>KMFC-GWO</th>
<th>Total time (s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Twitter</td>
<td>0.48</td>
<td>1.52</td>
<td>2.00</td>
</tr>
<tr>
<td>Facebook</td>
<td>0.62</td>
<td>2.14</td>
<td>2.76</td>
</tr>
<tr>
<td>YouTube</td>
<td>0.70</td>
<td>2.60</td>
<td>3.30</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Structural Utility Preservation</title>
<p>To further examine utility, we evaluate how anonymization impacts downstream data-mining tasks, including community detection and link prediction. The modularity and AUC scores show that KMFC-GWO preserves the major community structure of the graph while introducing minimal distortion to link-formation patterns. This confirms that the structural adjustments introduced by GWO do not compromise the interpretability of the anonymized network. <xref ref-type="table" rid="table-10">Table 10</xref> reports the preservation of key graph&#x2013;topology metrics before and after anonymization.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Structural utility preservation</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th></th>
<th>Original graph</th>
<th>KMFC-GWO</th>
<th>Change (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clustering coefficient</td>
<td></td>
<td>0.214</td>
<td>0.207</td>
<td>&#x2212;3.3%</td>
</tr>
<tr>
<td>Average path length</td>
<td></td>
<td>3.91</td>
<td>4.02</td>
<td>&#x002B;2.8%</td>
</tr>
<tr>
<td>Diameter</td>
<td></td>
<td>8</td>
<td>8</td>
<td>0%</td>
</tr>
<tr>
<td>Degree distribution (KL-divergence)</td>
<td></td>
<td>&#x2212;0.612</td>
<td>0.031, 0.598</td>
<td>&#x2212;2.3%</td>
</tr>
<tr>
<td>Modularity</td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>Link prediction (AUC)</td>
<td></td>
<td>0.842</td>
<td>0.819</td>
<td>&#x2212;2.7%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To further evaluate the utility of the anonymized graphs, several structural metrics were measured before and after applying KMFC-GWO. <xref ref-type="table" rid="table-10">Table 10</xref> reports the preservation of key topological properties including clustering coefficient, path length, diameter, modularity, and link-prediction performance.</p>

</sec>
<sec id="s4_8">
<label>4.8</label>
<title>Discussion&#x2014;Dynamic Networks</title>
<p>We also analyze the scalability of KMFC-GWO by reporting execution times for KMFC clustering, GWO optimization, and the overall anonymization process across three datasets. The runtime grows approximately linearly with the number of nodes, consistent with the complexity analysis presented in <xref ref-type="sec" rid="s3">Section 3</xref>. The results demonstrate that KMFC-GWO is suitable for medium-scale social networks and can be extended to larger datasets with parallelization. Although KMFC-GWO is designed for static social networks, it can be extended to dynamic environments where nodes and edges evolve over time. Incremental KMFC techniques can update cluster membership efficiently without full recalculation, while GWO can be re-initialized with partial populations to optimize only affected regions of the network. These extensions enable efficient anonymization of evolving graphs, consistent with dynamic anonymization approaches in recent literature.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>In this paper, a novel anonymization method (KMFC-GWO) has been presented. It combines a K-Member fuzzy clustering with a metaheuristic-driven optimization algorithm to enhance the resilience of anonymized graph-based social networks against various threats while minimizing information loss of the published attribute and graph matrices. By presenting a K-member variant of the fuzzy c-means clustering algorithm to achieve K-anonymity and GWO to further optimize the anonymity conditions, the proposed framework effectively anonymizes graph-based social networks. The simulation experiments performed on datasets sourced from Facebook, Twitter, and YouTube confirm the effectiveness of the KMFC-GWO algorithm proposed in reducing information loss while simultaneously fulfilling requirements for K-anonymity, L-diversity, and T-closeness conditions. Compared to existing anonymity methods, KMFC-GWO demonstrates superior performance, notably in minimizing information loss and generating balanced clusters. Despite the promising results, the KMFC-GWO method has some limitations. First, the two-phase structure of the framework&#x2014;clustering followed by optimization&#x2014;introduces additional computational complexity, which may limit its scalability to extremely large-scale social networks. Second, while the algorithm achieves K-anonymity, L-diversity, and T-closeness, it assumes static graphs and does not account for dynamic updates or evolving network structures. Additionally, the reliance on specific parameter settings in both the clustering and optimization phases may affect its generalizability across highly diverse datasets. To address these limitations, future research can explore more scalable and dynamic approaches that adapt to real-time changes in social networks. Integrating adaptive or incremental clustering techniques could enhance the method&#x2019;s applicability to evolving datasets. Moreover, hybridizing KMFC and GWO into a unified or parallel framework could improve efficiency and eliminate the need for a sequential two-step process. Extending the model to consider additional privacy metrics or adversarial scenarios may also further strengthen its robustness in practical applications. The KMFC-GWO framework strengthens privacy protection by simultaneously mitigating identity, attribute, and link disclosure risks. KMFC ensures that each anonymized cluster contains at least K similar users, reducing the likelihood of identity disclosure. By distributing diverse sensitive attributes within clusters, the method reduces attribute inference risks. Additionally, GWO-driven optimization introduces minimal structural modifications to the graph while preserving essential topological patterns, significantly limiting link re-identification attacks. Together, these mechanisms provide a multi-layered defense aligned with practical attack models in social networks. Finally, although users may voluntarily share personal content such as photos and updates, they rarely anticipate the unintended use of their data at scale. Automated profiling, targeted advertising, and cross-network re-identification are all realistic threats. Hence, privacy-preserving algorithms such as KMFC-GWO are essential to safeguard individuals&#x2019; rights and prevent misuse of their data, even in open and highly interactive environments like online social networks. To further enhance privacy evaluation, several complementary metrics can be integrated into KMFC-GWO. These include &#x0394;-Disclosure Risk, Structural Similarity Leakage Index, and Mutual Information Leakage. Incorporating these metrics would expand the evaluation scope beyond KA, LD, and TC, enabling deeper analysis of adversarial risks.</p>
</sec>
</body>
<back>
<ack>
<p>This work was partially funded by grant PID2022-141045OB-{C41,C42,C43} funded by MCIN/AEI/10.13039/501100011033/ and Feder a way of making Europe in artifacts project: Generation of Reliable Synthetic Health Data for Federated Learning in Secure Data Spaces.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This publication is part of the project PID2022-141045OB-C42, funded by MICIU/AEI/10.13039/501100011033 and by ERDF/EU.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and methodology: Saeideh Memarian, Gloria Mir&#x00F3;-Amarante; algorithm design and implementation: Saeideh Memarian; experimental evaluation and analysis: Saeideh Memarian, Andreea M. Oprescu; interpretation of results: Saeideh Memarian, Natalia Moreno-Naranjo, M. Carmen Romero-Ternero; manuscript drafting: Saeideh Memarian; manuscript review and editing: all authors. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The datasets used in this study are publicly available benchmark social network datasets obtained from their original sources. No new datasets were generated during the current study.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>This study does not involve human participants or animal subjects. All experiments were conducted using publicly available datasets in compliance with relevant data usage policies and ethical guidelines.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Duan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>CN</given-names></string-name>, <string-name><surname>Shokouhifar</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Impacts of social media advertising on purchase intention and customer loyalty in E-commerce systems</article-title>. <source>ACM Trans Asian Low-Resour Lang Inf Process</source>. <year>2024</year>;<volume>23</volume>(<issue>8</issue>):<fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3613448</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jain</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Sahoo</surname> <given-names>SR</given-names></string-name>, <string-name><surname>Kaubiyal</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Online social networks security and privacy: comprehensive review and analysis</article-title>. <source>Complex Intell Syst</source>. <year>2021</year>;<volume>7</volume>(<issue>5</issue>):<fpage>2157</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s40747-021-00409-7</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gangarde</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>A</given-names></string-name>, <string-name><surname>Pawar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Joshi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gonge</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Privacy preservation in online social networks using multiple-graph-properties-based clustering to ensure k-anonymity, l-diversity, and t-closeness</article-title>. <source>Electronics</source>. <year>2021</year>;<volume>10</volume>(<issue>22</issue>):<fpage>2877</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics10222877</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sweeney</surname> <given-names>L</given-names></string-name></person-group>. <article-title>K-anonymity: a model for protecting privacy</article-title>. <source>Int J Unc Fuzz Knowl Based Syst</source>. <year>2002</year>;<volume>10</volume>(<issue>5</issue>):<fpage>557</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1142/s0218488502001648</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Panda</surname> <given-names>BS</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>MN</given-names></string-name>, <string-name><surname>Patro</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Apply rough set methods to preserve social networks privacy&#x2014;a review</article-title>. In: <conf-name>Proceedings of 3rd International Conference on Artificial Intelligence: Advances and Applications</conf-name>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer Nature</publisher-name>; <year>2023</year>. p. <fpage>427</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-19-7041-2_34</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Langari</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Sardar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Amin Mousavi</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Radfar</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Combined fuzzy clustering and firefly algorithm for privacy preserving in social networks</article-title>. <source>Expert Syst Appl</source>. <year>2020</year>;<volume>141</volume>:<fpage>112968</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2019.112968</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Du</surname> <given-names>X</given-names></string-name>, <string-name><surname>Guizani</surname> <given-names>N</given-names></string-name></person-group>. <article-title>LocJury: an IBN-based location privacy preserving scheme for IoCV</article-title>. <source>IEEE Trans Intell Transport Syst</source>. <year>2021</year>;<volume>22</volume>(<issue>8</issue>):<fpage>5028</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tits.2020.2970610</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Privacy preservation method based on clustering interference algorithm in social networks</article-title>. <source>J Eng Sci Technol Rev</source>. <year>2022</year>;<volume>15</volume>(<issue>2</issue>):<fpage>191</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.25103/jestr.152.22</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Singh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Social networks privacy preservation: a novel framework</article-title>. <source>Cybern Syst</source>. <year>2024</year>;<volume>55</volume>(<issue>8</issue>):<fpage>2356</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1080/01969722.2022.2151966</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Casas-Roma</surname> <given-names>J</given-names></string-name></person-group>. <chapter-title>Privacy-preserving on graphs using randomization and edge-relevance</chapter-title>. In: <source>Modeling decisions for artificial intelligence</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2014</year>. p. <fpage>204</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-12054-6_18</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>HH</given-names></string-name>, <string-name><surname>Imine</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rusinowitch</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Anonymizing social graphs via uncertainty semantics</article-title>. In: <conf-name>Proceedings of the 10th ACM Symposium on Information, Computer and Communications Security; 2015 Apr 14&#x2013;17</conf-name>; <publisher-loc>Singapore</publisher-loc>. p. <fpage>495</fpage>&#x2013;<lpage>506</lpage>. doi:<pub-id pub-id-type="doi">10.1145/2714576.2714584</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Upper approximation based privacy preserving in online social networks</article-title>. <source>Expert Syst Appl</source>. <year>2017</year>;<volume>88</volume>:<fpage>276</fpage>&#x2013;<lpage>89</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2017.07.010</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kiabod</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dehkordi</surname> <given-names>MN</given-names></string-name>, <string-name><surname>Barekatain</surname> <given-names>B</given-names></string-name></person-group>. <article-title>TSRAM: a time-saving k-degree anonymization method in social network</article-title>. <source>Expert Syst Appl</source>. <year>2019</year>;<volume>125</volume>:<fpage>378</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2019.01.059</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yazdanjue</surname> <given-names>N</given-names></string-name>, <string-name><surname>Fathian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Amiri</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Evolutionary algorithms for k-anonymity in social networks based on clustering approach</article-title>. <source>Comput J</source>. <year>2020</year>;<volume>63</volume>(<issue>7</issue>):<fpage>1039</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1093/comjnl/bxz069</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rajabzadeh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shahsafi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Khoramnejadi</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A graph modification approach for k-anonymity in social networks using the genetic algorithm</article-title>. <source>Soc Netw Anal Min</source>. <year>2020</year>;<volume>10</volume>(<issue>1</issue>):<fpage>38</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s13278-020-00655-6</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Bezdek</surname> <given-names>JC</given-names></string-name></person-group>. <source>Pattern recognition with fuzzy objective function algorithms</source>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Plenum Press</publisher-name>; <year>1981</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shokouhifar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jalali</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Optimized Sugeno fuzzy clustering algorithm for wireless sensor networks</article-title>. <source>Eng Appl Artif Intell</source>. <year>2017</year>;<volume>60</volume>:<fpage>16</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2017.01.007</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mirjalili</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mirjalili</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Lewis</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Grey wolf optimizer</article-title>. <source>Adv Eng Softw</source>. <year>2014</year>;<volume>69</volume>:<fpage>46</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.advengsoft.2013.12.007</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Memarian</surname> <given-names>S</given-names></string-name>, <string-name><surname>Behmanesh-Fard</surname> <given-names>N</given-names></string-name>, <string-name><surname>Aryai</surname> <given-names>P</given-names></string-name>, <string-name><surname>Shokouhifar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mirjalili</surname> <given-names>S</given-names></string-name>, <string-name><surname>del Carmen Romero-Ternero</surname> <given-names>M</given-names></string-name></person-group>. <article-title>TSFIS-GWO: metaheuristic-driven Takagi-Sugeno fuzzy system for adaptive real-time routing in WBANs</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>155</volume>:<fpage>111427</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.111427</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Holland</surname> <given-names>J</given-names></string-name></person-group>. <source>Adaptation in natural and artificial systems</source>. <publisher-loc>Ann Arbor, MI, USA</publisher-loc>: <publisher-name>Universityof Michigan Press</publisher-name>; <year>1975</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>El-kenawy</surname> <given-names>EM</given-names></string-name>, <string-name><surname>Khodadadi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mirjalili</surname> <given-names>S</given-names></string-name>, <string-name><surname>Abdelhamid</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Eid</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Ibrahim</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Greylag goose optimization: nature-inspired optimization algorithm</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>238</volume>:<fpage>122147</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2023.122147</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>