<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">74033</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.074033</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Fuzzy C-Means Clustering-Driven Pooling for Robust and Generalizable Convolutional Neural Networks</article-title>
<alt-title alt-title-type="left-running-head">Fuzzy C-Means Clustering-Driven Pooling for Robust and Generalizable Convolutional Neural Networks</alt-title>
<alt-title alt-title-type="right-running-head">Fuzzy C-Means Clustering-Driven Pooling for Robust and Generalizable Convolutional Neural Networks</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Byeon</surname><given-names>Seunggyu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Lee</surname><given-names>Jung-hun</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kim</surname><given-names>Jong-Deok</given-names></name><xref ref-type="aff" rid="aff-3">3</xref><email>kimjd@pusan.ac.kr</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Engineering, Dong-eui University</institution>, <addr-line>Busan, 47340</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-2"><label>2</label><institution>Division of Electrical and Electronic Engineering, Korea Maritime &#x0026; Ocean University</institution>, <addr-line>Busan, 49112</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Information Convergence Engineering, Pusan National University</institution>, <addr-line>Busan, 46241</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Jong-Deok Kim. Email: <email>kimjd@pusan.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>12</day><month>3</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>2</issue>
<elocation-id>24</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>09</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>11</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_74033.pdf"></self-uri>
<abstract>
<p>This paper introduces a fuzzy C-means-based pooling layer for convolutional neural networks that explicitly models local uncertainty and ambiguity. Conventional pooling operations, such as max and average, apply rigid aggregation and often discard fine-grained boundary information. In contrast, our method computes soft memberships within each receptive field and aggregates cluster-wise responses through membership-weighted pooling, thereby preserving informative structure while reducing dimensionality. Being differentiable, the proposed layer operates as standard two-dimensional pooling. We evaluate our approach across various CNN backbones and open datasets, including CIFAR-10/100, STL-10, LFW, and ImageNette, and further probe small training set restrictions on MNIST and Fashion-MNIST. In these settings, the proposed pooling consistently improves accuracy and weighted F1 over conventional baselines, with particularly strong gains when training data are scarce. Even with less than 1% of the training set, our method maintains reliable performance, indicating improved sample efficiency and robustness to noisy or ambiguous local patterns. Overall, integrating soft memberships into the pooling operator provides a practical and generalizable inductive bias that enhances robustness and generalization in modern CNN pipelines.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Fuzzy logic</kwd>
<kwd>fuzzy c-means clustering</kwd>
<kwd>membership-based pooling</kwd>
<kwd>convolutional neural networks</kwd>
<kwd>downsampling</kwd>
<kwd>feature extraction</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Korea government (MSIT)</funding-source>
<award-id>IITP-2025-RS-2023-00260098</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Gyeongsangnam-do and the Gyeongnam Techno park</funding-source>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Deep learning has brought remarkable progress in many fields by learning to extract and organize useful features from data step by step. Within this trend, convolutional neural networks (CNNs) have become the main driving force behind top performance in tasks such as image classification, object detection, and segmentation [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. They are now actively applied to demanding areas like medical imaging and industrial inspection. In real-world applications, however, what matters is not only high accuracy but also predictions that can be reliable. This requires probability estimates that are well-calibrated and models that degrade gracefully when data are noisy or perturbed [<xref ref-type="bibr" rid="ref-3">3</xref>]. Such properties strongly depend on how the intermediate feature representations are formed [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. In short, these trends highlight the need for architectural improvements that prevent information loss when intermediate representations are spatially subsampled [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>CNN is a neural architecture introduced by LeNet-5, proposed by LeCun et al. in 1998, consisting of convolutional layers, downsampling stages, and a subsequent classification head [<xref ref-type="bibr" rid="ref-1">1</xref>] as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Since then, subsampling has served as a key design axis for reducing the spatial dimensionality of inputs at each stage. With the introduction of ImageNet, AlexNet and VGG established max pooling between blocks as the de facto standard for rapidly shrinking per-stage inputs [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>], while Network in Network popularized global average pooling (GAP) at the head for the final reduction [<xref ref-type="bibr" rid="ref-10">10</xref>]. GoogLeNet combined stride and pooling within multi-scale Inception paths to control per-stage input size, refined in v2/v3 with factorized convolutions and stronger normalization/optimization [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Inception-v4 and Inception-ResNet fused Inception blocks with residual connections, further strengthening learnable downsampling [<xref ref-type="bibr" rid="ref-13">13</xref>]. The All-CNN family showed that stride-2 convolutions alone can replace fixed pooling with little loss [<xref ref-type="bibr" rid="ref-14">14</xref>], and ResNet made such stride-based transitions standard via identity shortcuts [<xref ref-type="bibr" rid="ref-15">15</xref>]. Under efficiency goals, EfficientNet used depthwise stride-2 operators for spatial reduction [<xref ref-type="bibr" rid="ref-16">16</xref>] and squeeze-and-excitation (SE) and spatial attention modules for channel recalibration and feature refinement [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>], leaving only a final GAP. More recently, ConvNeXt and ViT minimize or remove local pooling and perform dimensionality reduction via patch embeddings and hierarchical merging [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. In summary, subsampling is a core axis that decides how much to reduce each layer&#x2019;s input, and its mechanism has shifted from fixed rules to learnable operators.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of convolutional neural network and its major components</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-1.tif"/>
</fig>
<p>Among the core components of CNNs, the subsampling stage reduces the spatial size of convolutional feature maps, lowers computational cost, and helps mitigate overfitting [<xref ref-type="bibr" rid="ref-7">7</xref>]. Although many modern backbones now downsample primarily with strided convolutions, the operation that condenses local evidence remains important. Max and average pooling summarize each neighborhood into a single scalar, which can blur boundary details, suppress weak yet informative signals, and miss cues near ambiguous boundaries [<xref ref-type="bibr" rid="ref-8">8</xref>]. To address these issues, adaptive reductions have been proposed. Representative examples include <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>-norm/generalized-mean (GeM) pooling, which interpolates between average-like and max-like behaviors [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>], and stochastic/mixed pooling, which use randomization or blending to avoid overcommitment [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>]. These methods aim to preserve fine structure, maintain discriminative cues under distribution shift or noise, and make the reduction step more data-aware.</p>
<p>Nonetheless, important limitations remain. Approaches that learn the pooling exponent or mixing weights, such as <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>/GeM, are prone to overfitting to training-distribution statistics; moreover, when <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>p</mml:mi></mml:math></inline-formula> drifts to extreme values, the power/root operations can induce exploding or vanishing gradients, making training numerically unstable [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>]. More fundamentally, many methods still collapse each neighborhood to a single scalar and weight primarily by magnitude or empirical rules, thereby failing to explicitly model local ambiguity and uncertainty (e.g., overlapping boundaries or weak yet informative signals) [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Fuzzy logic provides a principled mathematical framework for quantifying and handling ambiguity and uncertainty [<xref ref-type="bibr" rid="ref-26">26</xref>]. In the view of fuzzy set theory, a sample is not forced into a single category; instead, it is represented by a membership vector whose components lie in [0, 1] and sum to one. Thus, information that conventional pooling would collapse to a single scalar is first decomposed into a compact, multi-aspect code that can capture overlap among feature groups. Among fuzzy clustering methods, fuzzy C-means (FCM) is a widely used soft partitioning scheme: it alternates centroid estimation with membership updates, assigns higher grades to points closer to a centroid, and uses a fuzzifier <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>m</mml:mi></mml:math></inline-formula> to control assignment sharpness [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]. Viewed this way, fuzzy clustering complements the reduction stage in CNNs. Rather than compressing a neighborhood directly along a fixed aggregation axis, one can first obtain a low-dimensional, uncertainty-aware membership code and then summarize with type-specific strengths, allowing the summary to adapt to local mixtures while remaining drop-in compatible with standard backbones. In the design studied here, each location acquires memberships to a small set of latent clusters, and the layer aggregates with cluster-specific pooling exponents.</p>
<p>This paper proposes a fuzzy pooling layer that estimates per-window soft memberships with FCM and aggregates activations using <italic>cluster-specific</italic> <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> exponents <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>. Memberships weight the contributions of each cluster while the per-centroid exponents modulate pooling sharpness, enabling the layer to explicitly capture local ambiguity and uncertainty near boundaries. The module is drop-in compatible with standard CNNs and maintains strong performance under limited training data and unclear inter-class boundaries.</p>
<p><bold>Contributions.</bold>
<list list-type="simple">
<list-item><label>1.</label><p><bold>Uncertainty-aware pooling.</bold> Introduces a FCM&#x2013;driven pooling that encodes per-window soft memberships and aggregates with <italic>cluster-specific</italic> exponents, explicitly modeling local ambiguity that magnitude-based <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>/GeM or stochastic/mixed pooling do not capture.</p></list-item>
<list-item><label>2.</label><p><bold>Drop-in gains across data scarcity.</bold> Demonstrates consistent improvements in accuracy and weighted F1 over conventional pooling across backbones and datasets, with the largest margins under limited supervision and subtle class boundaries.</p></list-item>
</list></p>
<p><bold>Paper organization.</bold> <xref ref-type="sec" rid="s2">Section 2</xref> reviews pooling methods and related work and discusses their limitations. <xref ref-type="sec" rid="s3">Section 3</xref> details the design and operation of the proposed fuzzy clustering&#x2013;based pooling layer. <xref ref-type="sec" rid="s4">Section 4</xref> evaluates performance across diverse setups, and <xref ref-type="sec" rid="s5">Section 5</xref> concludes with future directions.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Background and Related Work</title>
<p>Pooling in CNNs plays a structural role in shaping the feature hierarchy. By summarizing responses within a local window, it reduces spatial resolution, lowers memory footprint and computational burden, and progressively enlarges the effective receptive field. The resulting compact intermediate representations make subsequent layers more tractable and enable stable scaling from early, high-resolution maps to deeper stages [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>]. In modern architectures, this reduction is realized either by explicit max/average pooling or by strided convolutions; in both cases, the intent is to transform dense neighborhood activations into concise, higher-level features that support efficient learning and inference [<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<sec id="s2_1">
<label>2.1</label>
<title>Value-Based Pooling: From Static to Learned Sharpness</title>
<p><bold>Static operators (max/average).</bold> For a window <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:math></inline-formula> with activations <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>,
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mtext>avg</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>As shown in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>, max pooling selects the largest activation in each window, accentuating strong responses and providing limited shift invariance as in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>, often leading to aliasing artifacts [<xref ref-type="bibr" rid="ref-30">30</xref>]. During backpropagation, the gradient is routed only to the argmax entry, yielding sparse updates but increasing sensitivity to outliers. In contrast, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2b</xref>, average pooling spreads gradients uniformly across entries and stabilizes optimization, though it can oversmooth edges and lose boundary-level evidence [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>].</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Traditional static pooling. (<bold>a</bold>) Max pooling keeps only the strongest response; (<bold>b</bold>) Average pooling takes the mean. Both collapse a local patch to a single scalar</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-2.tif"/>
</fig>
<p><bold>Generalized-mean (<italic>L<sub>p</sub></italic>) pooling.</bold> The generalized mean
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:msub><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>p</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">(</mml:mo></mml:mrow></mml:mstyle><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:munder><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>p</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">)</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mi>p</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:math></disp-formula>controls pooling sharpness with a single parameter <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>p</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]. In <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> when <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, it equals average pooling; for <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula>, it approximates max pooling; and for <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, it damps outliers like a geometric mean. As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3a</xref>, the normalized <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>L</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> case averages squared responses and then takes the square root.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Soft pooling variants on a shared <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula> neighborhood (yellow box). Each row squares the inputs and aggregates via: (<bold>a</bold>) normalized <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>L</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>; (<bold>b</bold>) Top-<italic>K</italic> averaging; (<bold>c</bold>) trimmed mean; (<bold>d</bold>) Type-1 fuzzy pooling with membership functions</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-3.tif"/>
</fig>
<p>In practice, power-mean pooling benefits from proper normalization for both accuracy and numerical stability. Prior work on <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>/GeM reports improved performance when the pooled output is normalized (e.g., power/mean normalization or downstream normalization layers) [<xref ref-type="bibr" rid="ref-23">23</xref>]. More broadly, Batch Normalization (BN) is known to stabilize training by smoothing the loss landscape and conditioning gradients [<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>].</p>
<p><bold>Order-statistic pooling.</bold> This family summarizes responses by <italic>rank</italic> rather than magnitude. As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3b</xref>, <italic>Top-K</italic> pooling averages the <italic>K</italic> largest activations (e.g., Top-3 average pooling on a <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula> neighborhood selects the three largest values and averages them). In <xref ref-type="fig" rid="fig-3">Fig. 3c</xref>, a <italic>trimmed mean</italic> discards the <italic>T</italic> largest and <italic>T</italic> smallest values (e.g., <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> on a <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula> window, i.e., T-1 trimmed pooling), then averages the remainder. Quantile pooling instead selects a specific order statistic, such as the median. With small windows, these rank-based selection masks can change discretely when ranks swap, yielding piecewise-smooth behavior [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p><bold>T-max-avg pooling.</bold> This hybrid keeps the maximum if the <italic>Top-K</italic> responses exceed a threshold <italic>T</italic>; otherwise, it falls back to the average of <italic>Top-K</italic> [<xref ref-type="bibr" rid="ref-33">33</xref>]. It balances noise-robustness and peak preservation, though with small windows, it often degenerates to simpler forms.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Membership-Based Pooling</title>
<p>Unlike value-only reductions, membership-based pooling explicitly represents local ambiguity by assigning each activation a degree of belonging in [0, 1], [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
<p><bold>Type-1 (membership/rule&#x2013;driven).</bold> Given activations and fixed/learned memberships, pooling is performed as a weighted average. As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3d</xref>, learned membership functions (with centers at approximately 0.2, 0.5, and 0.8 in this example) assign weights to entries according to their degrees of belonging before aggregation. This operator is simple and differentiable, but often depends on hand-crafted or task-tuned membership functions, which can limit generality [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>Fuzzy logic has also been integrated into CNN pooling mechanisms and clustering frameworks to enhance feature preservation [<xref ref-type="bibr" rid="ref-36">36</xref>]. FP-CNN [<xref ref-type="bibr" rid="ref-37">37</xref>] introduced a fuzzy pooling operator that adaptively combines max and average pooling through a membership function derived from local feature intensities. Specifically, each pooling region computes a fuzzy membership <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula> and aggregates responses as <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>&#x03BC;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>max</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>avg</mml:mtext></mml:mrow></mml:math></inline-formula>, thus balancing sharpness and smoothness while mitigating spatial information loss. Unlike this hybrid formulation, our proposed method generalizes the pooling process through a trainable, membership-driven <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> aggregation that adjusts its exponent per location, offering a more flexible and differentiable fuzzy adaptation mechanism.</p>
<p><bold>Type-2 (uncertain memberships).</bold> Here, memberships themselves are uncertain, modeled as intervals with type-reduction prior to defuzzification [<xref ref-type="bibr" rid="ref-38">38</xref>]. It can improve robustness under distribution shifts, but adds complexity and hyperparameters.</p>
<p><bold>Position of this work.</bold> Type-1 and Type-2 show the importance of uncertainty-aware pooling, but each faces practical challenges. Our approach addresses these by estimating memberships with FCM and aggregating with cluster-specific exponents, combining adaptability with drop-in compatibility.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Fuzzy Clustering&#x2013;Based Soft Pooling Layer</title>
<p><bold>Design rationale.</bold> The analysis in <xref ref-type="sec" rid="s2">Section 2</xref> suggests three desiderata for a modern pooling operator: (i) avoid collapsing each neighborhood to a single scalar purely by magnitude, so that ambiguous boundary evidence is not discarded; (ii) adapt the <italic>sharpness</italic> of aggregation to local patterns without the discrete switches of rank-based rules; and (iii) remain numerically stable and efficient, without incurring the overhead of heavy attention mechanisms [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>] or spectral transforms.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Overview of the Proposed Layer</title>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> summarizes the workflow of the proposed fuzzy clustering-driven soft pooling layer. Given an input feature map, we compute fuzzy memberships <italic>per sample and per location</italic> via FCM, using cluster centers shared across the batch. We then realize two pooling variants: (a) Membership Averaging and (b) Membership Maxing. Both variants ultimately perform a normalized generalized-mean (<inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>) aggregation whose exponent is adapted per spatial location.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Overview of the proposed fuzzy clustering&#x2013;based soft pooling layer. FCM yields a per-location membership map <italic>U</italic> over <italic>K</italic> latent types. Two branches set the pooling exponent: (<bold>a</bold>) Membership Averaging and (<bold>b</bold>) Membership Maxing. The resulting <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> guides a normalized <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> aggregation, followed by BN-style normalization</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-4.tif"/>
</fig>
<p>In practice, the proposed fuzzy pooling layer directly follows each convolutional block, taking its feature map as input without any intermediate transformation. The output retains the same channel dimension and is passed to the next convolution or fully connected layer. Thus, it functions as a drop-in replacement for conventional Max or Average pooling within standard CNN architectures.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Input Structure and Clustering</title>
<p>Assume a feature tensor <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and let <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the <italic>C</italic>-dimensional feature at spatial location <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> in sample <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>b</mml:mi></mml:math></inline-formula>. And also let <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">c</mml:mtext></mml:mrow></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denote trainable cluster centers in <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> (shared across the batch). Define the per-sample, per-location, per-cluster distance as in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">c</mml:mtext></mml:mrow></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> avoids division by zero. With fuzzifier <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>m</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, the memberships are defined following the standard FCM formulation [<xref ref-type="bibr" rid="ref-28">28</xref>] as in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.470em" minsize="2.470em">(</mml:mo></mml:mrow></mml:mstyle><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>&#x2113;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">(</mml:mo></mml:mrow></mml:mstyle><mml:mfrac><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x2113;</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:msup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">)</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mfrac><mml:mn>2</mml:mn><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow></mml:msup><mml:msup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.470em" minsize="2.470em">)</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mrow><mml:mtext>s.t.</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:math></disp-formula>yielding a membership tensor <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">U</mml:mtext></mml:mrow></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The centers <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">c</mml:mtext></mml:mrow></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> are optimized end-to-end via backpropagation.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Computation of the Pooling Exponent</title>
<p>Each cluster <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>k</mml:mi></mml:math></inline-formula> is associated with a learnable exponent <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>. We constrain <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> via a squashed parameterization as in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is the unbounded learnable parameter and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the sigmoid function.</p>
<p>Given the memberships <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> and cluster exponents <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, the per-location exponent <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is determined by one of the following two schemes as in <xref ref-type="disp-formula" rid="eqn-6">Eqs. (6)</xref> and <xref ref-type="disp-formula" rid="eqn-7">(7)</xref>.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">(a) Membership Averaging:</mml:mtext></mml:mrow></mml:mrow><mml:mspace width="1em" /><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">(b) Membership Maxing:</mml:mtext></mml:mrow></mml:mrow><mml:mspace width="1em" /><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>&#x22C6;</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mrow><mml:mtext>where&#xA0;</mml:mtext></mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>&#x22C6;</mml:mo></mml:msup><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mrow><mml:mtext>arg max</mml:mtext></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><bold>Interpretation of cluster exponents.</bold> Implicitly, clusters act as latent pattern types (e.g., edges, textures, flat regions). Learning a distinct <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> per type lets the layer tune aggregation sharpness to the local pattern (e.g., larger <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> near salient edges, smaller <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> in smooth areas) with low overhead.</p>
<p><bold>Commonality vs. difference.</bold> Both rules rely on the same memberships <italic>U</italic> and type-wise exponents <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. Averaging is smooth and differentiable, while Maxing is simpler but introduces a discrete <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:math></inline-formula>; in practice we use Averaging as default and report Maxing as an ablation.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Integration and Output</title>
<p>Pooling is applied <italic>per channel</italic>. Let <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the pooling window centered at <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> with cardinality <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula>. Using the computed exponent <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the output is defined as in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mspace width="thinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">|</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">|</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the input scalar activation at location <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and channel <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>c</mml:mi></mml:math></inline-formula>.</p>
<p>Optionally, a membership-weighted variant replaces the uniform average (<inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula>) with normalized weights <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2265;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> satisfying <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>BN-Inspired Stabilization after the Proposed Fuzzy Pooling</title>
<p>After the proposed fuzzy pooling, a BN-style normalization is appended to stabilize both feature scales and gradient dynamics. In fuzzy clustering-based pooling, the membership distributions <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> evolve spatially and temporally during training, and their interaction with the adaptive exponents <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can amplify response variance within each mini-batch. This phenomenon is analogous to the local-rule dominance observed in fuzzy systems under non-uniform data distributions [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>Applying BN immediately after pooling normalizes channel-wise statistics, effectively damping variance propagation and smoothing gradient flow across iterations. From a theoretical perspective, the adaptive <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> operator embedded in the proposed fuzzy pooling acts as a non-linear scaling mechanism whose sensitivity to input magnitude increases with <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>p</mml:mi></mml:math></inline-formula>. BN counteracts this by rescaling activations toward unit variance, thereby maintaining consistent learning dynamics regardless of the spatially varying sharpness <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. This interpretation aligns with prior findings that BN smooths the optimization landscape and stabilizes gradient propagation in deep networks [<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>].</p>
<p>Empirical evidence of this stabilizing effect is presented in <xref ref-type="app" rid="app-1">Appendix A</xref>. Across all architectures (LeNet-5, AlexNet, and VGG-16), the BN-applied variants consistently achieved 3%&#x2013;8% higher accuracy and weighted F1 scores with reduced variance, demonstrating that BN functions as a <italic>structural stabilizer</italic> within the proposed fuzzy-adaptive pooling rather than merely as a statistical normalization step. All experiments were conducted with a batch size of 32, providing sufficient samples for reliable batch statistics while avoiding the instability often observed in small-batch BN.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Illustrative Example</title>
<p>To clarify the computation process of the proposed fuzzy pooling layer, <xref ref-type="fig" rid="fig-5">Fig. 5</xref> visualizes the step-by-step operation from clustering to adaptive pooling. An input feature map is clustered by FCM into <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> latent types, yielding a membership map <italic>U</italic>. Two paths are illustrated: (a) <italic>Membership Averaging</italic>, where the cluster-specific exponents <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1.5</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> are averaged with memberships to form <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, then used for normalized <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> pooling; (b) <italic>Membership Maxing</italic>, where the argmax membership selects a cluster per location, followed by pooling with the corresponding <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In both cases, BN-style normalization is applied after pooling. For reference, conventional average and max pooling outputs are also shown.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Step-by-step computation of the proposed fuzzy pooling layer. Top: input feature map, FCM memberships, and learned centroids. Bottom: (<bold>a</bold>) Membership Averaging and (<bold>b</bold>) Membership Maxing, each followed by BN-style normalization. For comparison, the outputs of conventional average and max pooling are also shown</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-5.tif"/>
</fig>
<p>To further highlight the difference from conventional pooling, <xref ref-type="fig" rid="fig-6">Fig. 6</xref> presents the visualized results after the 1st and 2nd pooling stages using various pooling methods. While Average and <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>-norm pooling tend to smooth out detailed edges, and Max pooling overemphasizes only the highest activations, the proposed fuzzy pooling variants (Mavg and Mmax) adaptively regulate contrast based on the learned exponents <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. Assuming the feature activations are normalized in <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, regions associated with clusters having <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> are suppressed (weakened), whereas those with <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> are enhanced (highlighted). Consequently, the darkest and brightest clusters become more prominent, while intermediate tones are attenuated, leading to stronger edge delineation and higher local contrast.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Visual comparison of feature maps after the first and second pooling stages across different pooling methods. Average and <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>-norm pooling blur local textures, while Max pooling saturates high responses. In contrast, the proposed fuzzy pooling (Mavg, Mmax) adaptively balances enhancement and suppression through learned <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>, resulting in more distinct edges and sharper feature contrast</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-6.tif"/>
</fig>
<p>This two-step illustration&#x2014;computation flow followed by visual comparison&#x2014;clarifies how the proposed fuzzy pooling mechanism translates its adaptive exponent control into perceivable structural differences. The enhanced local contrast observed here provides a visual rationale for the superior discriminative performance demonstrated in <xref ref-type="sec" rid="s4_3">Section 4.3</xref> and supports the need for BN stabilization discussed in <xref ref-type="sec" rid="s3_5">Section 3.5</xref>. For further visualizations concerning the effect of the cluster count <italic>K</italic> and different integration strategies, please refer to <xref ref-type="app" rid="app-2">Appendix B</xref>.</p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Computational Complexity Analysis</title>
<p>The proposed fuzzy pooling layer integrates three computational components: (1) membership estimation based on feature-to-center distances defined in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, (2) adaptive <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>-norm aggregation guided by the learned exponents, and (3) BN-style normalization for stabilization. Each part contributes differently to the overall computational cost, which we analyze below.</p>
<p>Note that a classical FCM algorithm requires iterative updates with a typical complexity of <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>K</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mi>C</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>I</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> or <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>K</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>C</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>I</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for improved variants [<xref ref-type="bibr" rid="ref-39">39</xref>], where <italic>I</italic> is the number of iterations. By contrast, the proposed pooling layer performs a single-pass computation without iteration.</p>
<p><bold>Membership computation.</bold> Using the distance formulation in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, each spatial location computes its distances to <italic>K</italic> cluster centers and normalizes them according to the membership rule in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>. This operation requires <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> computations per location, and thus <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>C</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for the entire feature map, where <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>. Since cluster centers are optimized by backpropagation rather than iterative FCM updates, no additional loop over iterations is needed, keeping the process single-pass.</p>
<p><bold>Adaptive pooling and normalization.</bold> Once the memberships are obtained, the per-location exponent <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is computed through either the weighted averaging rule in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref> or the hard selection rule in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>. Both require at most <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> operations. Subsequent adaptive <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> pooling in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref> and BN-style normalization are performed channel-wise, each with <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>C</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> complexity.</p>
<p><bold>Overall complexity.</bold> Summing these components, the overall computational cost of the proposed layer is <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>C</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, dominated by the membership computation step. For comparison, standard average or max pooling has <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>C</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> complexity. Therefore, the proposed fuzzy pooling introduces only a small linear factor of <italic>K</italic> to model fuzzy memberships and adaptive aggregation strength, while maintaining computational feasibility for modern CNN architectures.</p>
<p><bold>Empirical inference efficiency.</bold> To complement the asymptotic analysis, we measured the actual inference time per batch (32 samples) across representative CNNs and pooling methods. <xref ref-type="table" rid="table-1">Table 1</xref> summarizes the results averaged over five runs.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Inference time comparison (ms per batch, mean <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std) across representative pooling methods and backbones. &#x2018;FP&#x2019; denotes the Type-1 fuzzy pooling baseline (FP-CNN)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Backbone</th>
<th>Avg</th>
<th>Max</th>
<th><inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula></th>
<th>T-max-avg</th>
<th>FP</th>
<th>Mavg</th>
<th>Mmax</th>
</tr>
</thead>
<tbody>
<tr>
<td>LeNet-5</td>
<td><inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mn>4.11</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.31</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mn>4.05</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.08</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mn>4.53</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.08</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mn>7.70</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.36</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mn>10.48</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.21</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mn>9.72</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.23</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mn>9.89</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.15</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>AlexNet</td>
<td><inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mn>6.90</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.89</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mn>6.69</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.21</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mn>7.69</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.77</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mn>32.70</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>1.42</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mn>18.85</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.36</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mn>17.13</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>1.12</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mn>15.09</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.25</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>VGG-16</td>
<td><inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mn>35.83</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>10.10</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mn>28.51</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.32</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mn>36.43</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>1.15</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mn>504.33</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>2.19</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mn>138.00</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.67</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mn>53.92</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>1.61</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mn>53.18</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>0.82</mml:mn></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The inference latency of the proposed fuzzy pooling (Mavg and Mmax) is generally comparable to that of the Type-1 fuzzy baseline (FP) on smaller networks, while being <italic>substantially faster</italic> on the deeper VGG-16 architecture (e.g., <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mrow><mml:mo>&#x223C;</mml:mo></mml:mrow><mml:mn>2.5</mml:mn><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> speedup over FP). Both variants are also significantly more efficient than the sorting-intensive T-max-avg operator. Although slightly slower than classical average or max pooling due to membership estimation, the proposed methods maintain a favorable trade-off between representational robustness and computational efficiency. Importantly, the gap between Mavg and Mmax remains within statistical variance, suggesting that their inference complexity is nearly identical despite their different exponent selection rules. This confirms the proposed design&#x2019;s suitability for real-time or on-device deployment across CNN architectures of various depths.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Implementation and Evaluation</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Environment and Setup</title>
<p>We evaluate the pooling layers defined in <xref ref-type="sec" rid="s3">Section 3</xref> in <bold>two distinct experimental phases</bold> against <bold>seven</bold> alternatives.</p>
<p><italic>Hardware configuration</italic>.</p>
<p>All experiments were conducted on a workstation equipped with an Intel Core i9-13900F CPU (24 cores, 32 threads), an NVIDIA RTX 4090 GPU (24 GB VRAM), and 64 GB of main memory. TensorFlow (ver. 2.12) was used with the NVIDIA-recommended configuration for the RTX 4090, including CUDA 12.x and cuDNN 8.x libraries. The batch size was fixed at 32 across all experiments to ensure consistent batch-normalization statistics and reproducible inference-time comparisons.</p>
<p><bold>Proposed methods. Mavg:</bold> Membership-Averaging fuzzy pooling; <bold>Mmax:</bold> Membership-Maxing fuzzy pooling. Both variants incorporate the BN-style stabilization described in <xref ref-type="sec" rid="s3">Section 3</xref> (ablation study provided in <xref ref-type="app" rid="app-1">Appendix A</xref>).</p>
<p><bold>Baselines.</bold> To validate effectiveness, we compare against seven pooling strategies: <bold>Max</bold> and <bold>Avg</bold> pooling (standard); <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mi mathvariant="bold-italic">L</mml:mi><mml:mi mathvariant="bold-italic">p</mml:mi></mml:msub></mml:math></inline-formula>: generalized-mean pooling with learnable <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mi>p</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-21">21</xref>]; <bold>T-max-avg:</bold> thresholded Top-<italic>K</italic> hybrid discussed in <xref ref-type="sec" rid="s2_1">Section 2.1</xref> [<xref ref-type="bibr" rid="ref-33">33</xref>]; <bold>Type-1 fixed:</bold> Type-1 fuzzy pooling with a <italic>fixed</italic> membership function (e.g., triangular/Gaussian) [<xref ref-type="bibr" rid="ref-34">34</xref>]; <bold>Type-1 learnable:</bold> Type-1 fuzzy pooling with <italic>trainable</italic> membership parameters; and <bold>FP:</bold> the fuzzy-pooling module from FP-CNN [<xref ref-type="bibr" rid="ref-37">37</xref>], representing a prior convolution-integrated fuzzy pooling framework.</p>
<p><italic>Backbones and insertion points</italic>.</p>
<p>We use three CNN backbones, as illustrated in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>: (a) LeNet-5 [<xref ref-type="bibr" rid="ref-1">1</xref>], (b) AlexNet [<xref ref-type="bibr" rid="ref-2">2</xref>], and (c) VGG-16 [<xref ref-type="bibr" rid="ref-9">9</xref>]. In all models, the native pooling layers (highlighted in green) are replaced <italic>one-for-one</italic> by each candidate pooling operator, preserving the original window size, stride, and padding configuration.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Backbone CNNs and pooling insertion points. (<bold>a</bold>) LeNet-5-style network for <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula> inputs (two conv blocks <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> two pooling stages <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> FC). (<bold>b</bold>) AlexNet-style network for <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula> inputs. (<bold>c</bold>) VGG-16-style network with five pooling stages. Green blocks indicate where native pooling is swapped with one of {Max, Avg, <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>, T-max-avg, Type-1 fixed, Type-1 learnable, Mavg, Mmax}</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-7.tif"/>
</fig>
<p><italic>Why hold-out (HO) instead of</italic> <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi>k</mml:mi></mml:math></inline-formula><italic>-fold CV</italic>.</p>
<p>While <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mi>k</mml:mi></mml:math></inline-formula>-fold cross-validation is standard practice, we observed that several <italic>traditional</italic> pooling baselines occasionally exhibited <italic>numerical instability</italic> or <italic>convergence failure</italic> (e.g., loss divergence or stagnation) under identical training protocols. Since such instability renders the corresponding folds statistically unreliable and practically undeployable, including them introduces unnecessary volatility into the comparison. Therefore, we adopt a repeated <italic>hold-out</italic> evaluation protocol and aggregate results only from <italic>successfully converged</italic> runs to ensure a fair assessment of peak capability (convergence criteria detailed below).</p>
<p><italic>Experiment 1 (Multi-backbone comparison)</italic>.</p>
<p>This phase evaluates general classification performance across diverse domains.
<list list-type="bullet">
<list-item>
<p><bold>Datasets:</bold> CIFAR-10 and CIFAR-100 [<xref ref-type="bibr" rid="ref-40">40</xref>], LFW (subset) [<xref ref-type="bibr" rid="ref-41">41</xref>], STL-10 [<xref ref-type="bibr" rid="ref-42">42</xref>], and ImageNette (a subset of ImageNet [<xref ref-type="bibr" rid="ref-43">43</xref>]), as summarized in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
</list-item>
<list-item>
<p><bold>Split:</bold> A single random hold-out (HO) partition with a ratio of <bold>train:val:test &#x003D; 0.6:0.2:0.2</bold>.</p></list-item>
<list-item>
<p><bold>Backbones:</bold> LeNet-5, AlexNet, and VGG-16.</p></list-item>
</list></p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Experiment 1 datasets. Input sizes show original resolution <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> training resolution per backbone</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Original size</th>
<th>Train Size (LeNet-5/AlexNet/VGG-16)</th>
<th>#Classes</th>
<th>Note</th>
</tr>
</thead>
<tbody>
<tr>
<td>CIFAR-10</td>
<td><inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula></td>
<td>10</td>
<td>Standard split</td>
</tr>
<tr>
<td>CIFAR-100</td>
<td><inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula></td>
<td>100</td>
<td>Standard split</td>
</tr>
<tr>
<td>STL-10</td>
<td><inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mn>96</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>96</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula></td>
<td>10</td>
<td>Standard split</td>
</tr>
<tr>
<td>LFW (subset)</td>
<td><inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mn>62</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>47</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula></td>
<td>41</td>
<td><monospace>min_faces_per_ person</monospace><inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mo>=</mml:mo><mml:mn>25</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>ImageNette</td>
<td><inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula>/<inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula></td>
<td>10</td>
<td>v2-320 subset</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><italic>Experiment 2 (Low-resolution, low-data probe)</italic>.</p>
<p>This phase probes robustness under data scarcity using lightweight models.
<list list-type="bullet">
<list-item>
<p><bold>Backbone:</bold> LeNet-5 only (input resized to <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula>).</p></list-item>
<list-item>
<p><bold>Datasets:</bold> MNIST [<xref ref-type="bibr" rid="ref-1">1</xref>] and Fashion-MNIST [<xref ref-type="bibr" rid="ref-44">44</xref>], as summarized in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
</list-item>
<list-item>
<p><bold>Split:</bold> For each training fraction <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mi>r</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0.001</mml:mn><mml:mo>,</mml:mo><mml:mn>0.002</mml:mn><mml:mo>,</mml:mo><mml:mn>0.005</mml:mn><mml:mo>,</mml:mo><mml:mn>0.01</mml:mn><mml:mo>,</mml:mo><mml:mn>0.02</mml:mn><mml:mo>,</mml:mo><mml:mn>0.05</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, we sample the <bold>training set</bold> at proportion <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>r</mml:mi></mml:math></inline-formula> and split the <italic>remainder</italic> equally into <bold>val:test &#x003D; 0.5:0.5</bold>.</p></list-item>
</list></p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Experiment 2 datasets (LeNet-5 only). For each train fraction <inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mi>r</mml:mi></mml:math></inline-formula>, the remainder is split val:test &#x003D; 0.5:0.5</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Original size</th>
<th>Train size (LeNet-5)</th>
<th>#Classes</th>
</tr>
</thead>
<tbody>
<tr>
<td>MNIST</td>
<td><inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula> (gray)</td>
<td><inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula></td>
<td>10</td>
</tr>
<tr>
<td>Fashion-MNIST</td>
<td><inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula> (gray)</td>
<td><inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula></td>
<td>10</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><italic>Repeated HO with convergence filtering</italic>.</p>
<p>For each (dataset, backbone, pooling) configuration, we draw independent random HO partitions and train until obtaining a fixed number of <italic>converged runs</italic>: 10 runs for Experiment 1 and 5 runs for Experiment 2. Runs that do not converge (e.g., due to divergence or instability) are discarded and retried with a new random seed and split. We report the mean <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> standard deviation calculated over these retained, successful runs.</p>
<p><italic>Preprocessing and training</italic>.</p>
<p>Inputs are min-max normalized to <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, and targets use one-hot encoding with a label smoothing factor of <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:mn>0.25</mml:mn></mml:math></inline-formula>. Training employs the Adam optimizer with a learning rate of <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:mn>4</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> (decaying by <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:mn>0.1</mml:mn></mml:math></inline-formula> on plateaus) for 100 epochs, utilizing categorical cross-entropy loss. Experiment 1 resolutions vary by backbone: LeNet-5 uses <inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> RGB (first conv adapted to 3 channels), while AlexNet and VGG-16 use <inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula>. Experiment 2 uses native <inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula> grayscale for LeNet-5. Unless otherwise noted, the number of clusters <italic>K</italic> in Mavg/Mmax is set equal to the number of classes for the dataset.</p>
<p><italic>Hyperparameter setting for</italic> <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mi>m</mml:mi></mml:math></inline-formula> <italic>and K</italic>.</p>
<p>The fuzzifier was fixed at <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>2.0</mml:mn></mml:math></inline-formula> in all experiments, following common practice in FCM-based models [<xref ref-type="bibr" rid="ref-28">28</xref>], which provides stable and moderately soft memberships for visual feature maps.</p>
<p>To determine a suitable range for the number of clusters <italic>K</italic>, we conducted preliminary experiments under conditions similar to the two main setups (Training ratio <inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:msub><mml:mi>T</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0.6</mml:mn></mml:math></inline-formula> in Experiment 1 and <inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:msub><mml:mi>T</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula> in Experiment 2) on representative datasets. Varying <italic>K</italic> from 2 to 9 revealed that the proposed method is generally robust to <italic>K</italic>, showing no statistically significant performance degradation across this range.</p>
<p>In the main experiments, to ensure optimal model selection, <italic>K</italic> was determined based on the validation loss during the <italic>stabilization phase</italic> of training&#x2014;the intermediate stage where training and validation losses oscillate around equilibrium before overfitting begins. This phase, which follows the initial rapid convergence and precedes the divergence between training and validation losses, best reflects the model&#x2019;s saturated generalization capability. Accordingly, we monitored the validation loss over the last 10 epochs of this phase and selected the <italic>K</italic> yielding the lowest validation loss. This procedure provides a data-driven and reproducible determination of <italic>K</italic>, avoiding bias from transient states.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Results</title>
<p><italic>Experiment 1 (Multi-backbone)</italic>.</p>
<p><xref ref-type="fig" rid="fig-8">Figs. 8</xref>&#x2013;<xref ref-type="fig" rid="fig-13">13</xref> present the accuracy and weighted F1 distributions (10 converged hold-out runs) for <italic>LeNet-5</italic>, <italic>AlexNet</italic>, and <italic>VGG-16</italic>. Across all configurations, <bold>Mavg</bold> and <bold>Mmax</bold> consistently achieve the highest medians (or tie for the top position) with narrower interquartile ranges (IQRs), indicating robust convergence. Notably, these improvements are statistically significant (Wilcoxon signed-rank test, <inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>) against the strongest baselines in most cases. In contrast, the <italic>Type-1 fuzzy</italic> and <italic>FP-CNN</italic> baselines generally yield lower medians and larger variability, highlighting the advantage of the proposed cluster-specific exponent adaptation.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Classification accuracy of pooling methods with LeNet-5 on (<bold>a</bold>) CIFAR-10, (<bold>b</bold>) CIFAR-100, (<bold>c</bold>) STL-10, (<bold>d</bold>) LFW, and (<bold>e</bold>) Imagenette. FP denotes the fuzzy-pooling module in FP-CNN [<xref ref-type="bibr" rid="ref-37">37</xref>]. Horizontal bars with asterisks mark paired Wilcoxon signed-rank test results (<inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mo>&#x2217;</mml:mo><mml:mo>:</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mo>&#x2217;</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mo>:</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mo>&#x2217;</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mo>:</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>) comparing the proposed fuzzy pooling (Mavg/Mmax) with the strongest non-membership baseline in each dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-8.tif"/>
</fig><fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Weighted F1 scores of pooling methods with LeNet-5 on (<bold>a</bold>) CIFAR-10, (<bold>b</bold>) CIFAR-100, (<bold>c</bold>) STL-10, (<bold>d</bold>) LFW, and (<bold>e</bold>) Imagenette. Significance notation follows <xref ref-type="fig" rid="fig-8">Fig. 8</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-9a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-9b.tif"/>
</fig><fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Classification accuracy of pooling methods with AlexNet on (<bold>a</bold>) STL-10, (<bold>b</bold>) LFW, and (<bold>c</bold>) Imagenette. FP denotes FP-CNN. Asterisks indicate significance under paired Wilcoxon tests (&#x2217;: <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>, &#x2217;&#x2217;: <inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>, &#x2217;&#x2217;&#x2217;: <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-10.tif"/>
</fig><fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Weighted F1 scores of pooling methods with AlexNet on (<bold>a</bold>) STL-10, (<bold>b</bold>) LFW, and (<bold>c</bold>) Imagenette. Significance notation follows <xref ref-type="fig" rid="fig-10">Fig. 10</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-11.tif"/>
</fig><fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Classification accuracy of pooling methods with VGG-16 on (<bold>a</bold>) STL-10, (<bold>b</bold>) LFW, and (<bold>c</bold>) Imagenette. FP denotes FP-CNN. Significance bars follow the Wilcoxon test conventions in <xref ref-type="fig" rid="fig-8">Fig. 8</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-12.tif"/>
</fig><fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Weighted F1 scores of pooling methods with VGG-16 on (<bold>a</bold>) STL-10, (<bold>b</bold>) LFW, and (<bold>c</bold>) Imagenette. Significance notation follows <xref ref-type="fig" rid="fig-12">Fig. 12</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-13.tif"/>
</fig>
<p>As detailed in <xref ref-type="fig" rid="fig-8">Figs. 8</xref> and <xref ref-type="fig" rid="fig-9">9</xref>, for LeNet-5, <bold>Mavg</bold> consistently outperforms Max, Avg, <inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>, and T-max-avg on CIFAR-10 and CIFAR-100, with the gap particularly pronounced on CIFAR-100 given the higher class diversity. On STL-10, although the overall accuracy baseline is lower, both <bold>Mavg</bold> and <bold>Mmax</bold> retain the lead, while learnable Type-1 fuzzy and FP pooling show wider variance and unstable behavior. On LFW, <bold>Mavg</bold> attains the highest medians, though <inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> and T-max-avg approach closely&#x2014;suggesting that when class margins are subtle or inputs are structurally aligned, simpler pooling schemes can occasionally generalize competitively. On Imagenette, <bold>Mavg</bold> again dominates, achieving both the highest median and the tightest spread in accuracy and F1.</p>

<p>Referring to <xref ref-type="fig" rid="fig-10">Figs. 10</xref> and <xref ref-type="fig" rid="fig-11">11</xref>, for AlexNet, the superiority of <bold>Mavg</bold> and <bold>Mmax</bold> is also evident. On STL-10, <bold>Mmax</bold> performs best, closely followed by <bold>Mavg</bold>, both exhibiting small variance. On LFW, <inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> or T-max-avg can sometimes match or slightly exceed the proposed methods, implying that under limited data and subtle inter-class separability, more complex pooling does not always guarantee gains over simpler adaptive baselines. In contrast, on Imagenette, both <bold>Mavg</bold> and <bold>Mmax</bold> clearly outperform the baselines, with significantly smaller variance across repeated runs.</p>
<p>As shown in <xref ref-type="fig" rid="fig-12">Figs. 12</xref> and <xref ref-type="fig" rid="fig-13">13</xref>, for VGG-16, the deeper backbone raises overall accuracy and F1 relative to LeNet-5 and AlexNet, yet <bold>Mavg</bold> consistently maintains a performance advantage&#x2014;most visibly on LFW and Imagenette. While <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> and T-max-avg appear competitive in certain configurations, they exhibit wider variability and more outliers, whereas <bold>Mavg</bold> and <bold>Mmax</bold> offer stable and robust convergence.</p>
<p>Comparing accuracy and weighted F1 across all backbones, the relative rankings of pooling methods remain nearly identical. Weighted F1, which better accounts for class imbalance, further highlights the superiority of the proposed methods, especially on CIFAR-100 and Imagenette where the number of classes is large and intra-class variance is substantial. This indicates that membership-based pooling aggregates feature responses more effectively than either extremal selection (Max) or uniform averaging (Avg).</p>
<p><bold>Statistical significance.</bold> To confirm that the observed performance gains are not due to random variation, paired two-sided Wilcoxon signed-rank tests were conducted over 10 independent runs per dataset and backbone. Horizontal bars and asterisks in <xref ref-type="fig" rid="fig-8">Figs. 8</xref>&#x2013;<xref ref-type="fig" rid="fig-13">13</xref> denote significance levels (<inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mo>&#x2217;</mml:mo><mml:mo>:</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mo>&#x2217;</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mo>:</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mo>&#x2217;</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mo>&#x2217;</mml:mo><mml:mo>:</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>). In nearly all cases, the proposed <bold>Mavg</bold> and <bold>Mmax</bold> achieve statistically significant improvements over both conventional (Avg, Max, <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>, T-max-avg) and fuzzy-type (Type-1, FP) pooling baselines, demonstrating that the observed advantages arise from the membership-based mechanism rather than random initialization.</p>

<p>In summary, Experiment 1 demonstrates that <bold>Mavg</bold> is the most consistently strong pooling method, with <bold>Mmax</bold> providing comparable or complementary benefits. Both converge reliably across diverse datasets and backbones, as evidenced by reduced variance and statistically significant improvements in most cases. While <inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula> or T-max-avg may serve as lightweight alternatives in constrained settings (e.g., LFW with AlexNet), membership-based pooling offers more robust and reproducible performance overall. Practically, <bold>Mavg</bold> is recommended as a default choice, whereas <bold>Mmax</bold> may be preferred when preserving high-frequency or edge-dominant features is critical, as qualitatively visualized in <xref ref-type="app" rid="app-2">Appendix B</xref>.</p>
<p><italic>Experiment 2 (Performance under limited supervision)</italic>.</p>
<p>We investigate the impact of training-data scarcity by varying the fraction of training samples <inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mi>r</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0.001</mml:mn><mml:mo>,</mml:mo><mml:mn>0.002</mml:mn><mml:mo>,</mml:mo><mml:mn>0.005</mml:mn><mml:mo>,</mml:mo><mml:mn>0.01</mml:mn><mml:mo>,</mml:mo><mml:mn>0.02</mml:mn><mml:mo>,</mml:mo><mml:mn>0.05</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.5</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> for LeNet-5 on MNIST and Fashion-MNIST. <xref ref-type="fig" rid="fig-14">Figs. 14</xref> and <xref ref-type="fig" rid="fig-15">15</xref> summarize the trends in classification accuracy and weighted F1, respectively (reported as mean <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std over 5 converged HO runs).</p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>Classification accuracy across training fractions for MNIST and Fashion-MNIST using LeNet-5 (Experiment 2). Shaded bands denote <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mo>&#x00B1;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> standard deviation over converged runs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-14.tif"/>
</fig><fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Weighted F1 across training fractions for MNIST and Fashion-MNIST using LeNet-5 (Experiment 2). Shaded bands denote <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mo>&#x00B1;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> standard deviation over converged runs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-15.tif"/>
</fig>
<p>Overall, <bold>membership-based pooling (Mavg/Mmax)</bold> consistently outperforms conventional pooling operators on both datasets. On MNIST, the advantage is evident even at extremely small training ratios (<inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>), where <bold>Mavg</bold> and <bold>Mmax</bold> sustain an accuracy <inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow><mml:mn>0.65</mml:mn></mml:math></inline-formula> and weighted F1 <inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow><mml:mn>0.66</mml:mn></mml:math></inline-formula>, while Avg and Max fall below <inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:mn>0.40</mml:mn></mml:math></inline-formula> on both metrics&#x2014;indicating that the proposed pooling extracts salient features even under severe supervision constraints.</p>
<p>As <inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:mi>r</mml:mi></mml:math></inline-formula> increases, all methods improve, but the <italic>relative ordering remains stable</italic>. <bold>Mavg</bold> and <bold>Mmax</bold> exceed <inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:mn>0.95</mml:mn></mml:math></inline-formula> accuracy and <inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:mn>0.93</mml:mn></mml:math></inline-formula> weighted F1 by <inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>0.02</mml:mn></mml:math></inline-formula>, whereas Avg and Max lag by a significant margin (<inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow><mml:mn>10</mml:mn></mml:math></inline-formula> percentage points). Type-1 fuzzy baselines (fixed or learnable) improve over Avg/Max but exhibit high variance, especially at small <inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:mi>r</mml:mi></mml:math></inline-formula>, signaling less stable convergence.</p>
<p>On Fashion-MNIST&#x2014;which is more challenging due to higher intra-class variability&#x2014;the gaps are even more pronounced at small fractions (<inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mi>r</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>0.005</mml:mn></mml:math></inline-formula>): conventional pooling shows wider standard deviation bands and degraded F1, whereas <bold>Mavg</bold> and <bold>Mmax</bold> improve steadily with narrower uncertainty bands. As the training set grows (<inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:mi>r</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>), all methods approach saturation, yet the proposed methods retain a notable edge, surpassing 0.90 weighted F1 at <inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>.</p>
<p>Taken together, Experiment 2 validates the robustness of membership-based pooling under limited supervision, a common real-world scenario where large annotated datasets are unavailable. By aggregating responses via soft memberships&#x2014;avoiding the pitfalls of both extremal selection (Max) and uniform averaging (Avg)&#x2014;the proposed pooling family offers a stronger inductive bias that translates into superior generalization, particularly under data scarcity.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Analysis and Discussion</title>
<p><italic>Key observations</italic>.
<list list-type="simple">
<list-item><label>(1)</label><p>Across LeNet-5, AlexNet, and VGG-16, <bold>Mavg</bold>/<bold>Mmax</bold> typically achieve higher medians and tighter IQRs than Max/Avg/<inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>/T-max-avg and both Type-1 fuzzy variants on CIFAR-100 and ImageNette, indicating simultaneous gains in accuracy and stability&#x2014;critical when reproducibility matters.</p></list-item>
<list-item><label>(2)</label><p>On LFW, improvements are present but smaller. With AlexNet, conventional schemes may occasionally tie or slightly outperform the proposed layers. This suggests that when early features are already strongly structured (e.g., by large receptive fields or strong low-level inductive biases) and decision boundaries are subtle, pooling contributes less to overall discriminability.</p></list-item>
<list-item><label>(3)</label><p><bold>Type-1 fuzzy pooling</bold> (fixed and learnable) consistently shows lower central tendency and higher variance. This aligns with the difficulty of specifying or stably learning crisp membership functions under noisy local statistics. By contrast, our FCM-style soft memberships adapt smoothly within each window, reducing sensitivity to local outliers and improving robustness&#x2014;consistent with Experiment 2 results under limited supervision.</p></list-item>
<list-item><label>(4)</label><p>Weighted F1 mirrors accuracy across all settings, indicating that gains are not artifacts of class-frequency imbalance but reflect genuine improvements in balanced classification.</p></list-item>
</list></p>
<p><italic>Overall implications</italic>.</p>
<p><bold>Membership-based pooling provides a consistent inductive bias across architectures and datasets</bold>, particularly when data are noisy, imbalanced, or scarce. While its relative advantage can diminish under strongly regularized architectures or inherently separable feature spaces (e.g., LFW with AlexNet), the method remains robust and avoids the instability observed in Type-1 fuzzy baselines. These properties make it a practical drop-in replacement for standard pooling in real-world scenarios where dataset conditions are rarely ideal.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>This work proposed two fuzzy C-means (FCM)-based pooling layers, <bold>Mavg</bold> (Membership-Averaging) and <bold>Mmax</bold> (Membership-Maxing), that bring <italic>soft, data-driven</italic> aggregation into convolutional neural networks. By leveraging fuzzy memberships computed per pooling region and converting them into a <italic>location-adaptive</italic> pooling exponent with BN-style stabilization (cf. <xref ref-type="sec" rid="s3">Section 3</xref>), the layers preserve boundary ambiguity and reduce information loss typical of static operators (Max/Avg) while remaining drop-in compatible with standard CNNs.</p>
<p><italic>Empirical findings</italic>.</p>
<p>Across <bold>Experiment 1</bold> (three backbones: LeNet-5, AlexNet, VGG-16; five datasets: CIFAR-10/100, STL-10, LFW, ImageNette), membership-based pooling attained higher median accuracy and weighted F1 with tighter variability than Max/Avg/<inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:msub><mml:mi>L</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>/T-max-avg and Type-1 fuzzy baselines, with the largest gains on CIFAR-100 and ImageNette where class diversity and ambiguity are pronounced. In <bold>Experiment 2</bold> (severe data scarcity on MNIST/Fashion-MNIST, down to <inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>), Mavg/Mmax retained meaningful performance where conventional pooling degraded sharply&#x2014;evidence that soft, membership-guided aggregation is robust in low-data regimes.</p>
<p><italic>Limitations and scope</italic>.</p>
<p>Although the proposed fuzzy pooling layers are lightweight and drop-in compatible with existing CNNs, several limitations and considerations remain.</p>
<p>First, the computation of fuzzy memberships introduces a slight linear overhead proportional to the number of clusters <italic>K</italic>. This cost is small compared to convolutional operations and involves no iterative optimization. Importantly, unlike rule-based fuzzy systems or conditional pooling strategies, the proposed layer contains no branching operations, which preserves GPU pipelining efficiency and allows highly parallel execution across pooling regions. Consequently, inference speed remains close to that of conventional pooling, as verified in our complexity analysis (<xref ref-type="sec" rid="s3_7">Section 3.7</xref>).</p>
<p>Second, the method&#x2019;s performance shows moderate sensitivity to hyperparameters such as the number of clusters <italic>K</italic> and the fuzzifier <inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:mi>m</mml:mi></mml:math></inline-formula>. Although stable results were obtained for <inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:mi>K</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>9</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>2.0</mml:mn></mml:math></inline-formula> across datasets, further work could explore adaptive or data-driven tuning schemes to improve robustness across architectures and domains.</p>
<p>Third, the BN-style normalization effectively stabilizes training but assumes sufficiently large batch sizes for reliable statistics. When batch size is limited or streaming inference is required, alternatives such as Group Normalization or Instance Normalization may provide better stability while retaining the same integration principle.</p>
<p>Finally, our evaluations were based on publicly available image benchmarks to ensure comparability with prior pooling studies. While these datasets provide valuable diversity, future validation on real-world or domain-specific data&#x2014;such as medical, environmental, or defense applications&#x2014;would further demonstrate the method&#x2019;s generalization and reliability under practical conditions.</p>
<p><italic>Practical implications</italic>.</p>
<p><bold>Mavg</bold> is a strong default due to its smoothness and stable convergence; <bold>Mmax</bold> can be preferable when preserving high-frequency or edge-dominant responses is critical. Using <italic>K</italic> close to the number of classes worked reliably across settings, and BN-style post-normalization consistently improved training stability and reproducibility. That said, benefits diminish when early features are already highly separable (e.g., AlexNet on LFW), suggesting that architecture capacity and dataset characteristics should inform the choice of variant and <italic>K</italic>.</p>
<p><italic>Future work</italic>.
<list list-type="bullet">
<list-item>
<p><italic>Hyperparameters and rules</italic>. Systematic study of <italic>K</italic>, fuzzifier <inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:mi>m</mml:mi></mml:math></inline-formula>, and exponent-composition rules (e.g., temperature-controlled averaging, entropy-aware mixing) to balance accuracy and efficiency.</p></list-item>
<list-item>
<p><italic>Learning strategies</italic>. Regularization/scheduling for the membership map, bilevel objectives for centroids vs. features, and calibration-aware training to improve reliability under shift.</p></list-item>
<list-item>
<p><italic>Architectural/generalization breadth</italic>. Extending to modern backbones (ResNets, ConvNeXts, Transformers) and tasks beyond classification (detection/segmentation), including dense prediction where spatial ambiguity is critical.</p></list-item>
<list-item>
<p><italic>Coupled optimization</italic>. Joint refinement of memberships and features (e.g., alternating updates, meta-learning of <inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:mi>m</mml:mi></mml:math></inline-formula>), and exploring probabilistic or mixture-of-experts views that unify clustering and pooling.</p></list-item>
</list></p>
<p><italic>Summary</italic>.</p>
<p>Embedding fuzzy memberships into pooling offers a principled and practical path to more robust, generalizable CNNs. By consistently improving accuracy and balanced metrics across architectures, datasets, and supervision levels&#x2014;and by retaining a simple drop-in form&#x2014;Mavg/Mmax illustrate the value of importing fuzzy-set principles into core deep learning operators.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to express their gratitude to all collaborators who provided constructive feedback during this study.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was supported by the Institute of Information &#x0026; Communications Technology Planning &#x0026; Evaluation (IITP)&#x2013;ITRC (Information Technology Research Center) grant funded by the Korea government (MSIT) (IITP-2025-RS-2023-00260098, 50%), and the Aerospace and ICT Localization &#x0026; Commercialization Technology Development Project funded by Gyeongsangnam-do and the Gyeongnam Techno park (50%).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization, Seunggyu Byeon, Jung-hun Lee and Jong-Deok Kim; methodology, Seunggyu Byeon; software, Seunggyu Byeon; validation, Seunggyu Byeon and Jung-hun Lee; formal analysis, Seunggyu Byeon; investigation, Seunggyu Byeon and Jung-hun Lee; writing&#x2014;original draft preparation, Seunggyu Byeon; writing&#x2014;review and editing, Seunggyu Byeon and Jong-Deok Kim; visualization, Seunggyu Byeon and Jung-hun Lee; supervision, Jong-Deok Kim; project administration, Jong-Deok Kim; funding acquisition, Jong-Deok Kim. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>All datasets used in this study are publicly available: CIFAR-10/100, STL-10, LFW (subset), ImageNette, MNIST, and Fashion-MNIST (see <xref ref-type="table" rid="table-2">Tables 2</xref> and <xref ref-type="table" rid="table-3">3</xref>). The implementation of the proposed Mavg/Mmax pooling layers (including training and evaluation scripts) is available at: <ext-link ext-link-type="uri" xlink:href="https://colab.research.google.com/drive/1u8S6Nyp8Ojciy28bMGnZXIoYuehoADCJ?usp=sharing">https://colab.research.google.com/drive/1u8S6Nyp8Ojciy28bMGnZXIoYuehoADCJ?usp=sharing</ext-link>. If any access issues occur, please contact the corresponding author.</p>

</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations</title>
<def-list>
<def-item>
<term>Avg</term>
<def>
<p>Average pooling</p>
</def>
</def-item>
<def-item>
<term>BN</term>
<def>
<p>Batch normalization</p>
</def>
</def-item>
<def-item>
<term>CIFAR</term>
<def>
<p>Canadian Institute for Advanced Research image datasets (CIFAR-10/100)</p>
</def>
</def-item>
<def-item>
<term>CNN</term>
<def>
<p>Convolutional neural network</p>
</def>
</def-item>
<def-item>
<term>CV</term>
<def>
<p>Cross-validation</p>
</def>
</def-item>
<def-item>
<term>FC</term>
<def>
<p>Fully connected (layer)</p>
</def>
</def-item>
<def-item>
<term>FCM</term>
<def>
<p>Fuzzy C-means clustering</p>
</def>
</def-item>
<def-item>
<term>FP</term>
<def>
<p>Fuzzy pooling (specifically the module in FP-CNN)</p>
</def>
</def-item>
<def-item>
<term>F1</term>
<def>
<p>F1-score (harmonic mean of precision and recall); &#x201C;weighted F1&#x201D; is class-frequency weighted</p>
</def>
</def-item>
<def-item>
<term>HO</term>
<def>
<p>Hold-out (train/validation/test split)</p>
</def>
</def-item>
<def-item>
<term>IQR</term>
<def>
<p>Interquartile range</p>
</def>
</def-item>
<def-item>
<term>Lp</term>
<def>
<p>Generalized-mean pooling with exponent p</p>
</def>
</def-item>
<def-item>
<term>LFW</term>
<def>
<p>Labeled Faces in theWild</p>
</def>
</def-item>
<def-item>
<term>Mavg</term>
<def>
<p>Membership-averaging fuzzy pooling (proposed)</p>
</def>
</def-item>
<def-item>
<term>Mmax</term>
<def>
<p>Membership-maxing fuzzy pooling (proposed)</p>
</def>
</def-item>
<def-item>
<term>MNIST</term>
<def>
<p>Modified National Institute of Standards and Technology dataset</p>
</def>
</def-item>
<def-item>
<term>RGB</term>
<def>
<p>Red&#x2013;Green&#x2013;Blue color channels</p>
</def>
</def-item>
<def-item>
<term>STL-10</term>
<def>
<p>STL-10 image dataset (10 classes, 96 &#x00D7; 96)</p>
</def>
</def-item>
<def-item>
<term>T-max-avg</term>
<def>
<p>Thresholded Top-K max&#x2013;average hybrid pooling</p>
</def>
</def-item>
<def-item>
<term>VGG</term>
<def>
<p>Visual Geometry Group (e.g., VGG-16)</p>
</def>
</def-item>
</def-list>
</glossary>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A Ablation Results on BN-Style Normalization</title>
<p>This appendix summarizes the detailed ablation results for BN-style normalization applied after the proposed fuzzy pooling layer. The normalization follows the standard batch normalization formulation described in [<xref ref-type="bibr" rid="ref-31">31</xref>]. The quantitative comparisons across datasets and backbones are presented in <xref ref-type="table" rid="table-A1">Table A1</xref>.</p>
<table-wrap id="table-A1">
<label>Table A1</label>
<caption>
<title>Ablation results of BN-style normalization across datasets and backbones. Each row shows mean <inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std over 10 runs. &#x2018;O&#x2019;: BN applied; &#x2018;X&#x2019;: BN omitted</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Backbone</th>
<th>Pooling</th>
<th>ACC (O)</th>
<th>ACC (X)</th>
<th>F1 (O)</th>
<th>F1 (X)</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2" align="center">CIFAR-10</td>
<td>LeNet-5</td>
<td>Mavg</td>
<td>89.6 <inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>83.4 <inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
<td>89.2 <inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>82.9 <inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.0</td>
</tr>
<tr>
<td>LeNet-5</td>
<td>Mmax</td>
<td>88.7 <inline-formula id="ieqn-212"><mml:math id="mml-ieqn-212"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>82.5 <inline-formula id="ieqn-213"><mml:math id="mml-ieqn-213"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.1</td>
<td>88.3 <inline-formula id="ieqn-214"><mml:math id="mml-ieqn-214"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>82.1 <inline-formula id="ieqn-215"><mml:math id="mml-ieqn-215"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.2</td>
</tr>
<tr>
<td rowspan="2" align="center">CIFAR-100</td>
<td>LeNet-5</td>
<td>Mavg</td>
<td>65.7 <inline-formula id="ieqn-216"><mml:math id="mml-ieqn-216"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
<td>58.9 <inline-formula id="ieqn-217"><mml:math id="mml-ieqn-217"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.3</td>
<td>65.2 <inline-formula id="ieqn-218"><mml:math id="mml-ieqn-218"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
<td>58.1 <inline-formula id="ieqn-219"><mml:math id="mml-ieqn-219"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.3</td>
</tr>
<tr>
<td>LeNet-5</td>
<td>Mmax</td>
<td>64.5 <inline-formula id="ieqn-220"><mml:math id="mml-ieqn-220"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
<td>58.2 <inline-formula id="ieqn-221"><mml:math id="mml-ieqn-221"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.2</td>
<td>64.0 <inline-formula id="ieqn-222"><mml:math id="mml-ieqn-222"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
<td>57.7 <inline-formula id="ieqn-223"><mml:math id="mml-ieqn-223"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.3</td>
</tr>
<tr>
<td rowspan="6">STL-10</td>
<td rowspan="2">LeNet-5</td>
<td>Mavg</td>
<td>75.9 <inline-formula id="ieqn-224"><mml:math id="mml-ieqn-224"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>70.2 <inline-formula id="ieqn-225"><mml:math id="mml-ieqn-225"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
<td>75.4 <inline-formula id="ieqn-226"><mml:math id="mml-ieqn-226"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>69.7 <inline-formula id="ieqn-227"><mml:math id="mml-ieqn-227"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
</tr>
<tr>
<td>Mmax</td>
<td>74.5 <inline-formula id="ieqn-228"><mml:math id="mml-ieqn-228"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>69.1 <inline-formula id="ieqn-229"><mml:math id="mml-ieqn-229"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
<td>74.1 <inline-formula id="ieqn-230"><mml:math id="mml-ieqn-230"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>68.7 <inline-formula id="ieqn-231"><mml:math id="mml-ieqn-231"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
</tr>
<tr>
<td rowspan="2">AlexNet</td>
<td>Mavg</td>
<td>78.4 <inline-formula id="ieqn-232"><mml:math id="mml-ieqn-232"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>73.1 <inline-formula id="ieqn-233"><mml:math id="mml-ieqn-233"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>78.0 <inline-formula id="ieqn-234"><mml:math id="mml-ieqn-234"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>72.6 <inline-formula id="ieqn-235"><mml:math id="mml-ieqn-235"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
</tr>
<tr>
<td>Mmax</td>
<td>77.5 <inline-formula id="ieqn-236"><mml:math id="mml-ieqn-236"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>72.2 <inline-formula id="ieqn-237"><mml:math id="mml-ieqn-237"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
<td>77.0 <inline-formula id="ieqn-238"><mml:math id="mml-ieqn-238"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>71.7 <inline-formula id="ieqn-239"><mml:math id="mml-ieqn-239"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
</tr>
<tr>
<td rowspan="2">VGG-16</td>
<td>Mavg</td>
<td>80.1 <inline-formula id="ieqn-240"><mml:math id="mml-ieqn-240"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>76.9 <inline-formula id="ieqn-241"><mml:math id="mml-ieqn-241"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>79.7 <inline-formula id="ieqn-242"><mml:math id="mml-ieqn-242"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>76.4 <inline-formula id="ieqn-243"><mml:math id="mml-ieqn-243"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
</tr>
<tr>
<td>Mmax</td>
<td>79.0 <inline-formula id="ieqn-244"><mml:math id="mml-ieqn-244"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>75.8 <inline-formula id="ieqn-245"><mml:math id="mml-ieqn-245"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>78.5 <inline-formula id="ieqn-246"><mml:math id="mml-ieqn-246"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>75.3 <inline-formula id="ieqn-247"><mml:math id="mml-ieqn-247"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
</tr>
<tr>
<td rowspan="6">LFW</td>
<td rowspan="2">LeNet-5</td>
<td>Mavg</td>
<td>82.5 <inline-formula id="ieqn-248"><mml:math id="mml-ieqn-248"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>79.1 <inline-formula id="ieqn-249"><mml:math id="mml-ieqn-249"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.0</td>
<td>82.1 <inline-formula id="ieqn-250"><mml:math id="mml-ieqn-250"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>78.8 <inline-formula id="ieqn-251"><mml:math id="mml-ieqn-251"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.0</td>
</tr>
<tr>
<td>Mmax</td>
<td>81.8 <inline-formula id="ieqn-252"><mml:math id="mml-ieqn-252"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
<td>78.6 <inline-formula id="ieqn-253"><mml:math id="mml-ieqn-253"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.1</td>
<td>81.3 <inline-formula id="ieqn-254"><mml:math id="mml-ieqn-254"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.8</td>
<td>78.2 <inline-formula id="ieqn-255"><mml:math id="mml-ieqn-255"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.1</td>
</tr>
<tr>
<td rowspan="2">AlexNet</td>
<td>Mavg</td>
<td>84.2 <inline-formula id="ieqn-256"><mml:math id="mml-ieqn-256"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>80.8 <inline-formula id="ieqn-257"><mml:math id="mml-ieqn-257"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
<td>83.8 <inline-formula id="ieqn-258"><mml:math id="mml-ieqn-258"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>80.3 <inline-formula id="ieqn-259"><mml:math id="mml-ieqn-259"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
</tr>
<tr>
<td>Mmax</td>
<td>83.7 <inline-formula id="ieqn-260"><mml:math id="mml-ieqn-260"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>80.0 <inline-formula id="ieqn-261"><mml:math id="mml-ieqn-261"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
<td>83.3 <inline-formula id="ieqn-262"><mml:math id="mml-ieqn-262"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>79.5 <inline-formula id="ieqn-263"><mml:math id="mml-ieqn-263"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.9</td>
</tr>
<tr>
<td rowspan="2">VGG-16</td>
<td>Mavg</td>
<td>86.5 <inline-formula id="ieqn-264"><mml:math id="mml-ieqn-264"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>83.7 <inline-formula id="ieqn-265"><mml:math id="mml-ieqn-265"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>86.1 <inline-formula id="ieqn-266"><mml:math id="mml-ieqn-266"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>83.3 <inline-formula id="ieqn-267"><mml:math id="mml-ieqn-267"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
</tr>
<tr>
<td>Mmax</td>
<td>85.9 <inline-formula id="ieqn-268"><mml:math id="mml-ieqn-268"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>83.2 <inline-formula id="ieqn-269"><mml:math id="mml-ieqn-269"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
<td>85.4 <inline-formula id="ieqn-270"><mml:math id="mml-ieqn-270"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>82.8 <inline-formula id="ieqn-271"><mml:math id="mml-ieqn-271"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.7</td>
</tr>
<tr>
<td rowspan="6">ImageNette</td>
<td rowspan="2">LeNet-5</td>
<td>Mavg</td>
<td>90.5 <inline-formula id="ieqn-272"><mml:math id="mml-ieqn-272"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>86.9 <inline-formula id="ieqn-273"><mml:math id="mml-ieqn-273"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>90.1 <inline-formula id="ieqn-274"><mml:math id="mml-ieqn-274"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>86.5 <inline-formula id="ieqn-275"><mml:math id="mml-ieqn-275"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
</tr>
<tr>
<td>Mmax</td>
<td>89.8 <inline-formula id="ieqn-276"><mml:math id="mml-ieqn-276"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>86.2 <inline-formula id="ieqn-277"><mml:math id="mml-ieqn-277"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>89.4 <inline-formula id="ieqn-278"><mml:math id="mml-ieqn-278"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>85.9 <inline-formula id="ieqn-279"><mml:math id="mml-ieqn-279"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
</tr>
<tr>
<td rowspan="2">AlexNet</td>
<td>Mavg</td>
<td>92.1 <inline-formula id="ieqn-280"><mml:math id="mml-ieqn-280"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>88.7 <inline-formula id="ieqn-281"><mml:math id="mml-ieqn-281"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>91.7 <inline-formula id="ieqn-282"><mml:math id="mml-ieqn-282"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>88.3 <inline-formula id="ieqn-283"><mml:math id="mml-ieqn-283"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
</tr>
<tr>
<td>Mmax</td>
<td>91.5 <inline-formula id="ieqn-284"><mml:math id="mml-ieqn-284"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>88.2 <inline-formula id="ieqn-285"><mml:math id="mml-ieqn-285"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>91.1 <inline-formula id="ieqn-286"><mml:math id="mml-ieqn-286"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.4</td>
<td>87.9 <inline-formula id="ieqn-287"><mml:math id="mml-ieqn-287"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
</tr>
<tr>
<td rowspan="2">VGG-16</td>
<td>Mavg</td>
<td>94.0 <inline-formula id="ieqn-288"><mml:math id="mml-ieqn-288"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>90.9 <inline-formula id="ieqn-289"><mml:math id="mml-ieqn-289"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>93.7 <inline-formula id="ieqn-290"><mml:math id="mml-ieqn-290"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>90.6 <inline-formula id="ieqn-291"><mml:math id="mml-ieqn-291"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
</tr>
<tr>
<td>Mmax</td>
<td>93.5 <inline-formula id="ieqn-292"><mml:math id="mml-ieqn-292"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>90.3 <inline-formula id="ieqn-293"><mml:math id="mml-ieqn-293"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>93.2 <inline-formula id="ieqn-294"><mml:math id="mml-ieqn-294"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.3</td>
<td>90.0 <inline-formula id="ieqn-295"><mml:math id="mml-ieqn-295"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>CIFAR-10 and CIFAR-100 were evaluated using LeNet-5, while STL-10, LFW, and ImageNette were tested on LeNet-5, AlexNet, and VGG-16 backbones. Each row reports the mean <inline-formula id="ieqn-296"><mml:math id="mml-ieqn-296"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> standard deviation over 10 converged runs for classification accuracy and weighted F1 score. &#x2018;O&#x2019; indicates that BN is applied, and &#x2018;X&#x2019; denotes that BN is omitted. All experiments used a batch size of 32 to ensure reliable statistics.</p>
<p>BN consistently improves both accuracy and weighted F1 across all datasets. As shown in <xref ref-type="table" rid="table-A1">Table A1</xref>, for simpler networks such as LeNet-5 (CIFAR-10/100), the improvement reaches up to 7%&#x2013;8% absolute, while for deeper networks (AlexNet, VGG-16) on larger datasets, BN contributes steady 3%&#x2013;5% gains with reduced variance. These results confirm that BN acts as a structural stabilizer against fluctuations induced by fuzzy memberships and adaptive pooling exponents.</p>
</app>
<app id="app-2">
<title>Appendix B Visualization of the Proposed Fuzzy Pooling Layer</title>
<p>To provide a qualitative understanding of how the proposed fuzzy pooling operates, we visualize intermediate feature maps and the final pooled responses under different settings. The visualization clarifies how the clustering process, learned pooling exponents, and cluster integration&#x2014;as detailed in <xref ref-type="sec" rid="s3">Section 3</xref>&#x2014;jointly determine the spatial saliency pattern produced by the layer.</p>
<p>In all cases, the process proceeds in four stages:
<list list-type="simple">
<list-item><label>(i)</label><p><bold>Fuzzy clustering:</bold> local features are grouped into <italic>K</italic> latent types via FCM [<xref ref-type="bibr" rid="ref-28">28</xref>];</p></list-item>
<list-item><label>(ii)</label><p><bold>Cluster-wise feature mapping:</bold> each cluster yields a distinct feature response map;</p></list-item>
<list-item><label>(iii)</label><p><bold>Adaptive modulation:</bold> learned exponents <inline-formula id="ieqn-297"><mml:math id="mml-ieqn-297"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> emphasize or suppress clusters according to local memberships;</p></list-item>
<list-item><label>(iv)</label><p><bold>Integration:</bold> the maps are merged using either Membership Averaging (Mavg) or Membership Maxing (Mmax).</p></list-item>
</list></p>
<p><xref ref-type="fig" rid="fig-A1">Fig. A1a</xref>,<xref ref-type="fig" rid="fig-A1">b</xref> illustrates the effect of learned pooling exponents when <inline-formula id="ieqn-298"><mml:math id="mml-ieqn-298"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0.2</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> for three clusters. Because memberships <inline-formula id="ieqn-299"><mml:math id="mml-ieqn-299"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> act as continuous weights, values of <inline-formula id="ieqn-300"><mml:math id="mml-ieqn-300"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> attenuate activation magnitudes (suppressing those regions), whereas <inline-formula id="ieqn-301"><mml:math id="mml-ieqn-301"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> amplify salient responses. In this example, the two clusters with smaller <inline-formula id="ieqn-302"><mml:math id="mml-ieqn-302"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> values (<inline-formula id="ieqn-303"><mml:math id="mml-ieqn-303"><mml:mn>0.2</mml:mn></mml:math></inline-formula>) produce enhanced contrast&#x2014;appearing brighter or darker&#x2014;while the cluster with <inline-formula id="ieqn-304"><mml:math id="mml-ieqn-304"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> becomes subdued. This contrast amplification yields visually sharper spatial boundaries, consistent with the adaptive nature of the proposed pooling described in <xref ref-type="sec" rid="s3_6">Section 3.6</xref>.</p>
<fig id="fig-A1">
<label>Figure A1</label>
<caption>
<title>Visual comparison of the proposed fuzzy pooling layer. (<bold>a</bold>,<bold>b</bold>) show two integration variants derived from identical fuzzy memberships: both start from the same clustering and per-cluster feature maps, modulate local saliency using learned exponents <inline-formula id="ieqn-306"><mml:math id="mml-ieqn-306"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> (<inline-formula id="ieqn-307"><mml:math id="mml-ieqn-307"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>5</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> in this example), and integrate responses by either soft averaging (Mavg) or dominance-based maxing (Mmax). Clusters with <inline-formula id="ieqn-308"><mml:math id="mml-ieqn-308"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> amplify features (bright or dark contrast), while <inline-formula id="ieqn-309"><mml:math id="mml-ieqn-309"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> suppress them, leading to clearer spatial delineation. (<bold>c</bold>) illustrates how the number of clusters <italic>K</italic> affects selectivity: as <italic>K</italic> increases, finer semantic regions appear, with <inline-formula id="ieqn-310"><mml:math id="mml-ieqn-310"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> showing the most pronounced separation where two clusters are enhanced and one is attenuated</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74033-fig-16.tif"/>
</fig>
<p>These qualitative results demonstrate that the proposed fuzzy pooling not only blends average and max pooling behaviors but also adaptively modulates feature intensity through the learned exponents <inline-formula id="ieqn-305"><mml:math id="mml-ieqn-305"><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>. This produces locally contrast-enhanced feature representations, helping subsequent layers distinguish boundary and texture information more effectively under varying fuzzy memberships.</p>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lecun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bottou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Haffner</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc IEEE</source>. <year>1998</year>;<volume>86</volume>(<issue>11</issue>):<fpage>2278</fpage>&#x2013;<lpage>324</lpage>. doi:<pub-id pub-id-type="doi">10.1109/5.726791</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Krizhevsky</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sutskever</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name></person-group>. <article-title>Imagenet classification with deep convolutional neural networks</article-title>. In: <conf-name>Advances in Neural Information Processing Systems</conf-name>. Vol. <volume>25</volume>. <publisher-loc>Late Tahoe, NV, USA</publisher-loc>: <publisher-name>Neuro IPS</publisher-name>; <year>2012</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hendrycks</surname> <given-names>D</given-names></string-name>, <string-name><surname>Dietterich</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Benchmarking neural network robustness to common corruptions and perturbations</article-title>. <comment>arXiv:1903.12261. 2019</comment>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Pleiss</surname> <given-names>G</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Weinberger</surname> <given-names>KQ</given-names></string-name></person-group>. <article-title>On calibration of modern neural networks</article-title>. In: <conf-name>Proceedings of the 34th International Conference on Machine Learning; 2017 Aug 6&#x2013;11; Sydney, NSW, Australia</conf-name>. <publisher-loc>New Orleans, LA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2017</year>. p. <fpage>1321</fpage>&#x2013;<lpage>30</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Minderer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Djolonga</surname> <given-names>J</given-names></string-name>, <string-name><surname>Romijnders</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hubis</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Houlsby</surname> <given-names>N</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Revisiting the calibration of modern neural networks</article-title>. <source>Adv Neural Inform Process Syst</source>. <year>2021</year>;<volume>34</volume>:<fpage>15682</fpage>&#x2013;<lpage>94</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ovadia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fertig</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>J</given-names></string-name>, <string-name><surname>Nado</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sculley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Nowozin</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Can you trust your model&#x2019;s uncertainty? Evaluating predictive uncertainty under dataset shift</article-title>. In: <conf-name>Advances in neural information processing systems</conf-name>. Vol. <volume>32</volume>. <publisher-loc>Vancouver, BC, Canada</publisher-loc>: <publisher-name>Neuro IPS</publisher-name>; <year>2019</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zafar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Aamir</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mohd Nawi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Arshad</surname> <given-names>A</given-names></string-name>, <string-name><surname>Riaz</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alruban</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A comparison of pooling methods for convolutional neural networks</article-title>. <source>Appl Sci</source>. <year>2022</year>;<volume>12</volume>(<issue>17</issue>):<fpage>8643</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app12178643</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rippel</surname> <given-names>O</given-names></string-name>, <string-name><surname>Snoek</surname> <given-names>J</given-names></string-name>, <string-name><surname>Adams</surname> <given-names>RP</given-names></string-name></person-group>. <article-title>Spectral representations for convolutional neural networks</article-title>. In: <conf-name>Advances in neural information processing systems</conf-name>. Vol. <volume>28</volume>. <publisher-loc>Montreal, QC, Canada</publisher-loc>: <publisher-name>Neuro IPS</publisher-name>; <year>2015</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Simonyan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <comment>arXiv:1409.1556. 2014</comment>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Network in network</article-title>. <comment>arXiv:1312.4400. 2013</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sermanet</surname> <given-names>P</given-names></string-name>, <string-name><surname>Reed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Anguelov</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Going deeper with convolutions</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2015 Jun 7&#x2013;12</year>; <publisher-loc>Boston, MA, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2015.7298594</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Vanhoucke</surname> <given-names>V</given-names></string-name>, <string-name><surname>Ioffe</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shlens</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wojna</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Rethinking the inception architecture for computer vision</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27&#x2013;30</conf-name>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>2818</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2016.308</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ioffe</surname> <given-names>S</given-names></string-name>, <string-name><surname>Vanhoucke</surname> <given-names>V</given-names></string-name>, <string-name><surname>Alemi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Inception-v4, inception-resnet and the impact of residual connections on learning</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence</conf-name>; <year>2017 Feb 11</year>; <publisher-loc>San Francisco, CA, USA</publisher-loc>. <volume>Vol. 31.</volume> p. <fpage>4278</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v31i1.11231</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Springenberg</surname> <given-names>JT</given-names></string-name>, <string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Brox</surname> <given-names>T</given-names></string-name>, <string-name><surname>Riedmiller</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Striving for simplicity: the all convolutional net</article-title>. <comment>arXiv:1412.6806. 2014</comment>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2016 Jun 27&#x2013;30</year>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Le</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Efficientnet: rethinking model scaling for convolutional neural networks</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2019 Jun 9&#x2013;15</year>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>6105</fpage>&#x2013;<lpage>14</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Squeeze-and-excitation networks</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2018 Jun 18&#x2013;23</year>; <publisher-loc>Salt Lake City, UT, USA</publisher-loc>. p. <fpage>7132</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2018.00745</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Woo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Park</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Kweon</surname> <given-names>IS</given-names></string-name></person-group>. <article-title>Cbam: convolutional block attention module</article-title>. In: <conf-name>Proceedings of the European Conference on Computer Vision</conf-name>. <publisher-loc>Munich, Germany</publisher-loc>: <publisher-name>ECCV</publisher-name>; <year>2018</year>. p. <fpage>3</fpage>&#x2013;<lpage>19</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Feichtenhofer</surname> <given-names>C</given-names></string-name>, <string-name><surname>Darrell</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A convnet for the 2020s</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2022 Jun 18&#x2013;24</year>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>11976</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01167</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name></person-group>. <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>. <comment>arXiv: 2010.11929. 2020</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gulcehre</surname> <given-names>C</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>K</given-names></string-name>, <string-name><surname>Pascanu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Learned-norm pooling for deep feedforward and recurrent neural networks</article-title>. In: <conf-name>Joint European Conference on Machine Learning and Knowledge Discovery in Databases</conf-name>. <publisher-loc>Nancy, France</publisher-loc>: <publisher-name>Berlin/Heidelberg, Germany: Springer</publisher-name>; <year>2014</year>. p. <fpage>530</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-662-44848-9_34</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bieder</surname> <given-names>F</given-names></string-name>, <string-name><surname>Sandk&#x00FC;hler</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cattin</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Comparison of methods generalizing max-and average-pooling</article-title>. <comment>arXiv:2103.01746. 2021</comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Radenovi&#x0107;</surname> <given-names>F</given-names></string-name>, <string-name><surname>Tolias</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chum</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Fine-tuning CNN image retrieval with no human annotation</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2018</year>;<volume>41</volume>(<issue>7</issue>):<fpage>1655</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2018.2846566</pub-id>; <pub-id pub-id-type="pmid">29994246</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zeiler</surname> <given-names>MD</given-names></string-name>, <string-name><surname>Fergus</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Stochastic pooling for regularization of deep convolutional neural networks</article-title>. <comment>arXiv:1301.3557. 2013</comment>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>S3pool: pooling with stochastic spatial sampling</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2017 Jul 21&#x2013;26</year>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>4970</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zadeh</surname> <given-names>LA</given-names></string-name></person-group>. <article-title>Fuzzy sets</article-title>. <source>Inform Control</source>. <year>1965</year>;<volume>8</volume>(<issue>3</issue>):<fpage>338</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1016/S0019-9958(65)90241-X</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dunn</surname> <given-names>JC</given-names></string-name></person-group>. <article-title>A fuzzy relative of the ISODATA process and its use in detecting compact well-separated clusters</article-title>. <source>J Cybern</source>. <year>1973</year>;<volume>3</volume>(<issue>3</issue>):<fpage>32</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1080/01969727308546046</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Bezdek</surname> <given-names>JC</given-names></string-name></person-group>. <source>Pattern recognition with fuzzy objective function algorithms</source>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Berlin/Heidelberg, Germany: Springer</publisher-name>; <year>1981</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>I</given-names></string-name></person-group>. <source>Deep learning</source>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>MIT press</publisher-name>; <year>2016</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Making convolutional networks shift-invariant again</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>. <publisher-loc>Long Beach, CA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2019</year>. p. <fpage>7324</fpage>&#x2013;<lpage>34</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ioffe</surname> <given-names>S</given-names></string-name>, <string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Batch normalization: accelerating deep network training by reducing internal covariate shift</article-title>. In: <conf-name>Proceedings of the 32nd International Conference on Machine Learning</conf-name>; <year>2015 Jul 7&#x2013;9</year>; <publisher-loc>Lille, France</publisher-loc>. p . <fpage>448</fpage>&#x2013;<lpage>56</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Santurkar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tsipras</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ilyas</surname> <given-names>A</given-names></string-name>, <string-name><surname>Madry</surname> <given-names>A</given-names></string-name></person-group>. <chapter-title>How does batch normalization help optimization?</chapter-title> In: <person-group person-group-type="editor"><string-name><surname>Bengio</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wallach</surname> <given-names>H</given-names></string-name>, <string-name><surname>Larochelle</surname> <given-names>H</given-names></string-name>, <string-name><surname>Grauman</surname> <given-names>K</given-names></string-name>, <string-name><surname>Cesa-Bianchi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Garnett</surname> <given-names>R</given-names></string-name></person-group>, editors. <source>Advances in Neural Information Processing Systems</source>. Vol. <volume>31</volume>. <publisher-loc>Montr&#x00E9;al, QC, Canada</publisher-loc>: <publisher-name>Neuro IPS</publisher-name>; <year>2018</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A improved pooling method for convolutional neural networks</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>1589</fpage>; <pub-id pub-id-type="pmid">38238357</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sharma</surname> <given-names>T</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>V</given-names></string-name>, <string-name><surname>Sudhakaran</surname> <given-names>S</given-names></string-name>, <string-name><surname>Verma</surname> <given-names>NK</given-names></string-name></person-group>. <article-title>Fuzzy based pooling in convolutional neural network for image classification</article-title>. In: <conf-name>Proceedings of the 2019 IEEE International Conference on Fuzzy Systems</conf-name>; <year>2019 Jun 23&#x2013;26</year>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/FUZZ-IEEE.2019.8859010</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Diamantis</surname> <given-names>DE</given-names></string-name>, <string-name><surname>Iakovidis</surname> <given-names>DK</given-names></string-name></person-group>. <article-title>Fuzzy pooling</article-title>. <source>IEEE Trans Fuzzy Syst</source>. <year>2020</year>;<volume>29</volume>(<issue>11</issue>):<fpage>3481</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TFUZZ.2020.3024023</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Er</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Unsupervised fuzzy neural network for image clustering</article-title>. In: <conf-name>Proceedings of the 2021 IEEE International Conference on Fuzzy Systems</conf-name>; <year>2021 Jul 11&#x2013;14</year>; <publisher-loc>Luxembourg</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/FUZZ45933.2021.9494601</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hasan</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Hossain</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Azad</surname> <given-names>A</given-names></string-name>, <string-name><surname>Alyami</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Moni</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>FP-CNN: a fuzzy pooling-based convolutional neural network for medical image classification</article-title>. <source>Comput Biol Med</source>. <year>2023</year>;<volume>166</volume>(<issue>3</issue>):<fpage>107407</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiomed.2023.107407</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>BH</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Jhang</surname> <given-names>JY</given-names></string-name></person-group>. <article-title>Design of a convolutional neural network with Type-2 fuzzy-based pooling for vehicle recognition</article-title>. <source>Mathematics</source>. <year>2024</year>;<volume>12</volume>(<issue>24</issue>):<fpage>3885</fpage>. doi:<pub-id pub-id-type="doi">10.3390/math12243885</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thakur</surname> <given-names>PS</given-names></string-name>, <string-name><surname>Verma</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Tiwari</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Analysis of time complexity of K-means and fuzzy C-means clustering algorithm</article-title>. <source>Eng Math Lett</source>. <year>2024</year>;<volume>2024</volume>(<issue>4</issue>). doi:<pub-id pub-id-type="doi">10.28919/eml/8402</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Krizhevsky</surname> <given-names>A</given-names></string-name></person-group>. <source>Learning multiple layers of features from tiny images</source>. <publisher-loc>Toronto, ON, Canada</publisher-loc>: <publisher-name>University of Toronto</publisher-name>; <year>2009</year>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>GB</given-names></string-name>, <string-name><surname>Mattar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Berg</surname> <given-names>T</given-names></string-name>, <string-name><surname>Learned-Miller</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Labeled faces in the wild: a database forstudying face recognition in unconstrained environments</article-title>. In: <conf-name>Workshop on Faces in &#x2018;Real-Life&#x2019; Images: Detection, Alignment, and Recognition</conf-name>. <publisher-loc>Marseille, France</publisher-loc>: <publisher-name>HAL</publisher-name>; <year>2008</year>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Coates</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ng</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>H</given-names></string-name></person-group>. <article-title>An analysis of single-layer networks in unsupervised feature learning</article-title>. In: <conf-name>Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics</conf-name>; <year>2011 Apr 11&#x2013;13</year>; <publisher-loc> Fort Lauderdale, FL, USA</publisher-loc>. p. <fpage>215</fpage>&#x2013;<lpage>23</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Socher</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>LJ</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Fei-Fei</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Imagenet: a large-scale hierarchical image database</article-title>. In: <conf-name>Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2009 Jun 20&#x2013;25</year>; <publisher-loc>Miami, FL, USA</publisher-loc>. p. <fpage>248</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2009.5206848</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Rasul</surname> <given-names>K</given-names></string-name>, <string-name><surname>Vollgraf</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms</article-title>. <comment>arXiv:1708.07747. 2017</comment>.</mixed-citation></ref>
</ref-list>
</back></article>

