<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">80241</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.080241</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An Explainability-Aware Transformer Framework for Brain Tumor Segmentation and Classification Using MRI</article-title>
<alt-title alt-title-type="left-running-head">An Explainability-Aware Transformer Framework for Brain Tumor Segmentation and Classification Using MRI</alt-title>
<alt-title alt-title-type="right-running-head">An Explainability-Aware Transformer Framework for Brain Tumor Segmentation and Classification Using MRI</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Jabbar</surname><given-names>Mamoona</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Jamil</surname><given-names>Uzma</given-names></name><email>uzma.jamil@gcuf.edu.pk</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Younas</surname><given-names>Muhammad</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Zafar</surname><given-names>Bushra</given-names></name></contrib>
<aff id="aff-1">
<institution>Department of Computer Science, Government College University</institution>, <addr-line>Faisalabad</addr-line>, <country>Pakistan</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Uzma Jamil. Email: <email>uzma.jamil@gcuf.edu.pk</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>4</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>1</issue>
<elocation-id>40</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>02</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>03</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_80241.pdf"></self-uri>
<abstract>
<p>Magnetic Resonance Imaging is one of the most commonly used neuro-oncology imaging modalities, which is a non-invasive mode of imaging and helps in detecting brain abnormalities in an effective way. Earlier researchers have demonstrated that brain tumor segmentation and classification can be effectively performed using deep learning techniques. Existing studies are primarily aimed at increasing prediction accuracy and provide insignificant consideration to model interpretability, limiting their practical application in clinical practice. To address this limitation, this research presents a two-stage explainable deep learning model, which combines transformer-based segmentation with an ensemble classification model that is consistent in explanations. The first stage introduces Swin-DS-HAFUNetv2, an enhanced transformer-based segmentation architecture that integrates hierarchical Swin Transformer encoders, refined hierarchical attention fusion, a contextual bottleneck transformer, and multi-scale deep supervision to improve tumor localization in T1-weighted MRI, particularly under low-contrast and irregular morphological conditions. The second stage includes the ECWMEv2 ensemble classifier, which integrates a perturbation analysis based on Grad-CAM (Gradient-weighted Class Activation Mapping) to systematically assess the consistency and clinical significance of visual explanations for candidate models. Only those architectures that exhibit stable and pathology-consistent explanations, such as ConvNeXt, Swin Transformer, and EVA02, are stored and merged by means of explanation-weighted soft voting with XGBoost-based meta-learning. Experimental evaluation on the BRISC2025 benchmark dataset indicates that Swin-DS-HAFUNetv2 has a mean Dice coefficient of 0.9782 and Intersection over Union (IoU) of 0.8656, with ECWMEv2 having a classification accuracy of 0.9917 and a Macro-F1 score of 0.9867. The mean Grad-CAM IoU of 0.692 reflects uniform and anatomically significant consistency of attention to tumor regions. These results demonstrate that the integration of explanation stability as a fundamental design principle significantly improves model robustness and interpretability to provide a methodologically validated and benchmark-level framework for future studies of multimodal and clinically oriented brain tumor analysis systems.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Brain tumor segmentation</kwd>
<kwd>classification</kwd>
<kwd>explainable artificial intelligence</kwd>
<kwd>Grad-CAM stability</kwd>
<kwd>medical image analysis</kwd>
<kwd>transformer networks</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Brain tumors can be considered one of the most aggressive and clinically complicated types of cancer that demand high levels of diagnostic accuracy to allow effective segmentation and classification of brain tumors to aid in treatment planning and prognosis evaluation [<xref ref-type="bibr" rid="ref-1">1</xref>]. The main type of imaging involved in neuro-oncology is Magnetic Resonance Imaging (MRI) because it is the only imaging modality that offers a superior soft-tissue contrast; it also has the capacity to represent anatomical and pathological structures in the brain without exposure to ionizing radiation [<xref ref-type="bibr" rid="ref-2">2</xref>]. Although it has its benefits, manual interpretation of MRI scans is time-consuming, labor-intensive, and prone to inter-observer variance, especially when dealing with tumor margins or non-uniform tissue appearance [<xref ref-type="bibr" rid="ref-3">3</xref>]. Such difficulties have increased the use of deep learning (DL) methods in automated brain tumor detection, segmentation, and classification [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>Convolutional Neural Networks (CNNs) are deep learning models that have been shown to achieve high performance in medical image recognition because they are capable of learning hierarchical visual feature representations [<xref ref-type="bibr" rid="ref-5">5</xref>]. Tumor localization and classification tasks have also been performed with models based on Convolutional Neural Network (CNN), which has contributed to better reproducibility as well as eliminated the need to perform manual interpretation [<xref ref-type="bibr" rid="ref-6">6</xref>]. Nonetheless, standard CNN designs frequently fail to represent the role of long-range spatial interactions and so perform poorly when using large tumor volumes or low-contrast MRI volumes [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>In order to mitigate such limitations, transformer-based architectures have recently become promising alternatives. Transformers can use self-attention processes to extract world contextual association in images and enhance the feature representation in challenging medical imaging tasks [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. However, several current transformer-based segmentation systems are two-dimensional and do not use successful multi-scale feature fusion and deep supervision, limiting their inference to different tumor shapes and structures [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>In classification, modern architectures, including CoAtNet, ConvNeXt, EVA02, Swin Transformer, and MaxViT, have been found to possess highly representative abilities, as compared to the standard CNN designs [<xref ref-type="bibr" rid="ref-11">11</xref>]. The models have better specificity in identifying subtypes of brain tumors and prognosis [<xref ref-type="bibr" rid="ref-12">12</xref>]. Nonetheless, segmentation and classification continue to be commonly considered as individual processes, with very few studies endeavoring to combine the two processes into common frameworks to perform automated brain tumor analysis [<xref ref-type="bibr" rid="ref-13">13</xref>].</p>
<p>The second major problem with medical artificial intelligence is the lack of interpretability. Most deep learning systems are black-box systems, and thus, a clinician can hardly comprehend how the predictions were made [<xref ref-type="bibr" rid="ref-14">14</xref>]. To solve this problem, explainable artificial intelligence (XAI) approaches like Gradient-weighted Class Activation mapping (Grad-CAM) and saliency mapping have been suggested to visualize image regions that affect model predictions [<xref ref-type="bibr" rid="ref-15">15</xref>]. Nevertheless, in the majority of studies, such explanations are not included in the modeling development or evaluation, but they are applied as post hoc visualization tools [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>Recent studies identify an important but understudied determinant of XAI method reliability in medical imaging models, including explanation stability, or the consistency of saliency maps with small perturbations of inputs [<xref ref-type="bibr" rid="ref-17">17</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>]. Though attention-based architectures like Swin Transformer have demonstrated increased robustness in MRI analysis [<xref ref-type="bibr" rid="ref-20">20</xref>], there are still numerous hybrid models that integrate deep learning and explanation consistency plus reliability, which lack formal quantitative evaluation mechanisms [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
<p>As a solution to these problems, this paper presents a proposal of an explainability-aware framework that unites tumor segmentation using a transformer and explanation-consistent ensemble classification in a unified pipeline. The proposed framework presents Swin-DS-HAFUNetv2, which is a transformer-based segmentation framework with hierarchical attention fusion and deep supervision to enhance the localization of the tumor boundary. Contrary to the current method of applying explainability to the post hoc visualization, an Explanation-Consistent Weighted Meta-Ensemble (ECWMEv2) classification framework takes into account Grad-CAM explanation stability as a quantitative metric in selecting a model and building an ensemble. The proposed solution should enhance both the model robustness and interpretability, as well as clinical reliability by adding consistency of the explanations as a quantitative measure of evaluation. The research objectives of this study are as follows.
<list list-type="simple">
<list-item><label>OBJ 1.</label>
<p>To improve segmentation accuracy for low-contrast brain MRI using hierarchical attention and deep supervision within transformer-based architectures.</p></list-item>
<list-item><label>OBJ 2.</label>
<p>To improve the robustness and reliability of classification through an ensemble strategy guided by explanation consistency under input perturbations.</p></list-item>
<list-item><label>OBJ 3.</label>
<p>To design an interpretable model selection and fusion process that incorporates quantifiable Grad-CAM stability metrics, including IoU, cosine similarity, and pass rate as selection criteria.</p></list-item>
</list></p>
<p>Overall, the key contributions of this work may be summarized in the following way:<list list-type="bullet">
<list-item>
<p><bold>Enhanced Transformer-Based Segmentation Architecture:</bold> The work presents Swin-DS-HAFUNetv2, an improved transformer-based segmentation framework that incorporates Swin Transformer encoders, hierarchical attention fusion, and deep supervision to enhance the multi-scale contextual feature coverage and capture more precise tumor boundary segmentation in the brain MRI images.</p></list-item>
<list-item>
<p><bold>Explanation-Consistent Ensemble Classification Framework:</bold> A Grad-CAM explanation stability is added to an Explanation-Consistent Ensemble Classification Framework (ECWMEv2), where Grad-CAM explanation stability is used as a model selection and weighting criterion. The models of candidate classification are tested on the input perturbation, and only those that generate stable and clinically meaningful attention maps are stored in the ensemble. Explanation-weighted soft voting and XGBoost-based meta-learning layer are used to combine the retained models to improve the reliability and interpretability of the classification.</p></list-item>
<list-item>
<p><bold>Perturbation-Based Quantitative Evaluation of Explanation Stability:</bold> This research proposes a perturbation-based Grad-CAM consistency analysis, which is used to provide a formal assessment of explanation stability. The quantitative metrics of similarity, such as Intersection over Union (IoU), cosine similarity, and pass rate, are used to gauge the stability of the explanations, which allows the aspect of consistency of the explanations to be used as a formal evaluation and filtering mechanism by the model, but not as a visual interpretation tool.</p></list-item>
</list></p>
<p>The rest of this paper will be structured as follows. <xref ref-type="sec" rid="s2">Section 2</xref> provides a review of the corresponding literature concerning the same deep learning-based brain tumor segmentation, classification, and explainable artificial intelligence in medical imaging. <xref ref-type="sec" rid="s3">Section 3</xref> provides the suggested methodology, such as the Swin-DS-HAFUNetv2 segmentation model and explanation-consistent ECWMEv2 ensemble classification strategy. <xref ref-type="sec" rid="s4">Section 4</xref> outlines the experimental design and gives the findings and discussion of the suggested framework. <xref ref-type="sec" rid="s5">Section 5</xref> concludes the overall findings of the proposed framework, discusses the limitations, and future work.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Extensive literature has been done on deep learning methodology on brain tumor segmentation and classification. The methods of CNN are still popular in the extraction of discriminative tumor characteristics of MRI images. It has been shown in a number of studies that deep convolutional architectures are effective in brain tumor classification and localization tasks [<xref ref-type="bibr" rid="ref-23">23</xref>&#x2013;<xref ref-type="bibr" rid="ref-25">25</xref>]. Moreover, some works include auxiliary clinical data, like patient age or demographic characteristics, to add the context of diagnosis and improve the classification performance [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
<p>Various researchers have tried to integrate segmentation and classification in hybrid learning models. As an example, joint tumor segmentation and classification have been carried out by clustering-based and probabilistic neural network algorithms [<xref ref-type="bibr" rid="ref-27">27</xref>]. In other words, Grad-CAM or related XAI methods are used to visualize regions of interest in tumor classification [<xref ref-type="bibr" rid="ref-28">28</xref>]. Nonetheless, such explanations are normally given in a qualitative format and are seldom substantiated in the quantitative type of similarity, like Intersection over Union or cosine similarity. Deep learning models that are based on optimization have also been suggested to enhance classification performance, but they do not typically provide mechanisms to evaluate the explanation robustness [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
<p>XAI approaches based on attribution, like SHAP (SHapley Additive explanations) and LIME (Local Interpretable Model-agnostic explanations), and Grad-CAM offer explanations on a feature or pixel level interpretations, but cannot be applied to dense spatial medical images when the consistency of an explanation needs to be at a region level [<xref ref-type="bibr" rid="ref-31">31</xref>]. Despite the extensive application of data augmentation to enhance model generalization, its effect on explanation fidelity has been poorly studied [<xref ref-type="bibr" rid="ref-32">32</xref>]. Other classical methods of machine learning that include clustering and SVM-based classifiers have also been studied, but these models do not provide the interpretability and strength needed in clinical practice [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>More recent works have investigated the hybrid designs of convolutional and transformer-based models to enhance the level of feature representation and classification accuracy [<xref ref-type="bibr" rid="ref-35">35</xref>&#x2013;<xref ref-type="bibr" rid="ref-38">38</xref>]. However, explainability within these models is often confined to visualization, and the stability of explanation when there is a perturbation in inputs is not often considered. Neuro-oncology surveys have also shown that XAI methods have been used predominantly as a visual interpretation, and not in training or model validation activities [<xref ref-type="bibr" rid="ref-39">39</xref>&#x2013;<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
<p>Other medical imaging areas, such as stroke detection and pulmonary disease analysis, have also reported similar limitations, with the model interpretability and the robustness of the model explanation having not yet been adequately studied [<xref ref-type="bibr" rid="ref-43">43</xref>&#x2013;<xref ref-type="bibr" rid="ref-47">47</xref>]. In spite of the high classification accuracy of such advanced architectures as EfficientNet, dilated convolutional networks, current evaluation frameworks mostly focus on predictive power without evaluating the stability of the explanation or the network&#x2019;s strength to perturbations in the inputs [<xref ref-type="bibr" rid="ref-48">48</xref>&#x2013;<xref ref-type="bibr" rid="ref-52">52</xref>].</p>
<p>Numerous segmentation architectures have been put forward to enhance the localization of tumor boundaries. Attention-based models, including ABANet, strive to improve the awareness of the boundaries but might not have enough contextual fusion mechanisms [<xref ref-type="bibr" rid="ref-53">53</xref>]. The use of traditional biomedical segmentation networks, e.g., U-Net and U-Net&#x002B;&#x002B;, continues to be quite common, but they are unable to capture long-range dependencies effectively unless using transformer-based encoders or deep supervision [<xref ref-type="bibr" rid="ref-54">54</xref>]. Other interpretable segmentation models have graphical interpretability, but they lack explanation measures in model choice or ensemble generation [<xref ref-type="bibr" rid="ref-55">55</xref>].</p>
<p>Moreover, the existing studies on transfer learning and cross-dataset generalization are primarily aimed at enhancing the accuracy of the classification but ignore the fidelity of the explanation [<xref ref-type="bibr" rid="ref-56">56</xref>]. Ensemble-based methods tend to increase the predictive accuracy without assessing the stability of visual explanations during imaging distortions in the real world [<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>]. Comparison of CNN with transformer models also does not focus on explaining stability, but still prioritizes performance measures [<xref ref-type="bibr" rid="ref-59">59</xref>]. More recent explainable diagnosis systems also do not formally test the reliability of the explanation in terms of perturbation [<xref ref-type="bibr" rid="ref-60">60</xref>&#x2013;<xref ref-type="bibr" rid="ref-62">62</xref>]. Even though, e.g., U-Net&#x002B;&#x002B; and modified SegNet can achieve better segmentation results, they cannot yet incorporate quantitative explanatory stability into their own decision processing pipelines [<xref ref-type="bibr" rid="ref-63">63</xref>,<xref ref-type="bibr" rid="ref-64">64</xref>].</p>
<p>Despite the large number of deep learning models that were proposed in brain tumor segmentation and classification, there are also a number of limitations in the current research. There are numerous methods that only concentrate on segmentation or classification, and only a small number of methods have attempted to combine both of them into a cohesive framework. In addition, the stability and robustness of visual explanations under input perturbations have not been systematically evaluated or utilized as a quantitative criterion for model selection in existing studies. To better highlight these limitations, <xref ref-type="table" rid="table-1">Table 1</xref> summarizes the key research gaps in representative existing studies.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Gap analysis of existing studies in MRI-based brain tumor segmentation and classification.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th rowspan="2">Ref.</th>
<th align="center" rowspan="2">Model</th>
<th colspan="3">Research Gap in Existing Studies</th>
</tr>
<tr>
<th>Segmentation &#x002B; Classification</th>
<th>XAI Used for Model Selection</th>
<th>Explanation Stability Evaluated</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>Swin-HAFUNet &#x002B; Pre-trained models</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>Explainable CNN</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>Hybrid CNN&#x2013;Transformer</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>Hybrid clustering with probabilistic neural networks</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>PSO-QESVM classifier</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>ResNet50</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Hybrid DL &#x002B; ML with SHAP explainability</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>XAI EfficientNetB0</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>Dilated parallel CNN</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>Transfer-learning CNNs (ResNet, DenseNet, EfficientNet)</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>CNN and Hybrid Transformer</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td>Modified SegNet &#x002B; Hybrid DL</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methodology</title>
<p>The proposed research presents a two-stage explainable deep learning system to be applied in automated brain tumor detection of MRI images. The framework aims to address two key challenges: (1) the accurate delineation of tumor boundaries in complex anatomical structures and (2) the classification of tumor types with predictable and consistent outcomes despite perturbation with inputs. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> presents the general structure with emphasis on the sequential work with raw MRI input by segmentation, classification, and explainability modules to generate the ultimate diagnostic output. With medical images, explainability methods must be spatially continuous and anatomically consistent to enable interpretation that is of clinical significance. Brain MRI classification requires clarification to determine where the discriminative areas are, within the tumor boundaries, as opposed to assigning significance to abstract dimensions of features. Accordingly, this study adopts saliency-based visualization through Grad-CAM, which generates spatially coherent heatmaps aligned with learned convolutional and transformer feature hierarchies. SHAP and LIME are attribution-based methods that can be applied to offer attribute-wise interpretability, especially when using tabular or low-dimensional data. Then, they are not so suited to the use of complex medical imaging when the precise localization of the areas and maintenance of the anatomical structure are essential. This is because they are not excluded from this study because of any shortage in their efficacy, but rather a methodological choice guided by the necessity of having spatially meaningful interpretability. Grad-CAM is therefore not used as only a post-hoc visualization model, but as a part of the whole framework, and it plays a role in model selection, ensemble weighting, and perturbation-based assessment of consistency of explanation.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of the proposed two-stage explainability-aware framework for brain tumor segmentation, classification, and explanation consistency analysis.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-1.tif"/>
</fig>
<p>The subsections that follow explain each component in detail, with the preprocessing stage being the first step to all the downstream learning activities.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset</title>
<p>The BRISC2025 dataset has been used in this study as a benchmark of brain tumor segmentation and classification by using MRI images [<xref ref-type="bibr" rid="ref-20">20</xref>]. The data is freely available on such official repositories as Kaggle [<xref ref-type="bibr" rid="ref-13">13</xref>], Figshare [<xref ref-type="bibr" rid="ref-16">16</xref>], and Zenodo, which makes it transparent and allows researchers to reproduce the obtained results. The BRISC2025 dataset was assembled by combining brain MRI images from various publicly available datasets and made consistent in a single standard reference dataset. Once integrated, there was a data refining process involving quality control, annotation control, and expert control. A certified radiologist and physician reviewed all tumor annotations and made sure that they were clinically plausible and consistent.</p>
<p>This data is 6000 brain MRI scans, a single slice, and T1-weighted. The current dataset release does not include other MRI modalities, including T2, FLAIR, or contrast-enhanced MRI sequences. The MRI images have been stored in the form of JPEG files, but the tumor segmentation masks of pixels have been stored as PNG files. The dataset consists of MRI slices with three anatomical orientations, i.e., axial, coronal, and sagittal, so that features that are invariant to imaging orientation can be learned. The data set favors two concomitant learning exercises. The initial challenge is pixel-level tumor segmentation, in which binary masks are given to define the entire tumor area. The second one is image-level multi-class tumor classification, which has four diagnostic groups: Glioma, Meningioma, Pituitary Tumor, and No Tumor. As opposed to volumetric datasets, which consider the glioma segmentation process only, BRISC2025 allows evaluating tumor localization and multi-class tumor diagnosis in a single framework.</p>
<p>To evaluate the experimental value, the data set was separated into training and testing data to prevent leakage of data. Precisely, the training (5000) and the testing (1000) images, which correspond to a data split ratio of 83% and 17%, respectively [<xref ref-type="bibr" rid="ref-20">20</xref>]. <xref ref-type="table" rid="table-2">Table 2</xref> gives a class-wise representation of the images in the training set and testing set.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Class-wise distribution of training and testing samples in the BRISC2025 dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Split</th>
<th>Images</th>
<th>Glioma</th>
<th>Meningioma</th>
<th>Pituitary</th>
<th>No Tumor</th>
</tr>
</thead>
<tbody>
<tr>
<td>Train</td>
<td>5000</td>
<td>1147</td>
<td>1329</td>
<td>1457</td>
<td>1067</td>
</tr>
<tr>
<td>Test</td>
<td>1000</td>
<td>254</td>
<td>306</td>
<td>300</td>
<td>140</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-3">Table 3</xref> shows sample segmentation cases in tumor types and imaging planes that are provided as original MRI slices, binary masks, and tumor overlays.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Examples of segmentation masks across different imaging planes and tumor types: (<bold>a</bold>) Glioma, (<bold>b</bold>) Meningioma, (<bold>c</bold>) Pituitary, and (<bold>d</bold>) No Tumor.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th></th>
<th>Original Image</th>
<th>Binary Mask</th>
<th>Tumor Region Overlay</th>
<th>Original Image</th>
<th>Binary Mask</th>
<th>Tumor Region Overlay</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2"><bold>Axial</bold></td>
<td align="center" colspan="3"><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMES_80241-inline-1.tif"/></td>
<td align="center" colspan="3"><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMES_80241-inline-2.tif"/></td>
</tr>
<tr>
<td/>
<td align="center" colspan="3">(<bold>a</bold>) Sample of Glioma segmentation</td>
<td align="center" colspan="2">(<bold>b</bold>) Sample of Meningioma segmentation</td>
</tr>
<tr>
<td rowspan="2"><bold>Coronal</bold></td>
<td align="center" colspan="3"><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMES_80241-inline-3.tif"/></td>
<td align="center" colspan="3"><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMES_80241-inline-4.tif"/></td>
</tr>
<tr>
<td/>
<td align="center" colspan="3">(<bold>c</bold>) Sample of Pituitary segmentation</td>
<td align="center" colspan="2">(<bold>d</bold>) Sample of No Tumor case</td>
</tr>
<tr>
<td rowspan="2"><bold>Sagittal</bold></td>
<td align="center" colspan="6"><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMES_80241-inline-5.tif"/></td>
</tr>
<tr>
<td/>
<td align="center" colspan="5">(<bold>e</bold>) Example illustrating minor annotation variability</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Although volumetric multimodal MRI data are conventional in clinical neuro-oncology workflows, the adoption of 2D T1-weighted images [<xref ref-type="bibr" rid="ref-25">25</xref>] in BRISC2025 indicates a limitation on the dataset level, as opposed to a limitation on the proposed framework. The dataset is categorically aimed at benchmarking algorithms, comparing models, and developing methodologies in brain tumor segmentation and classification, as opposed to actual clinical diagnostic use. Notably, the suggested framework and explainability-aware ensemble approach are modality-neutral and can be generalized to volumetric and multimodal MRI environments in case relevant annotated data is made accessible. Therefore, the findings reported in this study should be interpreted as internal methodological validation under the context determined by the dataset.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Data Preprocessing</title>
<p>The preprocessing stage is critical in determining the quality and reliability of segmentation and classification, as it determines the performance of all the subsequent learning stages directly [<xref ref-type="bibr" rid="ref-4">4</xref>]. The BRISC2025 dataset included 6000 T1-weighted MRI slices each; a four-step preprocessing pipeline has been used in this study [<xref ref-type="bibr" rid="ref-20">20</xref>]. It involved the skull stripping, standardization of intensities, scaling of images, and standardization of the space of images, which aims at addressing typical imaging artifacts in multi-institutional datasets of MRI [<xref ref-type="bibr" rid="ref-13">13</xref>]. Intensity thresholding, morphological filtering, and connected component analysis were combined to strip the skull to eliminate non-brain tissue, with peripheral tumor regions preserved [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. This was followed by intensity normalization (to zero mean and unit variance) and min-max normalization to [0, 1] to reduce variations caused by scanners and improve model convergence [<xref ref-type="bibr" rid="ref-16">16</xref>]. All MRI slices were also rescaled to 224 &#x00D7; 224 pixels using bilinear interpolation to maintain the size of the input and important anatomical features without making the computation intensive. A number of augmentation methods were included in the model training to enhance generalization and minimize overfitting. The augmentation pipeline included random rotations, Gaussian blur, and noise RNA, which are the changes that are common in clinical MRI images. This preprocessing system is anatomically consistent and can see tumors, but provides standardized information to the subsequent steps of segmentation and classification. The preprocessing steps, the parameters, are summarized as illustrated in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Preprocessing configuration and parameters used in this study.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Step</th>
<th>Operation</th>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" rowspan="4">Skull Stripping</td>
<td>Intensity Thresholding</td>
<td>Global Threshold</td>
<td>15%&#x2013;85% dynamic range of the original slice</td>
</tr>
<tr>
<td>Morphological Filtering</td>
<td>Structuring Element</td>
<td>Elliptical kernel, size 5 &#x00D7; 5</td>
</tr>
<tr>
<td>Connected Component Analysis</td>
<td>Region Selection</td>
<td>Largest contiguous region retained</td>
</tr>
<tr>
<td>Brain Mask Application</td>
<td>Masking Operation</td>
<td>Element-wise multiplication with a binary mask</td>
</tr>
<tr>
<td rowspan="2">Intensity Normalization</td>
<td>Standardization</td>
<td>Z-score Normalization</td>
<td>(x &#x2212; &#x03BC;)/&#x03C3; per slice</td>
</tr>
<tr>
<td>Min-Max Scaling</td>
<td>Pixel Rescaling</td>
<td>[0, 1]</td>
</tr>
<tr>
<td rowspan="2">Spatial Resizing</td>
<td>Image Resizing</td>
<td>Interpolation Method</td>
<td>Bilinear Interpolation</td>
</tr>
<tr>
<td>Target Resolution</td>
<td>Image Size</td>
<td>224 &#x00D7; 224</td>
</tr>
<tr>
<td rowspan="2">Format Standardization</td>
<td>Channel Configuration</td>
<td>Grayscale to RGB</td>
<td>Converted to 3 channels for compatibility</td>
</tr>
<tr>
<td>Data Type</td>
<td>Floating Point Format</td>
<td>32-bit Float</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Brain Tumor Segmentation Using Swin-DS-HAFUNetv2</title>
<p>After the preprocessing, brain tumor segmentation is the second most significant phase of the two-stage system suggested, since it should provide adequate localization of the tumor to ensure strong classification and soundness of visual explanations. Inaccurate segmentation may result in unreliable Activation of irrelevant backgrounds, misclassification, and untrustworthy explanatory visualizations.</p>
<p>To overcome these issues, Swin-DS-HAFUNetv2, an improved transformer-based segmentation architecture, was proposed as a part of the given study and developed to be able to produce tumor boundaries and long-range contextual interactions in T1-weighted brain MRI scans. The Swin-DS-HAFUNetv2 is built on top of the Swin-DS-HAFUNet [<xref ref-type="bibr" rid="ref-20">20</xref>], which has been previously presented and incorporates superior hierarchical attention fusion and contextual transformer bottleneck design along with multi-scale deep supervision.</p>
<p>The overall architecture (<xref ref-type="fig" rid="fig-2">Fig. 2</xref>) is in the form of an encoder-decoder architecture with attention-directed skip connections. Swin Transformer blocks will be used to extract the feature representations as hierarchies by the encoder and then reassemble the segmentation results gradually in high-resolution by adopting adaptive context aggregation and multi-scale feature integration by the decoder. Four primary elements are going to be included in the proposed architecture.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Block diagram of the proposed Swin DS HAFUNetv2 model for brain tumor segmentation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-2.tif"/>
</fig>
<p><list list-type="bullet">
<list-item>
<p>Hierarchical Swin Transformer encoding</p></list-item>
<list-item>
<p>Enhanced Hierarchical Attention Fusion (E-HAFv2)</p></list-item>
<list-item>
<p>Contextual Bottleneck Transformer (CBT)</p></list-item>
<list-item>
<p>Deep supervision-based decoder refinement</p></list-item>
</list></p>
<p>All these elements enhance the segmentation performance, especially in low contrast tumors and abnormal tumor morphologies.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Tokenization and Swin Transformer Encoder Design</title>
<p>The encoder consists of a hierarchical Swin Transformer backbone, which consists of extracting multi-scale contextual representations of the input MRI slices. Individual input MRI images are downsampled to 224 &#x00D7; 224 &#x00D7; 3 and broken down into 4 &#x00D7; 4 non-overlapping patches, which is the typical Swin Transformer tokenization approach. The result of this operation is: <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mfrac><mml:mn>224</mml:mn><mml:mn>4</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mn>224</mml:mn><mml:mn>4</mml:mn></mml:mfrac><mml:mo>=</mml:mo><mml:mn>56</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>56</mml:mn></mml:math></inline-formula> resulting in 3136 visual tokens. The patches are, in turn, embedded into an embedding space, where a linear embedding layer is used to give a feature representation of the number of channels of an embedding dimension of 96.</p>
<p>The encoder has four layers of Swin Transformer that are made of window-based multi-head self-attention blocks and feed-forward layers. The shifted-window attention mechanism enables the network to model both local and global contextual dependencies while maintaining linear computational complexity. Between stages, patch merging operations are applied to reduce spatial resolution and increase feature dimensionality. This hierarchical design enables the encoder to progressively capture richer semantic representations of tumor structures. The spatial resolution and channel dimensions across encoder stages are summarized in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Spatial channel dimensions across encoder stages.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Stage</th>
<th>Resolution</th>
<th>Channels</th>
</tr>
</thead>
<tbody>
<tr>
<td>Stage 1</td>
<td>56 &#x00D7; 56</td>
<td>96</td>
</tr>
<tr>
<td>Stage 2</td>
<td>28 &#x00D7; 28</td>
<td>192</td>
</tr>
<tr>
<td>Stage 3</td>
<td>14 &#x00D7; 14</td>
<td>384</td>
</tr>
<tr>
<td>Stage 4</td>
<td>7 &#x00D7; 7</td>
<td>768</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>This hierarchical feature extraction enables the model to acquire fine-tumor boundary information as well as global context relational information across the image.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Enhanced Hierarchical Attention Fusion (E-HAFv2)</title>
<p>Swin-DS-HAFUNetv2 uses an Enhanced Hierarchy Attention Fusion (E-HAFv2) module to minimize the semantic gap between the encoder and decoder representations. Unlike in the approaches of traditional skip connections, which simply concatenate encoder feature maps and decoder feature maps, E-HAFv2 uses attention-guided fusion to highlight the tumor-relevant areas and reduce the background noise.</p>
<p>As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, blocks of encoder and decoder characteristics with the same spatial resolution are initially optimized by Swin Transformer to enhance the strength of the contextual features. The maps of the features obtained are then concatenated, and a 1 &#x00D7; 1 convolution layer is applied to reduce the dimension. After this, a Squeeze-and-Excitation (SE) gating strategy is then used to recalibrate channel-wise feature responses adaptively, maximizing discriminative tumor features and reducing redundant/less relevant information.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>E-HAFv2 module with attention-guided skip integration and channel recalibration.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-3.tif"/>
</fig>
<p>Such a hierarchical approach of combining attention enhances the accuracy of tumor boundaries detection, especially in areas of the brain that are structurally complex, like the ventricles and cortical folds.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Contextual Bottleneck Transformer (CBT)</title>
<p>Swin-DS-HAFUNetv2 has a Contextual Bottleneck Transformer (CBT) in the deepest level of the network to obtain global contextual dependencies across the entire feature map. As illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, CBT substitutes the conventional convolutional bottleneck with a transformer-based module, which incorporates:
<list list-type="bullet">
<list-item>
<p>Multi-Head Self-Attention (MHSA)</p>
</list-item>
<list-item>
<p>Layer Normalization</p></list-item>
<list-item>
<p>Feed-Forward Networks (FFN)</p></list-item>
<list-item>
<p>Residual connections</p></list-item>
</list></p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Contextual Bottleneck Transformer module used at the bottleneck of Swin DS HAFUNetv2.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-4.tif"/>
</fig>
<p>This design allows the model to capture long-range dependencies across tumor structures that may span large spatial regions in MRI images. Additionally, CBT enhances the stability of feature propagation and allows the network to separate diffuse tumor boundaries more effectively than would be suitably captured by local convolutional filters.</p>
</sec>
<sec id="s3_3_4">
<label>3.3.4</label>
<title>Decoder Architecture and Adaptive Context Aggregation</title>
<p>The decoder gradually recombines high-resolution segmentation maps with the help of hierarchical encoder attributes by using patch expansion and attention-fused features. The three major operations that are carried out by each decoder stage are:<list list-type="order">
<list-item>
<p><bold>Patch Expansion:</bold> Progressively, patch expansion operations (i.e., transposed convolution or interpolation) are carried out on the feature maps to restore the spatial resolution.</p></list-item>
<list-item>
<p><bold>Attention-Guided Feature Fusion:</bold> The E-HAFv2 module is used to combine encoder features with decoder representations, which enables the decoder to take advantage of low-level spatial features and high-level semantic context.</p></list-item>
<list-item>
<p><bold>Contextual Feature Refinement:</bold> Feature maps go further to be refined with convolutional blocks with contextual enhancement features.</p></list-item>
</list></p>
<p>Deep supervision is used across various levels of the decoder to stabilise the training and enhance gradient propagation. The results of auxiliary segmentation (DS1, DS2, DS3) at intermediate resolutions are added to the total loss function, as shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Adaptive context aggregation with multi-level deep supervision in Swin DS HAFUNetv2.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-5.tif"/>
</fig>
<p>This multi-scale supervision strategy enhances the consistency of the segmentation and increases the speed of the network convergence.</p>
</sec>
<sec id="s3_3_5">
<label>3.3.5</label>
<title>Loss Function and Optimization</title>
<p>Given the class imbalance that is typically seen in brain tumor segmentation problems, a hybrid loss function is used to trade region overlap and pixel-level classification performance.</p>
<p>The segmentation loss is defined as in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>seg&#x00A0;</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Dice</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Focal Tversky&#x00A0;</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>To add deep supervision, the overall training goal can be a combination of the losses of results of the intermediate decoder as represented in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>DS</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.6</mml:mn><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>DS</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.3</mml:mn><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>DS</mml:mtext></mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>This weighted representation will guarantee that the high-resolution predictions will dominate the optimization process, with intermediate outputs giving auxiliary gradient signals. The Adam optimizer that employs a ReduceLROnPlateau learning rate schedule is used to optimize the network. <xref ref-type="table" rid="table-6">Table 6</xref> provides a summary of the training setup of Swin-DS-HAFUNetv2.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Training configuration for Swin DS HAFUNetv2.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Input resolution</td>
<td>224 &#x00D7; 224</td>
</tr>
<tr>
<td>Optimizer</td>
<td>Adam</td>
</tr>
<tr>
<td>Initial learning rate</td>
<td>1 &#x00D7; 10<sup>&#x2212;4</sup></td>
</tr>
<tr>
<td>Learning rate scheduler</td>
<td>ReduceLROnPlateau</td>
</tr>
<tr>
<td>Batch size</td>
<td>8</td>
</tr>
<tr>
<td>Number of epochs</td>
<td>50</td>
</tr>
<tr>
<td>Loss function</td>
<td>Dice &#x002B; CE &#x002B; Focal Tversky</td>
</tr>
<tr>
<td>Deep supervision weights</td>
<td>1.0, 0.6, 0.3</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The output segmentation of Swin-DS-HAFUNetv2 is considered to be the key element of the proposed explainability-aware framework, as it is the main input of the downstream classification phase. The framework provides classification models with pure tumor areas instead of unrelated background or imaging artifacts, as it explicitly isolates tumor areas and then classifies them. This tumor-centric representation is not only more accurate at making classification predictions, but also the visualizations provided by the corresponding Grad-CAM are more reliable. Specifically, Swin-DS-HAFUNetv2 generates a tumor mask that limits the spatial concentration of Grad-CAM explanations in choices, which ensures that saliency maps reflect anatomically significant parts of the tumor. The correct tumor segmentation is thus in direct correspondence with the explanation-consistency filtering strategy used in the ensemble classification phase, where attention maps are considered in anatomically relevant areas. This leads to saliency maps created that are more in line with tumor boundaries and have a lower rate of false activations, thus enhancing the interpretability and clinical accuracy of the prediction made by the model.</p>
<p>In general, Swin-DS-HAFUNetv2 offers a powerful and contextual segmentation backbone that can be seamlessly combined with an explainability-aware classification pipeline. The architecture shows a high level of scalability in terms of feature representation, interpretability, and performance of segmentation across the MRI slices in the BRISC2025 dataset, and this indicates the methodological and clinical viability of the proposed two-stage architecture.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Brain Tumor Classification with ECWMEv2 Framework</title>
<p>Its classification stage is based on the tumor-centered areas obtained on the segmentation result of Swin-DS-HAFUNetv2, and the classification models are only trained to capture the discriminative features of the tumor area instead of the background structures. The second phase of the proposed framework is devoted to the segmentation of the segmented brain tumors into four clinically significant categories: Glioma, Meningioma, Pituitary Tumor, and No Tumor. This phase uses a new ensemble architecture named ECWMEv2 (Explanation-Consistent Weighted Meta-Ensemble v2) that aims at maximizing predictive accuracy, as well as providing the property of interpretability, explanation stability, and resistance to perturbations in the input.</p>
<p>In contrast to traditional ensemble procedures that can only use predictive confidence, ECWMEv2 incorporates explanation-based criteria in several levels of model selection and fusion. The former is based on a Grad-CAM stability analysis, which is used to narrow down the models whose saliency maps are not sensitive to input perturbation. Models with stable, localized, and clinically meaningful forms of attention are the only ones that are kept when building up an ensemble. An explanation-weighted soft voting system is then used to aggregate the filtered models, with a higher voting weight being assigned to models with a more focused and consistent explanation. Finally, one more method that learns on the model level logit predictions to enhance the robustness and generalization of the final ensemble is a meta-learning layer, which is built on top of the final predictions and is based on XGBoost. The overall ECWMEv2 explainability-aware classification pipeline is illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Detailed architecture of the proposed Explanation-Consistent Weighted Meta-Ensemble (ECWMEv2) classification framework.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-6.tif"/>
</fig>
<sec id="s3_4_1">
<label>3.4.1</label>
<title>Candidate Models and Feature Extraction</title>
<p>The ECWMEv2 ensemble uses a pool of five transformer-based and convolutional neural architectures, which were chosen due to their variety of architectural biases and performance on medical imaging tasks. The model pool includes CoAtNet, ConvNeXt, Swin Transformer, MaxViT, and EVA02. CoAtNet is a combination of the convolution and self-attention features that help to compromise local spatial details and global dependencies. ConvNeXt is an updated design of CNNs that also considers design ideas of transformers but preserves the performance of convolutional operations. Swin Transformer applies window-based hierarchical self-attention, which allows learning fine-grained spatial structures in a scaled manner. MaxViT makes use of both block-based attention and grid-based attention to effectively combine local and global features. Lastly, EVA02 is a large-capacity vision transformer having trainable parameters with masked image models, which has high transfer learning in low-data medical imaging. All models are initialized with ImageNet-1K pretrained weights and fine-tuned on tumor-focused regions extracted from the BRISC2025 dataset using the output of the segmentation module. This guarantees both the localization of representation and the diagnostic relevance of the models. <xref ref-type="table" rid="table-7">Table 7</xref> gives a detailed description of the model specifications and training configurations.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Summary of candidate models for ECWMEv2.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Architecture Type</th>
<th>Parameters</th>
<th>Pretrained On</th>
<th>Key Strength</th>
</tr>
</thead>
<tbody>
<tr>
<td>CoAtNet</td>
<td>Conv &#x002B; Transformer</td>
<td>85 M</td>
<td>ImageNet-1K</td>
<td>Hybrid inductive bias</td>
</tr>
<tr>
<td>ConvNeXt</td>
<td>Convolutional</td>
<td>89 M</td>
<td>ImageNet-1K</td>
<td>Efficiency and generalization</td>
</tr>
<tr>
<td>Swin Transformer</td>
<td>Hierarchical Transformer</td>
<td>88 M</td>
<td>ImageNet-1K</td>
<td>Local feature modeling</td>
</tr>
<tr>
<td>MaxViT</td>
<td>Global-Local Transformer</td>
<td>120 M</td>
<td>ImageNet-1K</td>
<td>Comprehensive attention modeling</td>
</tr>
<tr>
<td>EVA02</td>
<td>ViT with Masked Pretraining</td>
<td>304 M</td>
<td>EVA Pretraining</td>
<td>Strong generalization</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>Grad-CAM Consistency Filtering for Explanation Reliability</title>
<p>Grad-CAM-based explanation consistency is used to assess whether candidate models have stable and clinically meaningful attention patterns, and only models that pass this test are included in the ensemble. Models that satisfy pre-determined consistency criteria are added to the final ensemble, and the other models are eliminated. The explanation consistency evaluation procedure, including similarity metrics, threshold criteria, and empirical justification, is detailed in <xref ref-type="sec" rid="s3_5">Section 3.5</xref>.</p>
</sec>
<sec id="s3_4_3">
<label>3.4.3</label>
<title>Explanation-Weighted Ensemble Construction</title>
<p>Following the explanation, consistency filtering, the retained models are not treated equally. Instead, their contributions to the final prediction are weighted according to the degree of Grad-CAM stability observed during the filtering stage.</p>
</sec>
<sec id="s3_4_4">
<label>3.4.4</label>
<title>Meta-Ensemble Layer Using XGBoost</title>
<p>To further refine the decision-making process, ECWMEv2 uses a meta-classifier that is trained on the logits of the retained models, allowing a refined prediction using model-level feature combination. The logit vectors of each backbone are obtained as raw and concatenated together and provided as features to an XGBoost classifier. The meta-layer is a non-linear learner of relationships between predictions made by a model and can be used to resolve conflicts between base models adaptively. It is trained on a stratified validation split on a binary logistic loss and L2 regularization. Such a two-level fusion, being an explanation-filtered soft voting and then a meta-classification, allows the framework to achieve high generalizability without the loss of explainability.</p>
</sec>
<sec id="s3_4_5">
<label>3.4.5</label>
<title>Training and Optimization Details</title>
<p>The backbone models are optimized with the help of uniform hyperparameters and data augmentation plans. The output layers associated with the model are retrained in a multiclass classification and are trained by the cross-entropy loss. The training is done on NVIDIA A100 GPUs with PyTorch and mixed-precision so as to speed up convergence. XGBoost meta-classifier is implemented using the validation split with the maximum depth of 3, the learning rate of 0.05, and the L2 regularization weight of 1.0. <xref ref-type="table" rid="table-8">Table 8</xref> lists the training setup in each of the models in the ECWMEv2 pool.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Training parameters for fine-tuning backbone models in ECWMEv2.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Input Image Size</td>
<td>224 &#x00D7; 224</td>
</tr>
<tr>
<td>Batch Size</td>
<td>16</td>
</tr>
<tr>
<td>Training and Testing Split</td>
<td>83:17</td>
</tr>
<tr>
<td>Learning Rate</td>
<td>0.0001</td>
</tr>
<tr>
<td>Optimizer</td>
<td>AdamW</td>
</tr>
<tr>
<td>Scheduler</td>
<td>Cosine Annealing</td>
</tr>
<tr>
<td>Number of Epochs</td>
<td>30</td>
</tr>
<tr>
<td>Loss Function</td>
<td>Cross-Entropy Loss</td>
</tr>
<tr>
<td>Early Stopping</td>
<td>Patience &#x003D; 5 epochs</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To reduce the risk of overfitting and improve model generalization, several strategies were employed during training. First, transfer learning was applied by initializing backbone networks with ImageNet-1K pretrained weights. Second, it was augmented with vast amounts of data by using rotation, flipping, blur, and noise injection techniques. Third, early stopping based on patience of five epochs was used to avoid excessive training when the validation performance stopped improving. Further to guarantee the robustness and minimize sampling bias, five-fold cross-validation was used in experimental assessment, and the dataset was split into five folds, with the model being trained and evaluated on alternate splits. In the case of the segmentation network, deep supervision has been added to enhance gradient propagation and minimize overfitting at intermediate layers of features.</p>
</sec>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Explainability Analysis and Consistency Evaluation Using Grad-CAM</title>
<sec id="s3_5_1">
<label>3.5.1</label>
<title>Explainability Integration within the Segmentation&#x2013;Classification Pipeline</title>
<p>The explainability aspect of the suggested framework is achieved utilizing Grad-CAM as a post hoc interpretability approach, as opposed to utilizing internal transformer attention maps employed directly. Although transformer self-attention offers implicit interactions between features, attention weights are not always spatially interpretable as saliency regions of medical images. Thus, Grad-CAM is used to produce spatially localized heatmaps, which show areas of the tumor that are used in classification predictions.</p>
<p>In contrast to the traditional methods, when Grad-CAM is applied to visualize the model once the inference is made, the offered framework incorporates the concept of explainability within the model selection and the model ensemble building. In particular, Grad-CAM maps obtained using a candidate model are checked on controlled input perturbations. Models that give unstable or inconsistent explanations are not included in the ensemble. Therefore, explainability is a clear reliability metric used to filter and weight models, and interpretability is an inseparable part of the decision-making pipeline and not a purely descriptive post hoc visualization. The saliency-based attribution with Grad-CAM is employed to create spatially consistent attention maps, which are appropriate to interpret the classification of brain tumors. Grad-CAM, in contrast to SHAP or LIME, which feature-level attribution, offers anatomically aligned heatmaps, which are required in medical imaging tasks where regional continuity is important.</p>
<p>For given class <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>c</mml:mi></mml:math></inline-formula>, Grad-CAM computes the importance weight of feature map <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>k</mml:mi></mml:math></inline-formula> in the final convolutional layer as:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>Z</mml:mi></mml:mfrac><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the score for class <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>c</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the <italic>k</italic>th feature map. <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>Z</mml:mi></mml:math></inline-formula> is the spatial normalization factor. The Grad-CAM heatmap is then computed as:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msup><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This produces a spatial saliency map indicating image regions contributing to the prediction. Each model in the classification ensemble produces Grad-CAM maps of clean and perturbed MRI inputs. Perturbations comprise: Rotation, Blur, and Noise injection simulating the clinically relevant variations of input. To measure the stability of the explanation, two measures are employed in the case of each pair of maps (clean <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and perturbed <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>):<list list-type="bullet">
<list-item>
<p><bold>Intersection over Union (IoU)</bold> is a measure of spatial overlap:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mtext>IoU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2229;</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x222A;</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x2265;</mml:mo><mml:mn>0.6</mml:mn></mml:math></disp-formula></p></list-item>
</list>
<list list-type="bullet">
<list-item>
<p><bold>Cosine Similarity</bold> is a measure of the directional consistency of attention distributions:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>Cosine</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x2265;</mml:mo><mml:mn>0.6</mml:mn></mml:math></disp-formula></p></list-item>
</list></p>
<p>The models that satisfy both thresholds (<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:mtext>IoU</mml:mtext></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>0.6</mml:mn></mml:math></inline-formula>, Cosine Similarity <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mo>&#x2265;</mml:mo><mml:mn>0.6</mml:mn></mml:math></inline-formula>) in 80% of the test samples are said to be explanation-consistent. The pass rate is calculated as in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>. This forms the Explanation Reliability Layer (ERL) in the ECWMEv2 framework, which filters out unstable models before ensemble construction.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>Pass Rate</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>Samples with Stable CAM</mml:mtext></mml:mrow><mml:mrow><mml:mtext>Total Test Samples</mml:mtext></mml:mrow></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mrow><mml:mtext>\%&#x00A0;</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>The retained models are not treated equally to ensure that there is fairness in the weighting of the models. Instead, each model is assigned a weight <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> b proportional to its explanation consistency, calculated in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mrow><mml:mtext>AvgIoU</mml:mtext></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mtext>AvgIoU</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:msub><mml:mrow><mml:mtext>AvgCos</mml:mtext></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mtext>AvgCos</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The final class probability is obtained via explanation-weighted soft voting expressed in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>. Here, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>Z</mml:mi></mml:math></inline-formula> ensures normalization and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x2223;</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the softmax probability from model <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>m</mml:mi></mml:math></inline-formula> for class <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>c</mml:mi></mml:math></inline-formula>.
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>ECWMEv</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x2223;</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>Z</mml:mi></mml:mfrac><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mtext>sclected</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x2223;</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_5_2">
<label>3.5.2</label>
<title>Threshold Selection and Trade-Offs</title>
<p>Thresholds of IoU &#x2265; 0.6 and Cosine Similarity &#x2265; 0.6 with a pass rate &#x2265; 80% were determined empirically to be able to compromise between strictness and diversity in the ensemble. <xref ref-type="table" rid="table-9">Table 9</xref> demonstrates the impact of stricter and stricter thresholds: extremely conservative thresholds (0.8 or higher) will weed out all stable models, whereas moderately balanced thresholds will weed out none. It is not aimed at imposing the same attention maps but providing consistency that is both anatomically significant and clinically interpretable.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Effect of increasing grad-CAM consistency thresholds on model retention.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Threshold Level</th>
<th>IoU Threshold</th>
<th>Cosine Similarity Threshold</th>
<th>Models Retained</th>
<th>Models Eliminated</th>
<th>Interpretation</th>
</tr>
</thead>
<tbody>
<tr>
<td>Balanced (60%)</td>
<td>&#x2265;0.60</td>
<td>&#x2265;0.60</td>
<td>ConvNeXt, Swin, EVA02</td>
<td>CoAtNet, MaxViT</td>
<td>Retains all stable explainers; filters noisy ones</td>
</tr>
<tr>
<td>Strict (70%)</td>
<td>&#x2265;0.70</td>
<td>&#x2265;0.70</td>
<td>Swin</td>
<td>ConvNeXt, EVA02, CoAtNet, MaxViT</td>
<td>Risks excluding explainable models due to minor instability</td>
</tr>
<tr>
<td>Stricter (80%)</td>
<td>&#x2265;0.80</td>
<td>&#x2265;0.80</td>
<td>None</td>
<td>All models</td>
<td>Too conservative; excludes even top-performing and consistent models</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_5_3">
<label>3.5.3</label>
<title>Results of Consistency Filtering</title>
<p>When all models were applied using the chosen thresholds, ConvNeXt, Swin Transformer, and EVA02 only had consistently stable and well-localized Grad-CAM heatmaps with tumor regions when perturbed. Conversely, CoAtNet as well as MaxViT were either diffuse or unstable and were eliminated in the final ensemble. The high consistency of the explanations by the average Grad-CAM IoU of 0.692 justifies the capability and transparency of the model. The quantitative analysis of the reliability of each model is provided in <xref ref-type="table" rid="table-10">Table 10</xref> to present the reliability of its explanation and retention.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Grad-CAM consistency metrics and model retention based on IoU &#x2265; 0.60 and cosine similarity &#x2265; 0.60.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Mean IoU &#x2265; 0.60</th>
<th>Cosine Similarity &#x2265; 0.60</th>
<th>Retained in ECWMEv2</th>
</tr>
</thead>
<tbody>
<tr>
<td>ConvNeXt</td>
<td>0.691</td>
<td>0.720</td>
<td>Yes</td>
</tr>
<tr>
<td>Swin Transformer</td>
<td>0.722</td>
<td>0.814</td>
<td>Yes</td>
</tr>
<tr>
<td>EVA02</td>
<td>0.664</td>
<td>0.713</td>
<td>Yes</td>
</tr>
<tr>
<td>CoAtNet</td>
<td>0.203</td>
<td>0.664</td>
<td>No</td>
</tr>
<tr>
<td>MaxViT</td>
<td>0.345</td>
<td>0.264</td>
<td>No</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The analysis has resulted in Grad-CAM being not just a visualization tool, but an important component of the ensemble in terms of filtering, weighting, and reliability of models. ECWMEv2 improves the readability, strength, and clinical credibility by ensuring that only models that are consistent and explanatory are used to make the predictions. The proposed explanation consistency filtering method is described in Algorithm 1.</p>
<fig id="fig-11">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-11.tif"/>
</fig>
</sec>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Performance Evaluation Metrics</title>
<p>The proposed framework is analyzed on three significant stages: tumor segmentation, tumor classification, and consistency of explanations (discussed in <xref ref-type="sec" rid="s3_5">Section 3.5</xref>). The test split of the BRISC2025 dataset is used in all the experiments according to <xref ref-type="table" rid="table-2">Table 2</xref>, and they are performed with subject-level separation to avoid data leakage between folds.</p>

<sec id="s3_6_1">
<label>3.6.1</label>
<title>Segmentation Evaluation</title>
<p>Evaluation of the segmentation model (Swin-DS-HAFUNetv2) is measured in terms of existing spatial overlap measures. The Dice coefficient is used to measure the overlap between the predicted segmentation <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>P</mml:mi></mml:math></inline-formula> and ground truth mask <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>G</mml:mi><mml:mo>,</mml:mo></mml:math></inline-formula> and it is given in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mtext>Dice Coefficient&#x00A0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref> is also used to compute the Intersection over Union or Jaccard Index:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mtext>IoU</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>To further analyze the model in terms of discriminative ability to divide into relevant areas, precision and recall are calculated using <xref ref-type="disp-formula" rid="eqn-12">Eqs. (12)</xref> and <xref ref-type="disp-formula" rid="eqn-13">(13)</xref>:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mtext>&#x00A0;Recall</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Here,<inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>,</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> denote true positives, false positives, and false negatives, respectively.</p>
</sec>
<sec id="s3_6_2">
<label>3.6.2</label>
<title>Classification Evaluation</title>
<p>The classification module is evaluated by measuring its accuracy and classes. The total classification precision is in <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mtext>Accuracy</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>To control the imbalance of classes, the F1-score is calculated as the harmonic mean of the precision and recall using <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>Precision&#x00A0;</mml:mtext></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>&#x00A0;Recall&#x00A0;</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>&#x00A0;Precision&#x00A0;</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>&#x00A0;Recall&#x00A0;</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results and Analysis</title>
<p>This section discusses in detail the performance of the explainability-integrated framework proposed, and the structure consists of Swin-DS-HAFUNetv2 to segment tumors and ECWMEv2 to classify. The assessments are in the form of evaluating qualitative and quantitative segmentation, ablation studies, consistency of explanations evaluation, and ultimate ensemble performance. The combination of the results proves the adequacy, clinical significance, and interpretability of the proposed framework.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Segmentation Results</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Qualitative Analysis of Segmentation Performance</title>
<p>To obtain a more in-depth understanding of the Swin-DS-HAFUNetv2 model, qualitative segmentation results on the validation set were viewed and clustered into good (&#x2265;0.90), borderline (0.75&#x2013;0.85), and bad (&#x2264;0.70) results as illustrated in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. This classification shows variability of segmentation performance and represents recurring patterns on clinical information.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Qualitative Investigation of the segmentation outcomes based on the offered Swin-DS-HAFUNetv2 model. (<bold>a</bold>) Good examples (High Dice &#x2265; 0.90): The prediction of masks is very close to ground truth, and there are few FP, and FN (<bold>b</bold>) Borderline examples (Medium Dice 0.75&#x2013;0.85): Prediction of masks is close to the ground truth, FP, and FN are moderate values, (<bold>c</bold>) Bad examples (Low Dice 0.70): There is a significant failure in segmentation, namely, either a missed tumor.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-7a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-7b.tif"/>
</fig>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Training Convergence and Stability across Folds</title>
<p>Mean IoU (mIoU) per epoch was used to monitor the training process and was performed on a four-fold cross-validation of segmentation. The model demonstrated a steady convergence of the model along the folds, as demonstrated in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, and optimality was reached at epoch 25; the output thereafter stabilized. The mean mIoU of the architecture was 0.865 &#x00B1; 0.16, which validates the consistency of the architecture on varying validation splits as well as the efficiency of deep supervision and integration of attention.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>mIoU vs. Epochs for all folds.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Quantitative Results of Cross-Validation and Statistical Testing</title>
<p><xref ref-type="table" rid="table-11">Table 11</xref> demonstrates the results of segmentation by cross-validation folds. The outcomes show a consistent performance, and the Dice scores are between 0.9780 and 0.9785. The model is relatively stable and reliably performed across all folds with an average precision of 0.9363 and a recall value of 0.9040 to 0.9091.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Cross-validation performance of the proposed Swin-DS-HAFUNetv2 model with mean, standard deviation, and 95% confidence interval to validate statistical reliability.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Fold</th>
<th>mIoU</th>
<th>Dice</th>
<th>Precision</th>
<th>Recall</th>
</tr>
</thead>
<tbody>
<tr>
<td>Fold 1</td>
<td>0.8657</td>
<td>0.9780</td>
<td>0.9358</td>
<td>0.9065</td>
</tr>
<tr>
<td>Fold 2</td>
<td>0.8671</td>
<td>0.9785</td>
<td>0.9358</td>
<td>0.9091</td>
</tr>
<tr>
<td>Fold 3</td>
<td>0.8640</td>
<td>0.9781</td>
<td>0.9372</td>
<td>0.9040</td>
</tr>
<tr>
<td>Mean &#x00B1; Std</td>
<td>0.8656 &#x00B1; 0.0016</td>
<td>0.9782 &#x00B1; 0.0003</td>
<td>0.9363 &#x00B1; 0.001</td>
<td>0.9065 &#x00B1; 0.0025</td>
</tr>
<tr>
<td>95% Confidence Interval (CI)</td>
<td>0.8638&#x2013;0.8674</td>
<td>0.9779&#x2013;0.9785</td>
<td>0.9352&#x2013;0.9374</td>
<td>0.9037&#x2013;0.9093</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Further statistical reliability testing was done on the proposed model, Swin-DS-HAFUNetv2, through cross-validation analysis of the confidence interval. The mean of the Dice score was 0.9782, and the standard deviation was 0.0003; hence, the 95% confidence interval of the score was 0.9779&#x2013;0.9785. The small range suggests that the performance of segmentation is close to the same value in the various validation splits, which proves the performance gains are statistically constant and not simply by chance.</p>
</sec>
<sec id="s4_1_4">
<label>4.1.4</label>
<title>Ablation Study of the Segmentation Pipeline</title>
<p>To quantify the individual contributions of the E-HAF and the E-CBE modules, a systematic ablation study was conducted, as summarized in <xref ref-type="table" rid="table-12">Table 12</xref>. The baseline Swin DS HAFUNetv2 model without either component achieved a weighted mIoU of 85.2%. The usage of E-HAF in isolation enhanced multi-scale feature integration, achieving consistent improvement in all tumor types and weighted mIoU to 86.5%. The added E-CBE further contributed to better contextual modeling of the bottleneck at the global scale, the weighted mIoU increased to 87.1%, and induced significant positive changes for glioma and meningioma cases. The overall performance of the model was the greatest when both E-HAF and E-CBE were used, and the glioma group had the highest increase. These findings indicate that E-HAF and E-CBE enhance each other in terms of benefits, and when they are combined, they provide a significant advantage to tumor boundary demarcation and contextual insight, especially of more complex and heterogeneous tumor structures. The results confirm the idea that E-HAF enhances the feature fusion process of different scales, and E-CBE enhances the bottleneck presentation. Together, they produce significant mIoU improvements, especially on the more complicated glioma subtype.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Ablation study of swin-DS-HAFUNetv2 on tumor types.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Baseline</th>
<th>E-HAF</th>
<th>E-CBE</th>
<th>mIoU (Glioma)</th>
<th>mIoU (Meningioma)</th>
<th>mIoU (Pituitary)</th>
<th>Weighted mIoU (Mean &#x00B1; Std)</th>
</tr>
</thead>
<tbody>
<tr>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>83.2</td>
<td>86.7</td>
<td>85.9</td>
<td>85.2 &#x00B1; 0.8</td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>84.8</td>
<td>87.6</td>
<td>87.1</td>
<td>86.5 &#x00B1; 0.7</td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>85.4</td>
<td>88.1</td>
<td>87.7</td>
<td>87.1 &#x00B1; 0.6</td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>86.5</td>
<td>89.4</td>
<td>88.7</td>
<td>88.2 &#x00B1; 0.5</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_1_5">
<label>4.1.5</label>
<title>Comparative Evaluation with Existing Models</title>
<p>The proposed Swin-DS-HAFUNetv2 model was tested and compared to some of the current segmentation models. Recent transformer-based segmentation models such as TransUNet, Swin-UNet, and UNETR have demonstrated strong performance in medical image segmentation tasks [<xref ref-type="bibr" rid="ref-66">66</xref>]. However, many of these architectures are primarily designed for volumetric multimodal MRI datasets such as BraTS. Since the BRISC2025 dataset contains single-slice T1-weighted images, direct comparison with these models requires substantial architectural adaptation. As shown in <xref ref-type="table" rid="table-13">Table 13</xref>, the proposed method significantly outperforms existing studies, particularly in glioma, meningioma, and pituitary cases.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Quantitative comparison of the proposed Swin-DS-HAFUNetv2 with representative CNN, attention-based, and transformer-based segmentation architectures.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Ref.</th>
<th>Model</th>
<th>Glioma</th>
<th>Meningioma</th>
<th>Pituitary</th>
<th>Weighted mIoU</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-53">53</xref>]</td>
<td>UNet</td>
<td>69.7</td>
<td>77.1</td>
<td>79.3</td>
<td>75.7</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
<td>UNet&#x002B;&#x002B;</td>
<td>71.7</td>
<td>74.2</td>
<td>79.7</td>
<td>75.3</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-52">52</xref>]</td>
<td>ABANet</td>
<td>72.4</td>
<td>80.4</td>
<td>84.7</td>
<td>79.5</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>Swin-HAFUNet</td>
<td>76.0</td>
<td>85.0</td>
<td>85.3</td>
<td>82.4</td>
</tr>
<tr>
<td>Proposed</td>
<td>Swin-DS-HAFUNetv2</td>
<td>86.5</td>
<td>89.4</td>
<td>88.7</td>
<td>88.2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Classification Results</title>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Single-Model Performance</title>
<p>Each backbone model was trained on the classification task. The results in <xref ref-type="table" rid="table-14">Table 14</xref> indicate that Swin Transformer and ConvNeXt outperformed the rest in terms of both predictive and explanation consistency.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Single-model classification performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Accuracy</th>
<th>Macro-F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>CoAtNet</td>
<td>0.993</td>
<td>0.9931</td>
</tr>
<tr>
<td>ConvNeXt</td>
<td>0.987</td>
<td>0.9877</td>
</tr>
<tr>
<td>Swin</td>
<td>0.992</td>
<td>0.9922</td>
</tr>
<tr>
<td>EVA02</td>
<td>0.817</td>
<td>0.8212</td>
</tr>
<tr>
<td>MaxViT</td>
<td>0.980</td>
<td>0.9828</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Grad-CAM Explanation Consistency Analysis&#x2014;Perturbation Stability</title>
<p>Explanation consistency was assessed under noise, rotation, and contrast jitter. Swin, ConvNeXt, and EVA02 passed the minimum threshold (IoU &#x2265; 0.6, cosine similarity &#x2265; 0.6, pass rate &#x2265; 0.8), as shown in <xref ref-type="table" rid="table-10">Tables 10</xref> and <xref ref-type="table" rid="table-15">15</xref>. Qualitative overlays of Grad-CAM on tumor regions of a sample data in <xref ref-type="fig" rid="fig-9">Fig. 9</xref> demonstrate moderate overlap (Dice &#x003D; 0.587, IoU &#x003D; 0.416), confirming alignment between model attention and pathology.</p>
<table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Grad-CAM stability metrics; models passing consistency thresholds were retained.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Retained in ECWMEv2 (IoU &#x2265; 0.6, Cosine Similarity &#x2265; 0.6)</th>
<th>Pass Rate &#x2265; 0.80</th>
<th>Status</th>
<th>Weight <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>ConvNeXt</td>
<td>Yes</td>
<td>0.812</td>
<td>Pass</td>
<td>0.304</td>
</tr>
<tr>
<td>Swin Transformer</td>
<td>Yes</td>
<td>0.861</td>
<td>Pass</td>
<td>0.437</td>
</tr>
<tr>
<td>EVA02</td>
<td>Yes</td>
<td>0.812</td>
<td>Pass</td>
<td>0.259</td>
</tr>
<tr>
<td>CoAtNet</td>
<td>No</td>
<td>0.002</td>
<td>Fail</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>MaxViT</td>
<td>No</td>
<td>0.010</td>
<td>Fail</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Visual and quantitative assessment of Grad-CAM alignment with tumor pathology, showing overlap between attention maps and ground-truth tumor regions under clean inputs: (<bold>a</bold>) original input MRI slice, (<bold>b</bold>) resized binary tumor ground truth mask, (<bold>c</bold>) tumor mask overlaid in red, (<bold>d</bold>) Grad-CAM heatmap contour overlaid on the MRI and tumor mask, and (<bold>e</bold>) full Grad-CAM heatmap with IoU and Dice values. The moderate overlap (IoU &#x003D; 0.416, Dice &#x003D; 0.587) indicates partial alignment between model attention and tumor region, validating the clinical relevance of explanations.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-9.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-10">Fig. 10</xref> contains Grad-CAM visualization with varying perturbations of inputs, which clearly shows that some models are more stable to explanations. It is also evident that Swin Transformer and ConvNeXt always have well-localized and strong attention maps, which do not change in brightness, noise, and rotation perturbation, meaning that they can consistently focus on tumor localities. EVA02 will show relatively wider and noisier activations, but they still are semantically meaningful and would yield valuable global contextual data. CoAtNet, in contrast, exhibits fragmented and dispersed attention, which is highly sensitive to perturbations, whereas MaxViT generates weak or diffuse activations that are of low anatomical relevance. These findings show that Swin Transformer, ConvNeXt, and EVA02 are the only ones that meet the quantitative explanation-consistency criteria (IoU &#x2265; 0.60, cosine similarity &#x2265; 0.60, and pass rate &#x2265; 0.80) and are thus included in the final ensemble.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Grad-CAM visualizations under various input perturbations for different models.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80241-fig-10.tif"/>
</fig>
</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Explanation-Consistent Meta-Ensemble (ECWMEv2)</title>
<p>The ECWMEv2 framework proposed was compared to the explanation-weighted soft voting and single classifiers. As shown in <xref ref-type="table" rid="table-16">Table 16</xref>, ECWMEv2 employing an XGBoost meta-learner achieved superior class-balanced performance, attaining a macro-F1 score of 0.9867 with an accuracy of 0.9917. In comparison, explanation-weighted soft voting achieved lower accuracy (0.9863) and lower macro-F1 (0.9734), indicating reduced robustness in handling inter-class variability.</p>
<table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Performance comparison between explanation-weighted soft voting and the proposed ECWMEv2 ensemble with XGBoost meta-learning.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Accuracy</th>
<th>Macro-F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>Explanation-Weighted (Soft Voting)</td>
<td>0.9863 &#x00B1; 0.0031</td>
<td>0.9734 &#x00B1; 0.0062</td>
</tr>
<tr>
<td>ECWMEv2 (XGBoost)</td>
<td>0.9917 &#x00B1; 0.0046</td>
<td>0.9885 &#x00B1; 0.0075</td>
</tr>
<tr>
<td>95% CI</td>
<td>0.9877&#x2013;0.9957</td>
<td>0.9819&#x2013;0.9951</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To verify that the performance improvement achieved by ECWMEv2 is statistically meaningful, statistical significance testing was conducted using five-fold cross-validation results. The proposed ECWMEv2 ensemble achieved an average accuracy of 0.9917 &#x00B1; 0.0046, compared with 0.9863 &#x00B1; 0.0031 for explanation-weighted soft voting. Similarly, the Macro-F1 score improved from 0.9734 &#x00B1; 0.0062 to 0.9885 &#x00B1; 0.0075. The stability of ECWMEv2 is also attested by the five-fold results of the cross-validation in <xref ref-type="table" rid="table-17">Table 17</xref>, with accuracy ranging between 0.9897 and 1.0000, and F1 scores ranging between 0.9789 and 1.0000. The mean of folds (accuracy &#x003D; 0.9917, macro-F1 &#x003D; 0.9867) indicates a high generalization and a low variance, which proves that the model selection based on explanation consistency is part of the consistent and high accuracy of classification.</p>
<table-wrap id="table-17">
<label>Table 17</label>
<caption>
<title>Five-fold cross-validation results for ECWMEv2 with statistical reliability indicators.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Fold</th>
<th>Accuracy</th>
<th>Macro-F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>Fold 1</td>
<td>0.9897</td>
<td>0.9867</td>
</tr>
<tr>
<td>Fold 2</td>
<td>0.9899</td>
<td>0.9789</td>
</tr>
<tr>
<td>Fold 3</td>
<td>0.9898</td>
<td>0.9888</td>
</tr>
<tr>
<td>Fold 4</td>
<td>1.0000</td>
<td>1.0000</td>
</tr>
<tr>
<td>Fold 5</td>
<td>0.9899</td>
<td>0.9879</td>
</tr>
<tr>
<td>Mean &#x00B1; Std</td>
<td>0.9917 &#x00B1; 0.0046</td>
<td>0.9885 &#x00B1; 0.0075</td>
</tr>
<tr>
<td>95% CI</td>
<td>0.9877&#x2013;0.9957</td>
<td>0.9819&#x2013;0.9951</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2_4">
<label>4.2.4</label>
<title>Ablation Study of the Classification Pipeline</title>
<p><xref ref-type="table" rid="table-18">Table 18</xref> summarizes the effect of each component in the ECWMEv2 classification framework. Using the best-performing single model (Swin Transformer) as a baseline, we observed an accuracy of 98.71% and a macro-F1 of 98.12%, with high Grad-CAM localization consistency (Mean IoU &#x003D; 0.722). When the Explanation Reliability Layer (ERL) was removed, and all models were included without Grad-CAM filtering, performance dropped to 98.42% accuracy and 97.81% macro-F1, with a significantly lower explanation quality (Mean IoU &#x003D; 0.461) due to the inclusion of unstable models like CoAtNet and MaxViT.</p>
<table-wrap id="table-18">
<label>Table 18</label>
<caption>
<title>Ablation of the classification pipeline.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Configuration</th>
<th>Accuracy</th>
<th>Macro-F1</th>
<th>Mean IoU (Grad-CAM)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Best single model (Swin Transformer)</td>
<td>0.9871</td>
<td>0.9812</td>
<td>0.722</td>
</tr>
<tr>
<td>Without ERL (all models included)</td>
<td>0.9842</td>
<td>0.9781</td>
<td>0.461</td>
</tr>
<tr>
<td>Without Explanation Weights (uniform voting)</td>
<td>0.9883</td>
<td>0.9829</td>
<td>0.672</td>
</tr>
<tr>
<td>Full ECWMEv2 (ERL &#x002B; Explanation-weighted voting)</td>
<td>0.9917</td>
<td>0.9867</td>
<td>0.692</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The fusion of explanation-based and uniform voting to the filtered models resulted in 98.83% accuracy and 98.29% macro-F1, indicating that explanation-weighted fusion enhances calibration and class balance. The overall performance is the best with ECWMEv2 architecture, both Grad-CAM consistency filtering (ERL) and explanation-weighted soft voting, with the final accuracy of 99.17, the macro-F1 of 98.67, and the Mean IoU of 0.692. These results confirm that the model selection that reduces problems with an explanation and focuses on the ensemble fusion helps to increase the reliability of prediction and the interpretability of the model in brain tumor classification.</p>
<p>Besides predictive performance, the computational complexity of the proposed framework was evaluated based on model parameters, training time, and inference latency.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Computational Complexity of Proposed Framework</title>
<p><xref ref-type="table" rid="table-19">Table 19</xref> provides the computational requirements of the proposed framework. Swin-DS-HAFUNetv2, the segmentation backbone consists of around 63 million parameters and takes 2.5&#x2013;3 h to train on one fold, the amount of computation needed to run a transformer-based architecture. The classification backbones (Swin, ConvNeXt, and EVA02) have 88 to 354 M parameters, and they can be trained in a matter of 1.5&#x2013;2 h per model. Such training costs imply that model development is computationally expensive and thus only centralized computing environments with GPUs are suitable.</p>
<table-wrap id="table-19">
<label>Table 19</label>
<caption>
<title>Computational complexity and inference latency of the proposed segmentation, classification, and ensemble components.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Component</th>
<th>Model(s)</th>
<th>Parameters (Millions)</th>
<th>Training Time</th>
<th>Inference Time</th>
</tr>
</thead>
<tbody>
<tr>
<td>Segmentation</td>
<td>Swin-DS-HAFUNetv2</td>
<td>63 M</td>
<td>2.5&#x2013;3 h/fold</td>
<td>&#x007E;35 ms/image</td>
</tr>
<tr>
<td>Classification</td>
<td>Swin, ConvNeXt, EVA02</td>
<td>88&#x2013;354 M</td>
<td>1.5&#x2013;2 h/model</td>
<td>&#x007E;25&#x2013;40 ms/image</td>
</tr>
<tr>
<td>Ensemble</td>
<td>XGBoost</td>
<td>&#x2013;</td>
<td>&#x003C;1 min</td>
<td>&#x007E;3 ms/image</td>
</tr>
<tr>
<td>Total Inference</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td>&#x007E;70&#x2013;90 ms/image</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Although the cost of training is still relatively high, the inference can be obtained with the minimum latency, which is essential in terms of practical implementation. The segmentation module takes 35 to 35 ms/image, and the classification models take 25&#x2013;40 ms/image. The ensemble module based on XGBoost is the one adding almost no overhead, as it takes less than one minute to train and about 3 ms per prediction.</p>
<p>Overall, the total time to process an image with the end-to-end pipeline is about 70&#x2013;90 ms, which indicates that segmentation, classification, and explainability-aware ensemble learning do not have prohibitive inference costs. These findings indicate that, despite the framework not being optimized to accommodate ultra-low-resource or edge devices, it can still be used in offline analysis or near-real-time analysis clinical research processes, where robustness and interpretability are desired over the use of a minimal computational footprint.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Comparative Analysis with Existing Studies</title>
<p><xref ref-type="table" rid="table-20">Table 20</xref> is a quantitative comparison of the proposed framework and the recent research on the segmentation and classification of brain tumors. The existing hybrid segmentation-classification models based on convolutional architectures typically obtain segmentation Dice scores of about 0.91 and classification accuracy 98% and are therefore limited to detecting complex tumor margins and overall contextual information [<xref ref-type="bibr" rid="ref-64">64</xref>]. They achieve predictive accuracy of the transfer-learning-based classification methods up to 98.49; however, only classification issues are addressed, and no explicit tumor localization and explanation-sensitive evaluation methods are applied [<xref ref-type="bibr" rid="ref-56">56</xref>]. Hybrid CNN Transformer classifiers extend the representation of features and also report accuracy levels of classification in the range of 96&#x2013;98, but do not incorporate stages of segmentation and do not quantitatively evaluate the stability of explanation [<xref ref-type="bibr" rid="ref-25">25</xref>]. Multi-class MRI tumor classification with ensemble CNN models, including InceptionV3 &#x002B; Xception, has a validation accuracy of about 98.30%, yet does not support segmentation or interpretability, which is essential when using in practice [<xref ref-type="bibr" rid="ref-58">58</xref>]. Transformer-based approaches analyzed using the BRISC dataset are more robust in terms of baseline, segmentation Dice scores of 0.94&#x2013;0.95, and classification accuracy of roughly 98%, but explainability in these works is viewed largely as a post hoc visualization tool and is not considered in model selection or ensemble building [<xref ref-type="bibr" rid="ref-20">20</xref>]. Comparative studies of CNN and hybrid Transformer models validate the effectiveness of Transformer-based classifiers and their accuracy of 98.2, being restricted to classification-only pipelines [<xref ref-type="bibr" rid="ref-59">59</xref>]. Explainable deep learning frameworks using NASNet with Grad-CAM and LIME achieve about 92.98% classification accuracy, improving transparency, but remain limited to binary classification without segmentation or quantitative explainability evaluation [<xref ref-type="bibr" rid="ref-42">42</xref>]. Hybrid explainable models that combine SHAP or Grad-CAM achieve classification rates of between 97% and 98%; however, explanation robustness to perturbations is not considered, and interpretability does not affect ensemble decisions [<xref ref-type="bibr" rid="ref-35">35</xref>]. The proposed Swin-DS-HAFUNetv2 &#x002B; ECWMEv2 framework also incorporates transformer-based segmentation, explanation-consistent ensemble classification, and quantitative explainability validation into one pipeline. The segmentation stage has a mean Dice score of 0.9782, which is better compared to the current CNN and Transformer-based segmentation approaches. To be classified, the ECWMEv2 ensemble with an XGBoost meta-learner has an average five-fold accuracy of 0.9917 and a macro-F1 score of 0.9867, which is better than the results of individual classifiers and the weighted voting using explanations. Notably, the consistency of the explanation is applied using perturbation-based Grad-CAM analysis, which results in a Grad-CAM IoU of 0.692, which allows the model&#x2019;s attention to be stable and clinically significant.</p>
<table-wrap id="table-20">
<label>Table 20</label>
<caption>
<title>Quantitative comparison with existing studies.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Ref.</th>
<th>Model</th>
<th>Segmentation Performance</th>
<th>Classification Performance</th>
<th>Explainability (for Visualization)</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
<td>Ensemble CNN (InceptionV3 &#x002B; Xception)</td>
<td>&#x2013;</td>
<td>Accuracy &#x003D; 98.30%</td>
<td>None</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>NASNet large &#x002B; XAI framework</td>
<td>&#x2013;</td>
<td>Accuracy &#x003D; 92.98%</td>
<td>LIME &#x002B; Grad-CAM</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td>Modified SegNet &#x002B; Hybrid DL (texture-enhanced)</td>
<td>Dice &#x2248; 0.91</td>
<td>Accuracy &#x2248; 97.6%</td>
<td>None</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>Transfer-learning CNNs (ResNet, DenseNet, EfficientNet)</td>
<td>&#x2013;</td>
<td>Accuracy &#x003D; 98.49%</td>
<td>None</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>Hybrid CNN&#x2013;Transformer (T1-weighted MRI)</td>
<td>&#x2013;</td>
<td>Accuracy &#x2248; 96%&#x2013;98%</td>
<td>None</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>Swin-HAFUNet</td>
<td>Dice &#x2248; 0.94&#x2013;0.95</td>
<td>Accuracy &#x2248; 98.0%</td>
<td>Visual only (Grad-CAM)</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>CNN vs. Hybrid Transformer Models</td>
<td>&#x2013;</td>
<td>Accuracy &#x2248; 98.2%</td>
<td>None</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Hybrid DL &#x002B; ML with SHAP</td>
<td>&#x2013;</td>
<td>Accuracy &#x2248; 97%&#x2013;98%</td>
<td>Post hoc</td>
</tr>
<tr>
<td>Proposed</td>
<td>Swin-DS-HAFUNetv2 &#x002B; ECWMEv2 (XGBoost)</td>
<td>Dice &#x003D; 0.9782</td>
<td>Accuracy &#x003D; 0.9917 (5-fold avg) Macro-F1 &#x003D; 0.9867</td>
<td>Integrated &#x0026; quantified (Grad-CAM IoU &#x003D; 0.692)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions and Future Directions</title>
<p>The research introduces a unified explainability-aware system of automated brain tumor analysis through the combination of transformer-based segmentation and explanation-consistent ensemble classification. The suggested Swin-DS-HAFUNetv2 showed the high accuracy of localization of the tumor boundaries through the combination of hierarchical self-attention, contextual bottleneck modeling, and deep supervision, leading to the accurate segmentation of various types of tumors. The framework, in combination with segmentation and downstream classification, was thus able to minimize bias in the background and, additionally, to make predictive decisions based on clinically relevant tumor regions and not spurious image characteristics. Besides, the ECWMEv2 classification framework added to the design group consistency of explanation as its core principle of ensemble construction. The framework attained 99.17% of classification accuracy and stability in response to various perturbations and imposed a metric of reliability on the explanations generated by the model selection and weighting process, leading to the development of a stable and interpretable decision-making process. In contrast to the traditional approach, where explainability is regarded as post hoc visualization, the research incorporates interpretability into the architectural design and optimization of the ensemble framework, which increases robustness, transparency, and clinical trustworthiness. The proposed framework demonstrates that predictive performance can be traded off with explanation stability as a promising way to generate reliable and understandable artificial intelligence systems that are reliable at a benchmark scale and can be applied as a general methodology in future clinical-scale studies.</p>
<p>The explainability-aware framework proposed presents a number of viable benefits to automated analysis of brain tumors. The framework can be used to combine transformer-based segmentation with explanation-consistent ensemble classification, allowing more precise localization of tumors and also predicting their type in a single pipeline. Grad-CAM stability as a model selection criterion will make sure that classification decisions are based on coherent and anatomically significant visual explanations, enhancing the openness of the decision-making process. Moreover, the models can improve the segmentation followed by classification, which enables the models to target tumor regions and not background tissues to minimize spurious activations and enhance the reliability of the diagnosis. These features render the framework appropriate to research-based clinical decision support systems, in which interpretability and stability of prediction are critical.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Research Limitations</title>
<p>Although the proposed explainability-aware framework demonstrates promising performance, the following limitations should be addressed in future work. The experiments were initially conducted on 2D single-slice T1-weighted MRI images, as the BRISC2025 dataset only provides this type of data. Consequently, the framework currently does not leverage multimodal MRI data, such as T2 or FLAIR images, which are widely used in clinical neuro-oncology to further characterize tumors. Additionally, the proposed method is based on 2D MRI slices rather than 3D volumetric scans, and therefore lacks the ability to capture inter-slice contextual information&#x2014;an important factor in achieving precise clinical diagnoses.</p>
<p>The experimental evaluation was performed using only one benchmark dataset, with no external validation conducted on other publicly available datasets such as BraTS, TCIA, or REMBRANDT. This limitation may hinder the generalization of the proposed framework to datasets with different imaging characteristics or tumor patterns.</p>
<p>Moreover, multi-center validation involving various scanners, institutions, or acquisition protocols was not carried out. Variations in MRI hardware, acquisition parameters, and imaging protocols in real-world clinical settings can lead to domain shifts, potentially degrading the performance of models trained on a specific dataset when applied to different clinical environments.</p>
<p>In addition, potential biases may exist in the BRISC2025 dataset due to the distribution of tumor types, imaging orientations, or patient demographics. If certain tumor features or imaging characteristics are over-represented, the trained models may learn dataset-specific patterns rather than generalizable diagnostic features. Although cross-validation and data augmentation were employed to mitigate overfitting, further testing on diverse clinical data is necessary to ensure fairness and robustness across different patient populations.</p>
<p>Furthermore, the proposed framework is built upon transformer-based architectures and ensemble learning, which entail considerable computational complexity and require a GPU-enabled training environment. While inference latency remains manageable, the training demands may limit deployment in resource-constrained settings, such as edge-based medical systems.</p>
<p>Finally, although Grad-CAM explanations were used to assess explanation stability and select the optimal ensemble, saliency-based explanation methods can be sensitive to model architecture and input perturbations. Therefore, additional approaches to complement interpretability techniques and uncertainty-aware explanations should be explored in the future to further assess potential biases and improve the reliability of visual explanations in clinical decision-support systems.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Future Research Directions</title>
<p>This research can be further advanced in the following directions. First, the proposed framework can be extended to accommodate multimodal and volumetric 3D MRI data, including T1, T2, and FLAIR modalities, to improve tumor characterization and localization. Second, incorporating uncertainty-aware learning could provide a more objective estimate of prediction confidence and help address ambiguous or borderline cases in clinical imaging. Third, model compression techniques such as pruning, knowledge distillation, or quantization can be employed to enhance computational efficiency, facilitating the integration of the system into clinical practice. Finally, multicenter validation studies using a diverse range of MRI data will be necessary to evaluate the robustness and clinical utility of the proposed explainability-aware system.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>We would like to express our sincere gratitude to all individuals.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Mamoona Jabbar, Uzma Jamil; data collection: Mamoona Jabbar; analysis and interpretation of results: Mamoona Jabbar, Uzma Jamil; draft manuscript preparation: Mamoona Jabbar, Bushra Zafar; review and editing: Mamoona Jabbar, Muhammad Younas; funding acquisition: Mamoona Jabbar. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The BRISC dataset is available at Kaggle (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/briscdataset/brisc2025/">https://www.kaggle.com/datasets/briscdataset/brisc2025/</ext-link>), Figshare (<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.6084/m9.figshare.30533120">https://doi.org/10.6084/m9.figshare.30533120</ext-link>), and Zenodo (<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.17524350">https://doi.org/10.5281/zenodo.17524350</ext-link>).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Anitha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nair</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kamaraj</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Brain tumor detection and classification using deep neural network and interpretation using XAI techniques</article-title>. In: <conf-name>Proceedings of the 2024 IEEE International Conference on Signal Processing, Informatics, Communication and Energy Systems (SPICES); 2024 Sep 20&#x2013;22</conf-name>; <publisher-loc>Kottayam, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/SPICES62143.2024.10779829</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Agrawal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chaki</surname> <given-names>J</given-names></string-name></person-group>. <article-title>CerebralNet meets explainable AI: brain tumor detection and classification with probabilistic augmentation and a deep learning approach</article-title>. <source>Biomed Signal Process Control</source>. <year>2025</year>;<volume>110</volume>(<issue>1</issue>):<fpage>108210</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2025.108210</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alqhtani</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Soomro</surname> <given-names>TA</given-names></string-name>, <string-name><surname>Ali Shah</surname> <given-names>A</given-names></string-name>, <string-name><surname>Aziz Memon</surname> <given-names>A</given-names></string-name>, <string-name><surname>Irfan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Improved brain tumor segmentation and classification in brain MRI with FCM-SVM: a diagnostic approach</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>4</issue>):<fpage>61312</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3394541</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alqhtani</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Soomro</surname> <given-names>TA</given-names></string-name>, <string-name><surname>Bin Ubaid</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>A</given-names></string-name>, <string-name><surname>Irfan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Asiri</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>Contrast normalization strategies in brain tumor imaging: from preprocessing to classification</article-title>. <source>Comput Model Eng Sci</source>. <year>2024</year>;<volume>140</volume>(<issue>2</issue>):<fpage>1539</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2024.051475</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zahoor</surname> <given-names>A</given-names></string-name>, <string-name><surname>Irfan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Usman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Haider</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Brain tumor detection in magnetic resonance images using swin transformer</article-title>. <source>Conclus Med</source>. <year>2025</year>;<volume>1</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.71107/kx24gt94</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ariful Islam</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mridha</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Safran</surname> <given-names>M</given-names></string-name>, <string-name><surname>Alfarhood</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mohsin Kabir</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Revolutionizing brain tumor detection using explainable AI in MRI images</article-title>. <source>NMR Biomed</source>. <year>2025</year>;<volume>38</volume>(<issue>3</issue>):<fpage>e70001</fpage>. doi:<pub-id pub-id-type="doi">10.1002/nbm.70001</pub-id>; <pub-id pub-id-type="pmid">39948696</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Asif</surname> <given-names>IH</given-names></string-name>, <string-name><surname>Bin Haque</surname> <given-names>S</given-names></string-name>, <string-name><surname>Nawaz</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MM</given-names></string-name></person-group>. <article-title>Brain tumor classification using deep learning with explainable AI (XAI)</article-title>. In: <conf-name>Proceedings of the 2025 8th International Conference on Electronics, Materials Engineering &#x0026; Nano-Technology (IEMENTech); 2025 Jan 31&#x2013;Feb 2</conf-name>; <publisher-loc>Kolkata, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iementech65115.2025.10959473</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Asiri</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Soomro</surname> <given-names>TA</given-names></string-name>, <string-name><surname>Ali Shah</surname> <given-names>A</given-names></string-name>, <string-name><surname>Pogrebna</surname> <given-names>G</given-names></string-name>, <string-name><surname>Irfan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Alqahtani</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Optimized brain tumor detection: a dual-module approach for MRI image enhancement and tumor classification</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>(<issue>1</issue>):<fpage>42868</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3379136</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Azeez</surname> <given-names>O</given-names></string-name>, <string-name><surname>Abdulazeez</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Classification of brain tumor based on machine learning algorithms: a review</article-title>. <source>J Appl Sci Technol Trends</source>. <year>2025</year>;<volume>6</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.38094/jastt61188</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Balamurugan</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gnanamanoharan</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Brain tumor segmentation and classification using hybrid deep CNN with LuNetClassifier</article-title>. <source>Neural Comput Appl</source>. <year>2023</year>;<volume>35</volume>(<issue>6</issue>):<fpage>4739</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-022-07934-7</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhaskaran</surname> <given-names>SB</given-names></string-name>, <string-name><surname>Datta</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Explainability of brain tumor classification model based on InceptionV3 using XAI tools</article-title>. <source>J Flow Vis Image Process</source>. <year>2024</year>;<volume>32</volume>(<issue>2</issue>):<fpage>35</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1615/jflowvisimageproc.2024054026</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Biratu</surname> <given-names>ES</given-names></string-name>, <string-name><surname>Schwenker</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ayano</surname> <given-names>YM</given-names></string-name>, <string-name><surname>Debelee</surname> <given-names>TG</given-names></string-name></person-group>. <article-title>A survey of brain tumor segmentation and classification algorithms</article-title>. <source>J Imaging</source>. <year>2021</year>;<volume>7</volume>(<issue>9</issue>):<fpage>179</fpage>. doi:<pub-id pub-id-type="doi">10.3390/jimaging7090179</pub-id>; <pub-id pub-id-type="pmid">34564105</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><article-title>Brain Tumor Classification (MRI)</article-title>. <comment>[cited 2026 Jan 1]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/sartajbhuvaji/brain-tumor-classification-mri">https://www.kaggle.com/datasets/sartajbhuvaji/brain-tumor-classification-mri</ext-link>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jetlin</surname> <given-names>CP</given-names></string-name>, <string-name><surname>Sherly Puspha Annabel</surname> <given-names>L</given-names></string-name></person-group>. <article-title>PyQDCNN: pyramid QDCNNet for multi-level brain tumor classification using MRI image</article-title>. <source>Biomed Signal Process Control</source>. <year>2025</year>;<volume>100</volume>(<issue>3</issue>):<fpage>107042</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2024.107042</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Charaabi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sayari</surname> <given-names>A</given-names></string-name>, <string-name><surname>El Hamdi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Njah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ben Slima</surname> <given-names>M</given-names></string-name></person-group>. <article-title>An XAI-infused multiclass MRI brain tumor classification using deep transfert learning (DTL)</article-title>. In: <conf-name>Proceedings of the 2024 10th International Conference on Control, Decision and Information Technologies (CoDIT); 2024 Jul 1&#x2013;4</conf-name>; <publisher-loc>Vallette, Malta</publisher-loc>. p. <fpage>1044</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/codit62066.2024.10708599</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Brain tumor dataset</article-title>. <source>Figshare</source>. <year>2017</year>. doi:<pub-id pub-id-type="doi">10.6084/m9.figshare.1512427</pub-id>; <pub-id pub-id-type="pmid">39653243</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deepa</surname> <given-names>S</given-names></string-name>, <string-name><surname>Janet</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sumathi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ananth</surname> <given-names>JP</given-names></string-name></person-group>. <article-title>Hybrid optimization algorithm enabled deep learning approach brain tumor segmentation and classification using MRI</article-title>. <source>J Digit Imag</source>. <year>2023</year>;<volume>36</volume>(<issue>3</issue>):<fpage>847</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10278-022-00752-2</pub-id>; <pub-id pub-id-type="pmid">36622465</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Elhadidy</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Elgohr</surname> <given-names>AT</given-names></string-name>, <string-name><surname>El-geneedy</surname> <given-names>M</given-names></string-name>, <string-name><surname>Akram</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kasem</surname> <given-names>HM</given-names></string-name></person-group>. <article-title>Comparative analysis for accurate multi-classification of brain tumor based on significant deep learning models</article-title>. <source>Comput Biol Med</source>. <year>2025</year>;<volume>188</volume>(<issue>1</issue>):<fpage>109872</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiomed.2025.109872</pub-id>; <pub-id pub-id-type="pmid">39970824</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ennab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mcheick</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Advancing AI interpretability in medical imaging: a comparative analysis of pixel-level interpretability and grad-CAM models</article-title>. <source>Mach Learn Knowl Extr</source>. <year>2025</year>;<volume>7</volume>(<issue>1</issue>):<fpage>12</fpage>. doi:<pub-id pub-id-type="doi">10.3390/make7010012</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fateh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rezvani</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Moayedi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rezvani</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fateh</surname> <given-names>F</given-names></string-name>, <string-name><surname>Fateh</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>BRISC: annotated dataset for brain tumor segmentation and classification</article-title>. <source>Sci Data</source>. <year>2026</year>;<volume>13</volume>(<issue>1</issue>):<fpage>361</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41597-026-06753-y</pub-id>; <pub-id pub-id-type="pmid">41644571</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Filvantorkaman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Torkaman</surname> <given-names>PM</given-names></string-name>, <string-name><surname>Filvan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zabihi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Moradi</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Fusion-based brain tumor classification using deep learning and explainable AI, and rule-based reasoning</article-title>. <comment>arXiv:2508.06891. 2025</comment>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zikry</surname> <given-names>TM</given-names></string-name>, <string-name><surname>Allen</surname> <given-names>GI</given-names></string-name></person-group>. <article-title>Are machine learning interpretations reliable? A stability study on global interpretations</article-title>. <comment>arXiv:2505.15728. 2025</comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mzoughi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Njeh</surname> <given-names>I</given-names></string-name>, <string-name><surname>BenSlima</surname> <given-names>M</given-names></string-name>, <string-name><surname>Farhat</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mhiri</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Vision transformers (ViT) and deep convolutional neural network (D-CNN)-based models for MRI brain primary tumors images multi-classification supported by explainable artificial intelligence (XAI)</article-title>. <source>Vis Comput</source>. <year>2025</year>;<volume>41</volume>(<issue>4</issue>):<fpage>2123</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00371-024-03524-x</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Iftikhar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Anjum</surname> <given-names>N</given-names></string-name>, <string-name><surname>Siddiqui</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Ur Rehman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ramzan</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Explainable CNN for brain tumor detection and classification through XAI based key features identification</article-title>. <source>Brain Inform</source>. <year>2025</year>;<volume>12</volume>(<issue>1</issue>):<fpage>10</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40708-025-00257-y</pub-id>; <pub-id pub-id-type="pmid">40304860</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ilani</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>D</given-names></string-name>, <string-name><surname>Banad</surname> <given-names>YM</given-names></string-name></person-group>. <article-title>T1-weighted MRI-based brain tumor classification using hybrid deep learning models</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>7010</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-92020-w</pub-id>; <pub-id pub-id-type="pmid">40016334</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tampu</surname> <given-names>IE</given-names></string-name>, <string-name><surname>Bianchessi</surname> <given-names>T</given-names></string-name>, <string-name><surname>Blystad</surname> <given-names>I</given-names></string-name>, <string-name><surname>Lundberg</surname> <given-names>P</given-names></string-name>, <string-name><surname>Nyman</surname> <given-names>P</given-names></string-name>, <string-name><surname>Eklund</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Pediatric brain tumor classification using deep learning on MR images with age fusion</article-title>. <source>Neuro Oncol Adv</source>. <year>2025</year>;<volume>7</volume>(<issue>1</issue>):<fpage>vdae205</fpage>. doi:<pub-id pub-id-type="doi">10.1093/noajnl/vdae205</pub-id>; <pub-id pub-id-type="pmid">39777258</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Javeed</surname> <given-names>MD</given-names></string-name>, <string-name><surname>Nagaraju</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chandrasekaran</surname> <given-names>R</given-names></string-name>, <string-name><surname>Rajulu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Tumuluru</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Brain tumor segmentation and classification with hybrid clustering, probabilistic neural networks</article-title>. <source>J Intell Fuzzy Syst</source>. <year>2023</year>;<volume>45</volume>(<issue>4</issue>):<fpage>6485</fpage>&#x2013;<lpage>500</lpage>. doi:<pub-id pub-id-type="doi">10.3233/jifs-232493</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pareek</surname> <given-names>KK</given-names></string-name>, <string-name><surname>Ameta</surname> <given-names>GK</given-names></string-name></person-group>. <article-title>Explainable AI (XAI) based verification for brain tumor classification using deep learning techniques</article-title>. <source>IET Conf Proc</source>. <year>2025</year>;<volume>2025</volume>(<issue>7</issue>):<fpage>1184</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1049/icp.2025.1569</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kanna</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Salau</surname> <given-names>AO</given-names></string-name></person-group>. <article-title>New cognitive computational strategy for optimizing brain tumour classification using magnetic resonance imaging Data</article-title>. <source>Intell Based Med</source>. <year>2025</year>;<volume>11</volume>:<fpage>100215</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ibmed.2025.100215</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Neamah</surname> <given-names>K</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>F</given-names></string-name>, <string-name><surname>Waheed</surname> <given-names>SR</given-names></string-name>, <string-name><surname>Kurdi</surname> <given-names>WHM</given-names></string-name>, <string-name><surname>Taha</surname> <given-names>AY</given-names></string-name>, <string-name><surname>Kadhim</surname> <given-names>KA</given-names></string-name></person-group>. <article-title>Utilizing deep improved ResNet50 for brain tumor classification based MRI</article-title>. <source>IEEE Open J Comput Soc</source>. <year>2024</year>;<volume>5</volume>(<issue>11</issue>):<fpage>446</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ojcs.2024.3453924</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Akg&#x00FC;ndo&#x011F;du</surname> <given-names>A</given-names></string-name>, <string-name><surname>&#x00C7;elikba&#x015F;</surname> <given-names>&#x015E;</given-names></string-name></person-group>. <article-title>Explainable deep learning framework for brain tumor detection: integrating LIME, Grad-CAM, and SHAP for enhanced accuracy</article-title>. <source>Med Eng Phys</source>. <year>2025</year>;<volume>144</volume>(<issue>1</issue>):<fpage>104405</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.medengphy.2025.104405</pub-id>; <pub-id pub-id-type="pmid">40925692</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>T</given-names></string-name>, <string-name><surname>Brennan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Mileo</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bendechache</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Image data augmentation approaches: a comprehensive survey and future directions</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>187536</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3470122</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mandle</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Sahu</surname> <given-names>SP</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Brain tumor segmentation and classification in MRI using clustering and kernel-based SVM</article-title>. <source>Biomed Pharmacol J</source>. <year>2022</year>;<volume>15</volume>(<issue>2</issue>):<fpage>699</fpage>&#x2013;<lpage>716</lpage>. doi:<pub-id pub-id-type="doi">10.13005/bpj/2409</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tonmoy</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Shams</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Adnan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Mridha</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Safran</surname> <given-names>M</given-names></string-name>, <string-name><surname>Alfarhood</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>X-Brain: explainable recognition of brain tumors using robust deep attention CNN</article-title>. <source>Biomed Signal Process Control</source>. <year>2025</year>;<volume>100</volume>(<issue>18</issue>):<fpage>106988</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2024.106988</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nahiduzzaman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Abdulrazak</surname> <given-names>LF</given-names></string-name>, <string-name><surname>Kibria</surname> <given-names>HB</given-names></string-name>, <string-name><surname>Khandakar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ayari</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Ahamed</surname> <given-names>MF</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A hybrid explainable model based on advanced machine learning and deep learning models for classifying brain tumors using MRI images</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>1649</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-85874-7</pub-id>; <pub-id pub-id-type="pmid">39794374</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Naseer</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yasir</surname> <given-names>T</given-names></string-name>, <string-name><surname>Azhar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shakeel</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zafar</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Computer-aided brain tumor diagnosis: performance evaluation of deep learner CNN using augmented brain MRI</article-title>. <source>Int J Biomed Imaging</source>. <year>2021</year>;<volume>2021</volume>(<issue>2</issue>):<fpage>5513500</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2021/5513500</pub-id>; <pub-id pub-id-type="pmid">34234822</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nassar</surname> <given-names>SE</given-names></string-name>, <string-name><surname>Yasser</surname> <given-names>I</given-names></string-name>, <string-name><surname>Amer</surname> <given-names>HM</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>A robust MRI-based brain tumor classification via a hybrid deep learning technique</article-title>. <source>J Supercomput</source>. <year>2024</year>;<volume>80</volume>(<issue>2</issue>):<fpage>2403</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11227-023-05549-w</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>KT</given-names></string-name>, <string-name><surname>Park</surname> <given-names>HM</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Vankerschaver</surname> <given-names>J</given-names></string-name>, <string-name><surname>De Neve</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Towards improved cervical cancer screening: vision transformer-based classification and interpretability</article-title>. In: <conf-name>Proceedings of the 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI); 2025 Apr 14&#x2013;17</conf-name>; <publisher-loc>Houston, TX, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ISBI60581.2025.10981006</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>Asmita</collab>, <string-name><surname>Mittal</surname> <given-names>P</given-names></string-name></person-group>. <article-title>From black box AI to XAI in neuro-oncology: a survey on MRI-based tumor detection</article-title>. <source>Discov Artif Intell</source>. <year>2025</year>;<volume>5</volume>(<issue>1</issue>):<fpage>30</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s44163-025-00247-3</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ugbomeh</surname> <given-names>O</given-names></string-name>, <string-name><surname>Yiye</surname> <given-names>V</given-names></string-name>, <string-name><surname>Ibeke</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ezenkwu</surname> <given-names>CP</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>V</given-names></string-name>, <string-name><surname>Alkhayyat</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Machine learning algorithms for stroke risk prediction leveraging on explainable artificial intelligence techniques (XAI)</article-title>. In: <conf-name>Proceedings of the 2024 International Conference on Electrical Electronics and Computing Technologies (ICEECT); 2024 Aug 29&#x2013;31</conf-name>; <publisher-loc>Greater Noida, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iceect61758.2024.10739320</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Akinlade</surname> <given-names>O</given-names></string-name>, <string-name><surname>Vakaj</surname> <given-names>E</given-names></string-name>, <string-name><surname>Dridi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tiwari</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ortiz-Rodriguez</surname> <given-names>F</given-names></string-name></person-group>. <chapter-title>Semantic segmentation of the lung to examine the effect of COVID-19 using UNET model</chapter-title>. In: <source>Applied machine learning and data analytics</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>. p. <fpage>52</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-34222-6_5</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Adnan</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Ghazal</surname> <given-names>TM</given-names></string-name>, <string-name><surname>Saleem</surname> <given-names>M</given-names></string-name>, <string-name><surname>Farooq</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Yeun</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep learning driven interpretable and informed decision making model for brain tumour prediction using explainable AI</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>19223</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-03358-0</pub-id>; <pub-id pub-id-type="pmid">40451921</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kanchanamala</surname> <given-names>P</given-names></string-name>, <string-name><surname>Kuppusamy</surname> <given-names>V</given-names></string-name>, <string-name><surname>Ganesan</surname> <given-names>G</given-names></string-name></person-group>. <article-title>QDCNN-DMN: a hybrid deep learning approach for brain tumor classification using MRI images</article-title>. <source>Biomed Signal Process Control</source>. <year>2025</year>;<volume>101</volume>(<issue>6</issue>):<fpage>107199</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2024.107199</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar Tiwary</surname> <given-names>P</given-names></string-name>, <string-name><surname>Johri</surname> <given-names>P</given-names></string-name>, <string-name><surname>Katiyar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chhipa</surname> <given-names>MK</given-names></string-name></person-group>. <article-title>Deep learning-based MRI brain tumor segmentation with EfficientNet-enhanced UNet</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>1</issue>):<fpage>54920</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3554405</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Narayankar</surname> <given-names>P</given-names></string-name>, <string-name><surname>Baligar</surname> <given-names>VP</given-names></string-name></person-group>. <article-title>Explainability of brain tumor classification based on region</article-title>. In: <conf-name>Proceedings of the 2024 International Conference on Emerging Technologies in Computer Science for Interdisciplinary Applications (ICETCS); 2024 Apr 22&#x2013;23</conf-name>; <publisher-loc>Bengaluru, India</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/icetcs61022.2024.10544289</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Qari</surname> <given-names>S</given-names></string-name>, <string-name><surname>Thafar</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>Brain stroke detection and classification using CT imaging with transformer models and explainable AI</article-title>. <comment>arXiv:2507.09630. 2025</comment>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qezelbash-Chamak</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hicklin</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A hybrid learnable fusion of ConvNeXt and swin transformer for optimized image classification</article-title>. <source>IoT</source>. <year>2025</year>;<volume>6</volume>(<issue>2</issue>):<fpage>30</fpage>. doi:<pub-id pub-id-type="doi">10.3390/iot6020030</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mahesh</surname> <given-names>TR</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>M</given-names></string-name>, <string-name><surname>Anupama</surname> <given-names>TA</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>VV</given-names></string-name>, <string-name><surname>Geman</surname> <given-names>O</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>VD</given-names></string-name></person-group>. <article-title>An XAI-enhanced efficientNetB0 framework for precision brain tumor detection in MRI imaging</article-title>. <source>J Neurosci Meth</source>. <year>2024</year>;<volume>410</volume>(<issue>3</issue>):<fpage>110227</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jneumeth.2024.110227</pub-id>; <pub-id pub-id-type="pmid">39038716</pub-id></mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rahman</surname> <given-names>T</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>J</given-names></string-name></person-group>. <article-title>MRI-based brain tumor classification using a dilated parallel deep convolutional neural network</article-title>. <source>Digital</source>. <year>2024</year>;<volume>4</volume>(<issue>3</issue>):<fpage>529</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.3390/digital4030027</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Raja</surname> <given-names>RV</given-names></string-name>, <string-name><surname>Jayashankari</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sheela</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jancy Sickory Daisy</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gokilam</surname> <given-names>GG</given-names></string-name>, <string-name><surname>Joel</surname> <given-names>MR</given-names></string-name></person-group>. <chapter-title>Metrics and techniques for evaluating machine learning models and optimization algorithms</chapter-title>. In: <source>AI model design and data management for disease prediction</source>. <publisher-loc>Hershey, PA, USA</publisher-loc>: <publisher-name>IGI Global Scientific Publishing</publisher-name>; <year>2025</year>. p. <fpage>193</fpage>&#x2013;<lpage>222</lpage>. doi:<pub-id pub-id-type="doi">10.4018/979-8-3373-5137-7.ch007</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rajendran</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rajagopal</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Thanarajan</surname> <given-names>T</given-names></string-name>, <string-name><surname>Shankar</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alsubaie</surname> <given-names>NM</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Automated segmentation of brain tumor MRI images using deep learning</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>64758</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2023.3288017</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rasa</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Talukder</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Khalid</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kazi</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Brain tumor classification using fine-tuned transfer learning models on magnetic resonance imaging (MRI) images</article-title>. <source>Digit Health</source>. <year>2024</year>;<volume>10</volume>:<fpage>20552076241286140</fpage>. doi:<pub-id pub-id-type="doi">10.1177/20552076241286140</pub-id>; <pub-id pub-id-type="pmid">39381813</pub-id></mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rezvani</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fateh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khosravi</surname> <given-names>H</given-names></string-name></person-group>. <article-title>ABANet: attention boundary-aware network for image segmentation</article-title>. <source>Expert Syst</source>. <year>2024</year>;<volume>41</volume>(<issue>9</issue>):<fpage>e13625</fpage>. doi:<pub-id pub-id-type="doi">10.1111/exsy.13625</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ronneberger</surname> <given-names>O</given-names></string-name>, <string-name><surname>Fischer</surname> <given-names>P</given-names></string-name>, <string-name><surname>Brox</surname> <given-names>T</given-names></string-name></person-group>. <article-title>U-Net: convolutional networks for biomedical image segmentation</article-title>. In: <conf-name>Proceedings of the Medical Image Computing and Computer-Assisted Intervention&#x2014;MICCAI 2015</conf-name>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2015</year>. p. <fpage>234</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-24574-4_28</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Saeed</surname> <given-names>T</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Hamza</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shabaz</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>WZ</given-names></string-name>, <string-name><surname>Alhayan</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Neuro-XAI: explainable deep learning framework based on deeplabV<sup>3&#x002B;</sup> and Bayesian optimization for segmentation and classification of brain tumor in MRI scans</article-title>. <source>J Neurosci Meth</source>. <year>2024</year>;<volume>410</volume>(<issue>6</issue>):<fpage>110247</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jneumeth.2024.110247</pub-id>; <pub-id pub-id-type="pmid">39128599</pub-id></mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shamshad</surname> <given-names>N</given-names></string-name>, <string-name><surname>Sarwr</surname> <given-names>D</given-names></string-name>, <string-name><surname>Almogren</surname> <given-names>A</given-names></string-name>, <string-name><surname>Saleem</surname> <given-names>K</given-names></string-name>, <string-name><surname>Munawar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rehman</surname> <given-names>AU</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Enhancing brain tumor classification by a comprehensive study on transfer learning techniques and model efficiency using MRI datasets</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>100407</fpage>&#x2013;<lpage>18</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3430109</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sharif</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tanvir</surname> <given-names>U</given-names></string-name>, <string-name><surname>Munir</surname> <given-names>EU</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Yasmin</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Brain tumor segmentation and classification by improved binomial thresholding and multi-features selection</article-title>. <source>J Ambient Intell Humaniz Comput</source>. <year>2024</year>;<volume>15</volume>(<issue>1</issue>):<fpage>1063</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s12652-018-1075-x</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Asif</surname> <given-names>RN</given-names></string-name>, <string-name><surname>Naseem</surname> <given-names>MT</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mazhar</surname> <given-names>T</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Brain tumor detection empowered with ensemble deep learning approaches from MRI scan images</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>15002</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-99576-7</pub-id>; <pub-id pub-id-type="pmid">40301625</pub-id></mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thahiruddin</surname> <given-names>M</given-names></string-name></person-group>. <article-title>CNNs vs. hybrid transformers for brain tumor classification on the BRISC dataset</article-title>. <source>J Aplikasi Teknologi Informasi Dan Manajemen</source>. <year>2025</year>;<volume>6</volume>(<issue>1</issue>):<fpage>24</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.31102/jatim.v6i1.3545</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Tran</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Le-Thi</surname> <given-names>ND</given-names></string-name></person-group>. <chapter-title>Development of an explainable AI system for brain tumor diagnosis from MRI scans</chapter-title>. In: <source>Future data and security engineering</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>63</fpage>&#x2013;<lpage>77</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-95-4721-0_5</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yiye</surname> <given-names>V</given-names></string-name>, <string-name><surname>Ugbomeh</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ezenkwu</surname> <given-names>CP</given-names></string-name>, <string-name><surname>Ibeke</surname> <given-names>E</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>V</given-names></string-name>, <string-name><surname>Alkhayyat</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Investigating key contributors to hospital appointment no-shows using explainable AI</article-title>. In: <conf-name>Proceedings of the 2024 International Conference on Electrical Electronics and Computing Technologies (ICEECT); 2024 Aug 29&#x2013;31</conf-name>; <publisher-loc>Greater Noida, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iceect61758.2024.10739123</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu </surname> <given-names>Y</given-names></string-name>, <string-name><surname>Owais</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kateb</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chaddad</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Deep modeling and optimization of medical image classification</article-title>. In: <conf-name>Proceedings of the 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI); 2025 Apr 14&#x2013;17</conf-name>; <publisher-loc>Houston, TX, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/isbi60581.2025.10981184</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Rahman Siddiquee</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Tajbakhsh</surname> <given-names>N</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>J</given-names></string-name></person-group>. <chapter-title>UNet&#x002B;&#x002B;: a nested U-Net architecture for medical image segmentation</chapter-title>. In: <source>Deep learning in medical image analysis and multimodal learning for clinical decision support</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2018</year>. p. <fpage>3</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-00889-5_1</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kusuma</surname> <given-names>PV</given-names></string-name>, <string-name><surname>Reddy</surname> <given-names>SCM</given-names></string-name></person-group>. <article-title>Brain tumor segmentation and classification using MRI: modified segnet model and hybrid deep learning architecture with improved texture features</article-title>. <source>Comput Biol Chem</source>. <year>2025</year>;<volume>117</volume>(<issue>1</issue>):<fpage>108381</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiolchem.2025.108381</pub-id>; <pub-id pub-id-type="pmid">40020564</pub-id></mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ranjbarzadeh</surname> <given-names>R</given-names></string-name>, <string-name><surname>Keles</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bendechache</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Comparative analysis of real-clinical MRI and BraTS datasets for brain tumor segmentation</article-title>. <source>IET Conf Proc</source>. <year>2024</year>;<volume>2024</volume>(<issue>10</issue>):<fpage>39</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1049/icp.2024.3274</pub-id>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>M</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Sparse dynamic volume TransUNet with multi-level edge fusion for brain tumor segmentation</article-title>. <source>Comput Biol Med</source>. <year>2024</year>;<volume>172</volume>(<issue>8</issue>):<fpage>108284</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiomed.2024.108284</pub-id>; <pub-id pub-id-type="pmid">38503086</pub-id></mixed-citation></ref>
</ref-list>
</back></article>
























