<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">59452</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2025.059452</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>ParMamba: A Parallel Architecture Using CNN and Mamba for Brain Tumor Classification</article-title>
<alt-title alt-title-type="left-running-head">ParMamba: A Parallel Architecture Using CNN and Mamba for Brain Tumor Classification</alt-title>
<alt-title alt-title-type="right-running-head">ParMamba: A Parallel Architecture Using CNN and Mamba for Brain Tumor Classification</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Su</surname><given-names>Gaoshuai</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Hongyang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>lihy@whu.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Huafeng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>College of Computer and Information Technology, China Three Gorges University</institution>, <addr-line>Yichang, 443000</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Computer Engineering, Jingchu University of Technology</institution>, <addr-line>Jingmen, 448000</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Hongyang Li. Email: <email>lihy@whu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>03</day><month>03</month><year>2025</year>
</pub-date>
<volume>142</volume>
<issue>3</issue>
<fpage>2527</fpage>
<lpage>2545</lpage>
<history>
<date date-type="received">
<day>08</day>
<month>10</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>12</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_59452.pdf"></self-uri>
<abstract>
<p>Brain tumors, one of the most lethal diseases with low survival rates, require early detection and accurate diagnosis to enable effective treatment planning. While deep learning architectures, particularly Convolutional Neural Networks (CNNs), have shown significant performance improvements over traditional methods, they struggle to capture the subtle pathological variations between different brain tumor types. Recent attention-based models have attempted to address this by focusing on global features, but they come with high computational costs. To address these challenges, this paper introduces a novel parallel architecture, ParMamba, which uniquely integrates Convolutional Attention Patch Embedding (CAPE) and the ConvMamba block including CNN, Mamba and the channel enhancement module, marking a significant advancement in the field. The unique design of ConvMamba block enhances the ability of model to capture both local features and long-range dependencies, improving the detection of subtle differences between tumor types. The channel enhancement module refines feature interactions across channels. Additionally, CAPE is employed as a downsampling layer that extracts both local and global features, further improving classification accuracy. Experimental results on two publicly available brain tumor datasets demonstrate that ParMamba achieves classification accuracies of 99.62% and 99.35%, outperforming existing methods. Notably, ParMamba surpasses vision transformers (ViT) by 1.37% in accuracy, with a throughput improvement of over 30%. These results demonstrate that ParMamba delivers superior performance while operating faster than traditional attention-based methods.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Brain tumor classification</kwd>
<kwd>convolutional neural networks</kwd>
<kwd>channel enhancement module</kwd>
<kwd>convolutional attention patch embedding</kwd>
<kwd>mamba</kwd>
<kwd>ParMamba</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Colleges and Universities in Hubei Province</funding-source>
<award-id>T201923</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Key Science and Technology Project of Jingmen</funding-source>
<award-id>2021ZDYF024</award-id>
<award-id>2022ZDYF019</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Cultivation Project of JingchuUniversity of Technology</funding-source>
<award-id>PY201904</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Brain tumor is a mass formed by the uncontrolled proliferation of brain cells, which occurs in males and females of all ages and is one of the most dangerous diseases in the world [<xref ref-type="bibr" rid="ref-1">1</xref>]. Therefore, early identification of brain tumors is particularly crucial to improve the treatment effect and survival rate of patients. Among various imaging techniques, magnetic resonance imaging (MRI) is currently the most popular method for detecting brain tumors [<xref ref-type="bibr" rid="ref-2">2</xref>]. From the perspective of MRI, brain tumors can be further classified into gliomas, meningiomas, and pituitary tumors based on their size, shape, and location [<xref ref-type="bibr" rid="ref-3">3</xref>]. Each type of tumor can be life-threatening. Therefore, it is crucial to classify different types of brain tumors effectively and then conduct targeted treatments for each type. However, due to the small structural differences between different brain tumors, accurately classifying them is challenging. Manual classification by doctors is inevitably prone to misdiagnosis and inefficient, while also increasing the burden on doctors. To overcome these difficulties, machine learning-based techniques have begun to be applied to MRI images for automatic brain tumor classification tasks and have played a crucial role in computer-aided diagnosis (CAD) systems.</p>
<p>So far, the emergence of CNN has revolutionized the field of medical image processing. Traditional machine learning algorithms such as Support Vector Machine (SVM) and Decision Tree often require manual feature design, which relies on domain knowledge and expert experience, and the process is cumbersome. In contrast, CNN can automatically learn feature representations from raw data, reducing the reliance on manual feature engineering. Therefore, CNN have significantly improved the performance of CAD systems. Although CNN have achieved better classification performance compared to traditional machine learning methods, due to the characteristics of local feature extraction, CNNS may ignore some key inherent tumor properties when processing brain tumor classification tasks, such as the contextual information surrounding the tumor region and the size variations of the tumor. The lack of these pieces of information can lead to limitations in the model&#x2019;s ability to identify tumors. Additionally, the high similarity between brain tumor categories add to additional difficulties for the practical application of CNN. To overcome these challenges, researchers [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>] have increasingly integrated attention mechanisms into CNN models, By incorporating attention mechanisms, CNNs are able to capture finer-grained details within tumor regions, which are crucial for distinguishing between different types of brain tumors.</p>
<p>The emergence of vision transformer (ViT) [<xref ref-type="bibr" rid="ref-7">7</xref>] has overcome the limitations of CNN, yet it require extensive computational resources during training, limiting the input sequence length and increasing training time. Recently, Mamba [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>] has not only overcome these difficulties but also possesses the ability to extract global features. The Spatial State Model (SSM) represented by Mamba operates in linear time with respect to sequence length. Benefiting from its linear time complexity, Mamba has a lower computational cost, potentially demonstrating higher computational efficiency in handling complex tasks.</p>
<p>Despite significant advancements in deep learning for medical image classification, existing studies still fall short in capturing subtle distinctions in brain tumors and reducing computational costs. Traditional CNN perform well in local feature extraction but are limited in handling long-range dependencies and global features. Additionally, while ViT overcome the limitations of CNN, they require substantial computational resources. The main objective of this study is to develop an efficient brain tumor classification model that achieves high classification accuracy, with the aim of enhancing the practical applicability of automatic brain tumor detection and diagnosis. To this end, we propose a novel parallel architecture named ParMamba, which utilizes ConvMamba blocks as its backbone and incorporates Convolutional Attention Patch Embedding (CAPE). The ConvMamba block combines the local feature extraction capabilities of CNN with the global feature capturing strengths of the Mamba, resulting in improved computational efficiency and classification accuracy. Furthermore, the channel enhancement module is introduced to enhance cross-channel feature interactions. CAPE captures brain tumor information from multiple perspectives by applying both max pooling and convolutional pooling in two directions, which enriches the feature representation. The main contributions of this paper are summarized as follows:
<list list-type="bullet">
<list-item>
<p>This paper propose a parallel architecture called ParMamba, designed for effective multi-class classification of brain tumors in MRI images.</p></list-item>
<list-item>
<p>This paper design a novel module, ConvMamba block, which can extract brain tumor image features from local and global contexts.</p></list-item>
<list-item>
<p>Comprehensive experiments were conducted on two brain tumor datasets, which verified the excellent ability of ParMamba in brain tumor classification.</p></list-item>
<list-item>
<p>Comparing ParMamba with the most advanced brain tumor classification methods demonstrates that ParMamba outperforming existing brain tumor classification methods.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Brain tumors exhibit a variety of types, each with distinct characteristics and manifestations. Furthermore, many brain tumors have indistinct boundaries, making them easily confusable with surrounding normal tissue, which adds to the complexity of classification. In early automated systems, people used machine learning methods to identify tumors of brain tumors from MRI images. El-Dahshan et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] used discrete wavelet transform (DWT) to obtain features related to MRI images and classified them using a k-nearest neighbor-based classifier. Shim et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] combined finite element analysis and machine learning methods for detecting brain injuries. Das et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] adopted Ripplet Transform Type-I (RT) to represent significant features of brain MRI images and used least squares support vector machines to classify brain MRI images. Zhang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] compared traditional training methods such as Scaled Chaotic Artificial Bee Colony (SCABC), momentum BP, genetic algorithms, and simulated annealing, indicating that the SCABC method is better. Shim et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] developed an efficient computational pipeline to generate finite element models of brain injury for specific individuals, predicting structural damage following trauma. However, traditional machine learning still poses challenges due to its reliance on manual intervention and the requirement for extensive domain knowledge and expertise.</p>
<p>CNN can automatically learn feature representations from raw data, reducing the reliance on manual feature engineering. Therefore, CNN has been widely used in CAD methods for medical images and has revolutionized the field of medical image analysis. Ayadi et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed a deep CNN model, utilizing ten different convolutional layers to extract features from brain tumor images, aiming to enhance the ability to capture brain tumor features. Atha et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] employed CNN as the base architecture and incorporated the idea of semi-supervised learning, combining the training process of labeled and unlabeled brain tumor data, enabling the network to learn from both types of data simultaneously. Rizwan et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] used Gaussian convolution kernels to extract local features of brain tumors, enhancing the accuracy of brain tumor feature extraction. Zhu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] adopted DenseNets as the base network structure and introduced a biologically inspired evolutionary algorithm to optimize the network structure of DenseNets, adapting it to the characteristics of medical image data. Aamir et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] used two pre-trained CNN networks, EfficientNet and ResNet50, to extract features from brain tumor images separately. Then, they employed partial least squares (PLS) to fuse the feature vectors extracted from the two models, forming a hybrid feature vector. Kumar et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed a multi-class brain tumor classification method based on ResNet-50 and global average pooling, which performs well in addressing gradient disappearance and overfitting issues in deep networks. Gursoy et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] integrated graph neural networks (GNNs) to capture relational dependencies between image regions and CNN to extract spatial features, enhancing the detection of brain tumors.</p>
<p>However, when using CNN models for brain tumor classification, limited data can lead to the over-fitting problem. Therefore, to address the issue of small brain tumor image samples, data augmentation techniques have been applied to enlarge the dataset in some works. Yaqub et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] utilized generative adversarial networks (GANs) for data augmentation. Ghassemi et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] employed GANs for preprocessing brain tumor images, in addition to data augmentation methods such as image rotation and mirroring. Li et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] adopted data augmentation techniques including rotation, horizontal flipping, and vertical flipping, as well as salt-and-pepper noise as a data augmentation method.</p>
<p>In recent years, with the introduction of attention mechanism into the field of computer vision, researchers have also combined attention mechanism with CNN and applied it to medical images. Dutta et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] adopted CNN as the base architecture and introduced a lightweight global attention mechanism after the CNN layers, enhancing the model&#x2019;s ability to extract more salient features from brain tumor images. Wang et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] utilized a pre-trained vision transformer as the base architecture, and combined with token merging techniques, to extract key information from brain tumor images. Isunuri et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] employed a pre-trained efficientNetB4 to extract brain tumor features and then used multi-path convolution and multi-head attention for feature enhancement.</p>
<p>More recently, Mamba based on state-space models (SSM) has emerged in people&#x2019;s vision. Mamba not only has the ability of global feature extraction, but also exhibits linear complexity related to the size of the input image. Yue et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] combined the local feature extraction capability of CNN with the ability of SSM to capture long-range dependencies, thereby modeling medical images in different modes. Ma et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] have utilized the integration of CNN&#x2019;s local pattern recognition capabilities with Mamba&#x2019;s global context understanding, enabling automatic adaptation to various datasets and applicability in segmentation tasks across diverse biomedical imaging fields. Ruan et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] have proposed a medical image segmentation model based on SSM. By leveraging Visual State Space (VSS) blocks to capture extensive contextual information, they have constructed an asymmetric encoder-decoder architecture, marking it as the first medical image segmentation model purely built upon SSM.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methodology</title>
<p>The architecture of the proposed ParMamba adopts four units for brain tumor classification, which is similar to numerous prior studies [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, Each unit incorporates a patch embedding layer, succeeded by multiple sequentially arranged ConvMamba blocks. Specifically, the patch embedding layer adopts convolutional attention patch embedding layer [<xref ref-type="bibr" rid="ref-32">32</xref>], while each ConvMamba block consists of a ConvBlock, a MambaBlock, and a channel enhancement module.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The overall architecture of the proposed ParMamba</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-1.tif"/>
</fig>
<p>In the first unit, The channel dimension of the input image <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is mapped to <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> channels via convolutional attention patch embedding layer, obtaining the embedded image <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula>. Then, the <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is fed into <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>N</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> sequentially stacked ConvMamba blocks to extract image features, yielding <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula>. The second, third, and fourth units repeat the above operations, resulting in the fourth unit, the <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>4</mml:mn></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:mn>32</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:mn>32</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula> is obtained. Finally, classifier predicts whether the input image is glioma, meningioma, notumor, or pituitary.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Convolutional Attention Patch Embedding</title>
<p>Convolutional Attention Patch Embedding (CAPE) [<xref ref-type="bibr" rid="ref-32">32</xref>] is a downsampling method, as shonw in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. CAPE combines convolutional patch embedding and channel attention module (CAM). CAPE processes the input image or the feature map <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> from the previous stage, with dimensions <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, which can be described in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mi>E</mml:mi><mml:mi>m</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Convolutional attention patch embedding</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-2.tif"/>
</fig>
<p>The input <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> will be patched into <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> represents the downsampling ratio, and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the adjustable number of channels for stage <italic>i</italic>.</p>
<p>In CAM, Overlap MaxPool captures global spatial information and downsamples the input, followed by a <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mn>11</mml:mn></mml:math></inline-formula> convolution to increase the number of channels. In convolutional patch embedding, a convolution with a kernel size of 7, stride of 4, and padding of 2 is used to capture local information, downsample, and increase the number of channels. Finally, the feature maps obtained from CAM and convolutional patch embedding are summed to calibrate the feature map. Utilizing the local and global feature extraction capabilities of CAPE and the ConvMamba block, the fusion of local and global features can adequately extract fine-grained and coarse-grained features from brain tumor images, thus enhancing the recognition ability for different types of brain tumors.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>ConvMamba Block</title>
<p>As depicted in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, using a parallel architecture with two separate branches for feature extraction enhances the model&#x2019;s ability to capture crucial features through feature fusion. This approach allows the model to leverage the complementary strengths of the individual branches, thereby bolstering its overall performance in identifying and extracting salient features [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>]. Therefore, the proposed ConvMamba block also adopts a parallel structure, consisting of a ConvBlock, a MambaBlock and a channel enhancement module. The ConvBlock, comprised of multiple convolutional layers and activation functions extracts brain tumor characteristics. Prior work [<xref ref-type="bibr" rid="ref-29">29</xref>] has validated the feasibility of the MambaBlock. As a state-space model (SSM) [<xref ref-type="bibr" rid="ref-8">8</xref>], Mamba possesses the characteristic of linear time complexity, ensuring efficient feature extraction and model training when processing brain tumor images. Additionally, Mamba&#x2019;s ability to comprehend global context is particularly significant for identifying complex structures in brain tumor images, as brain tumors often exhibit diverse shapes, sizes, and locations. Therefore, the proposed ConvMamba block utilizes the MambaBlock to extract global features of brain tumors, allowing for better attention to the lesion locations.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The parallel architecture of ConvMamba block</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-3.tif"/>
</fig>
<p>In the ConvMamba block, the feature maps <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula> is obtained through a CAPE. The channels <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> of <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is then divided into two parts, which are processed by a ConvBlock and a MambaBlock, respectively. Afterward, the channels are merged, and a channel enhancement module is applied to explore the dependencies among the channels. This process can be mathematically expressed with the following equations:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>X</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>X</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>X</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>X</mml:mi><mml:mn>4</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>4</mml:mn></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mfrac><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>X</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mn>4</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula>. <italic>Norm(.)</italic> represents layer normalization, <italic>Split(.)</italic> represents splitting the channels <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. <italic>Conv(.)</italic> and <italic>Mamba(.)</italic> represent the ConvBlock and MambaBlock, respectively, used for extracting feature information. <italic>Cat(.)</italic> is the concatenation of the channels of <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>X</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>X</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>. <italic>CEM(.)</italic> represents the channel enhancement module, which facilitates the information exchange between different channels.</p>
<p>This section describes each component of the ConvMamba block in detail.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>ConvBlock</title>
<p>Traditional convolutions typically employ <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mn>11</mml:mn></mml:math></inline-formula> or <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mn>33</mml:mn></mml:math></inline-formula> kernels and directly perform convolutions on all channels of the input data. This results in a relatively large number of parameters and computational costs. The structure of ConvBlock is inspired by MobileNet [<xref ref-type="bibr" rid="ref-34">34</xref>], which utilizes depthwise separable convolutions. It consists of two steps: firstly, depthwise convolution, where each channel of the input data is convolved separately; secondly, pointwise convolution, where a <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mn>11</mml:mn></mml:math></inline-formula> kernel is used to convolve the output of the depthwise convolution to fuse information from different channels. With depthwise separable convolutions, the number of parameters and computational cost are relatively low, allowing for the use of larger size kernels, such as <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mn>77</mml:mn></mml:math></inline-formula>. Larger size kernels have a wider receptive field, enabling them to capture more extensive information, which can be advantageous when dealing with the complex structures of brain tumor images.</p>
<p>Inspired by ConvNeXt [<xref ref-type="bibr" rid="ref-33">33</xref>], the ConvBlock employs an inverted bottleneck layer structure, where the middle is large and the ends are small. The inverted bottleneck layer structure, by expanding and contracting the number of channels, increases the nonlinearity of the network, enabling the model to have better generalization capabilities and effectively avoid information loss. This means that the model can better adapt to new data and improve performance on unseen data. Additionally, the ConvBlock incorporates layer normalization and the GELU activation function after the depthwise separable convolution and the first <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mn>11</mml:mn></mml:math></inline-formula> convolution. Layer normalization can enhance the stability of training, prevent issues such as gradient vanishing and gradient explosion, and provide a regularization effect that improves generalization performance. The GELU activation function enables the neural network to learn and represent complex nonlinear relationships.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>MambaBlock</title>
<p>First, a review of Structured State Space Sequence Models (S4) [<xref ref-type="bibr" rid="ref-35">35</xref>] is presented. S4 is based on the concept of hidden states, where an internal state variable <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>h</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> describes the system&#x2019;s state, and an input <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> drives the state transitions. It is typically defined as:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>B</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where, <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the state matrix, while <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>B</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>C</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> denote the projection parameters.</p>
<p>Next, this system is discretized by introducing a time scale parameter <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo>&#x25B3;</mml:mo></mml:math></inline-formula> and applying a fixed discretization rule, zero-order hold (ZOH), which transforms <italic>A</italic> and <italic>B</italic> into discrete parameters <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mover><mml:mi>A</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mover><mml:mi>B</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover></mml:math></inline-formula>, defined as:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mover><mml:mi>A</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x25B3;</mml:mo><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mover><mml:mi>B</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mtext>&#xA0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x25B3;</mml:mo><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x25B3;</mml:mo><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mo>&#x2212;</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>I</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mo>&#x22C5;</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mo>&#x25B3;</mml:mo><mml:mi>B</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mover><mml:mi>C</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>C</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref> can then be expressed with discrete parameters as:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mover><mml:mi>A</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mover><mml:mi>B</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mover><mml:mi>C</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Additionally, for an input sequence of length T, a global convolution with kernel <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mover><mml:mi>K</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover></mml:math></inline-formula> can be applied to compute the output of the equation as follows:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mover><mml:mi>K</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mover><mml:mi>B</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mover><mml:mrow><mml:mi>A</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:msup><mml:mover><mml:mi>A</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mover><mml:mi>B</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mover><mml:mi>K</mml:mi><mml:mo>&#x2212;</mml:mo></mml:mover></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The Selective State Space Model (S6) [<xref ref-type="bibr" rid="ref-8">8</xref>] is an extension of the S4 model. S6 dynamically adjusts certain parameters (such as <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mo>&#x25B3;</mml:mo></mml:math></inline-formula>, B, and C), enabling the model to selectively retain or disregard information based on input data. This flexibility allows the model to better accommodate diverse sequence characteristics.</p>
<p>The core of MambaBlock is the 2D selective scanning (SS2D), which was proposed by Vmamba [<xref ref-type="bibr" rid="ref-10">10</xref>]. As shown in the <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, SS2D comprises three components: scan expanding, S6 block, and scan merging. Firstly, scan expanding generates multiple sequences by extending the input image along four directions (upper left to the lower right, lower right to the upper left, upper right to the lower left, and lower left to the upper right). These sequences are then fed into the S6 block to extract and integrate detailed features from each direction. Finally, scan merging sums and merges the sequences from these four directions to restore an output image with the same size as the input image. This process aims to enable the model to distinguish and retain key information while filtering out irrelevant information, thereby enhancing the model&#x2019;s performance.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Illustration of 2D-Selective-Scan (SS2D). The input patches undergo scan expanding in four different directions, and each sequence is processed independently by distinct S6 blocks. Subsequently, the results are merged through scan merging to construct the final output</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-4.tif"/>
</fig>
<p>It is noteworthy that there are two residual connections in MambaBlock. In the first residual connection, the input is first normalized, then passes through a linear layer, a depthwise separable convolution, and a SiLU activation function to extract features. These features are then fed into the SS2D, followed by further feature extraction through normalization and a linear layer. In the second residual connection, the feature map obtained from the previous residual connection is normalized, and then a linear layer is applied for feature fusion. This yields the output of the MambaBlock branch. Furthermore, the normalization method utilized in MambaBlock is layer normalization.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Channel Enhancement Module</title>
<p>After merging parallel tokens, some previous methods [<xref ref-type="bibr" rid="ref-32">32</xref>] employed Multi-Layer Perceptrons (MLP) to facilitate information exchange between different channels, often resulting in a substantial computational burden. To reduce computational costs while maintaining classification performance, this paper proposes a lightweight channel enhancement module, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The channel enhancement module consists of a <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mn>11</mml:mn></mml:math></inline-formula> convolution with compressed channels to reduce computational costs and enable cross-channel communication, as well as a <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mn>11</mml:mn></mml:math></inline-formula> convolution with expanded channels to enhance local relationships. Normalization and activation functions are applied after each convolution to improve generalization performance, and residual connections are utilized to prevent gradient vanishing.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments and Results</title>
<p>This section provides a detailed description of the experiments conducted in this paper, including the experimental details, and the datasets used in the experiments. Comparisons are also made with existing brain tumor classification methods. Furthermore, ablation studies were performed to investigate the effectiveness of CAPE, MambaBlock, and the channel enhancement module.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset</title>
<p>To train and evaluate the proposed ParMamba in this paper, two publicly available brain tumor MRI datasets are utilized.</p>
<p>The first dataset (dataset 1) is available on the Kaggle website<xref ref-type="fn" rid="fn1"><sup>1</sup></xref><fn id="fn1"><label>1</label><p><ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset">https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset</ext-link>, accessed on 08 January 2024</p></fn>. Dataset 1 contains four different types of MRI images, including glioma, meningioma, pituitary, and no tumor, with 1621, 1645, 1757, and 2000 images respectively, totaling 7023 images. Among them, 5712 images are used for training, and 1311 images are used for testing. Some sample images are shown in <xref ref-type="fig" rid="fig-5">Fig. 5a</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Images in dataset 1 and dataset 2</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-5.tif"/>
</fig>
<p>The second dataset (dataset 2) is available on the Figshare website<xref ref-type="fn" rid="fn2"><sup>2</sup></xref><fn id="fn2"><label>2</label><p><ext-link ext-link-type="uri" xlink:href="https://figshare.com/articles/dataset/brain_tumor_dataset/1512427/5">https://figshare.com/articles/dataset/brain_tumor_dataset/1512427/5</ext-link>, accessed on 08 January 2024</p></fn>. Dataset 2 contains 3064 T1-weighted magnetic resonance imaging (MRI) images from 233 patients with different types of brain tumors. The images in dataset 2 are classified into three categories: 708 meningioma, 1462 glioma, and 930 pituitary tumors. In this paper, the dataset is split into training and testing sets with a ratio of 8:2. Some sample images are shown in <xref ref-type="fig" rid="fig-5">Fig. 5b</xref>.</p>
<p>Before the experiments, the datasets were first preprocessed. All images were resized to 224224 pixels and normalized for each dataset, scaling all pixel values to the range of [0, 1] by dividing by 255. Additionally, due to the limited data in dataset 2, two simple strategies for data augmentation were employed, including mirroring and rotation operations. Through data augmentation, dataset 2 was expanded to three times its original size, totaling 9192 images.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Implementation Details</title>
<p>All experiments in this paper were conducted in a Python 3.8 environment with the deep learning framework PyTorch 1.13.0. The CPU model is 96 Intel(R) Xeon(R) Platinum 8255C CPU, and the GPU model is RTX 3090 with 24 GB of video memory. The operating system is Ubuntu 20.04.01.</p>
<p>During the training process, an AdamW optimizer with a learning rate of 1e&#x2212;3 and momentum of 0.9 was utilized. A cross-entropy loss function with a weight decay of 0.05 was employed to optimize the model parameters. For training the model, the epoch was set to 500, with a batch size of 32. The channel dimensions of the four units [<inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>C</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>C</mml:mi><mml:mn>4</mml:mn></mml:msub></mml:math></inline-formula>] are [48, 96, 192, 384], and the number of stacked ConvMamba blocks [<inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>N</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>N</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>N</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>N</mml:mi><mml:mn>4</mml:mn></mml:msub></mml:math></inline-formula>] are [3, 3, 9, 3]. Additionally, no pre-trained weights were used. The specific configuration is presented in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Configurations of the ParMamba</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Hyper-parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Input size</td>
<td><inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mn>224</mml:mn><mml:mo>,</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Depth</td>
<td>18</td>
</tr>
<tr>
<td>Embedding dimension</td>
<td>384</td>
</tr>
<tr>
<td>Optimizer</td>
<td>AdamW</td>
</tr>
<tr>
<td>Weight decay</td>
<td>0.05</td>
</tr>
<tr>
<td>Learning rate</td>
<td>1e&#x2212;3</td>
</tr>
<tr>
<td>Momentum</td>
<td>0.9</td>
</tr>
<tr>
<td>Batch size</td>
<td>32</td>
</tr>
<tr>
<td>Max epoch</td>
<td>500</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation Metrics</title>
<p>Based on the characteristics of brain tumor images, this paper uses accuracy, precision, sensitivity, and F1 score as the evaluation metrics for the proposed model. These four metrics can directly reflect the effectiveness of the model. The expressions for calculating these metrics are as follows:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where, <italic>TP</italic> represents the number of true positive samples, <italic>FP</italic> is the number of false positive samples in the confusion matrix. <italic>TN</italic> stands for the number of true negatives, and <italic>FN</italic> represents the false negative samples in the confusion matrix.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Results</title>
<p>This section introduces the classification performance of the proposed ConvMemba on dataset 1 and dataset 2, and compares its results with the brain tumor classification models. It is noteworthy that the ablation study in this section is conducted on dataset 1.</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Ablation Study</title>
<p>This section compares Mamba with the attention mechanism by replacing the MambaBlock in ConvMamba block with a vanilla ViT [<xref ref-type="bibr" rid="ref-7">7</xref>]. As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, the accuracy of Mamba increased by 1.37% compared with ViT, while the parameters decreased by 1.7 M, FLOPs decreased by 55%, and throughput rate also increased by more than 30% compared with ViT. Therefore, it is demonstrated that Mamba not only outperforms ViT in terms of speed but also achieves a higher accuracy.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison between Mamba and ViT on dataset 1</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Models</th>
<th>Para (M)</th>
<th>FLOPs (G)</th>
<th>Throughput (img/s)</th>
<th>Accuracy (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>w/ViT</td>
<td>8.17</td>
<td>2.21</td>
<td>453</td>
<td>98.25</td>
</tr>
<tr>
<td>w/MambaBlock</td>
<td>6.47</td>
<td>0.98</td>
<td>594</td>
<td>99.62</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Furthermore, an ablation study was conducted to investigate the overall impact of CAPE, MambaBlock, and the proposed channel enhancement module. As indicated in <xref ref-type="table" rid="table-3">Table 3</xref>, without using CAPE and instead employing a single-branch convolutional patch embedding, the accuracy decreased by 1.22%, with only a 0.1 M reduction in parameters. This suggests that CAPE is crucial for this model. Disabling the channel enhancement module after combining the parallel branches of ParMamba resulted in a 0.62% decrease in accuracy, with only a 0.83 M reduction in parameters. This proves that the channel enhancement module effectively enhances the merged channels. Therefore, this paper adopts the channel enhancement module to improve the model&#x2019;s accuracy with a minimal increase in parameters. Removing MambaBlock and adopting a single-branch CNN structure led to a 0.31% drop in accuracy, while the number of parameters increased by 3.42 M. This demonstrates that the proposed parallel architecture of ParMamba is feasible, validating the effectiveness of MambaBlock&#x2019;s global context comprehension ability in handling complex brain tumor MRI images. At the same time, it reduces the number of parameters, making the model more lightweight.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Ablation study of ParMamba on dataset 1</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Models</th>
<th>Accuracy (%)</th>
<th>Precision (%)</th>
<th>Sensitivity (%)</th>
<th>F1 score (%)</th>
<th>Para (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td>w/o CAPE</td>
<td>98.40</td>
<td>98.31</td>
<td>98.37</td>
<td>98.34</td>
<td>6.37</td>
</tr>
<tr>
<td>w/o channel enhancement module</td>
<td>99.00</td>
<td>99.03</td>
<td>98.96</td>
<td>98.99</td>
<td>5.64</td>
</tr>
<tr>
<td>w/o MambaBlock</td>
<td>99.31</td>
<td>99.29</td>
<td>99.27</td>
<td>99.28</td>
<td>9.89</td>
</tr>
<tr>
<td>ParMamba</td>
<td>99.62</td>
<td>99.58</td>
<td>99.59</td>
<td>99.59</td>
<td>6.47</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Classification Performance On Dataset 1</title>
<p>The test loss and accuracy curves of the proposed ParMamba on dataset 1 are illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. It can be observed that ParMamba converges within 500 epochs and achieves perfect classification results in the classification test on dataset 1. The confusion matrix is shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. As seen from the confusion matrix, among the 1311 test images, only 5 images were misclassified, with 2 gliomas being wrongly labeled as meningiomas and pituitary tumors, 1 meningioma being misclassified as a pituitary tumor, and 2 pituitary tumors being mislabeled as gliomas and meningiomas. The classification performance of the four categories are listed in <xref ref-type="table" rid="table-4">Table 4</xref>. For the non-tumor category, it achieved 100% precision, sensitivity, and F1 score, indicating that the proposed ParMamba can accurately determine whether a tumor has developed in the brain. As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, the overall precision, sensitivity, F1 score, and accuracy for the four categories are 99.59%, 99.58%, 99.59%, and 99.62%, respectively. In summary, the proposed ParMamba can effectively learn the characteristics of brain tumor images, resulting in excellent classification performance on dataset 1.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Test loss and test accuracy on dataset 1</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-6.tif"/>
</fig><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Confusion matrix on dataset 1</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-7.tif"/>
</fig><table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance metrics on dataset 1</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Classification</th>
<th>Precision (%)</th>
<th>Sensitivity (%)</th>
<th>F1 score (%)</th>
<th>Accuracy (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Glioma</td>
<td>99.67</td>
<td>99.33</td>
<td>99.50</td>
<td>99.33</td>
</tr>
<tr>
<td>Meningioma</td>
<td>99.35</td>
<td>99.67</td>
<td>99.51</td>
<td>99.67</td>
</tr>
<tr>
<td>No tumor</td>
<td>100.00</td>
<td>100.00</td>
<td>100.00</td>
<td>100.00</td>
</tr>
<tr>
<td>Pituitart</td>
<td>99.33</td>
<td>99.33</td>
<td>99.33</td>
<td>99.33</td>
</tr>
<tr>
<td>Overall</td>
<td>99.59</td>
<td>99.58</td>
<td>99.59</td>
<td>99.62</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4_3">
<label>4.4.3</label>
<title>Classification Performance on Dataset 2</title>
<p>The test loss and accuracy curves of the proposed ParMamba on dataset 2 are depicted in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, and ParMamba converges within 500 epochs. It is evident that ParMamba achieved perfect classification results in the classification test on dataset 2. The confusion matrix is shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. As seen from the confusion matrix, out of 1837 test images, only 12 images were misclassified. Among the 855 glioma images, 5 were identified as meningiomas, and 2 were identified as pituitary tumors. In the 424 meningioma images, only 1 was misclassified as a glioma, and 3 were identified as pituitary tumors. Among the 558 pituitary tumor images, only 1 was misclassified as a glioma. The classification performance of the three categories are listed in <xref ref-type="table" rid="table-5">Table 5</xref>. Combined with the performance on dataset 1, the proposed ParMamba demonstrates excellent precision, sensitivity, and F1 score in the classification of meningiomas and pituitary tumors. As shown in <xref ref-type="table" rid="table-5">Table 5</xref>, the overall precision, sensitivity, F1 score, and accuracy for the three categories are 99.19%, 99.35%, 99.27%, and 99.35%, respectively. This indicates that the proposed ParMamba can effectively determine the type of tumor present in the brain.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Test loss and test accuracy on dataset 2</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-8.tif"/>
</fig><fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Confusion matrix on dataset 2</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_59452-fig-9.tif"/>
</fig><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance metrics on dataset 2</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Classification</th>
<th>Precision (%)</th>
<th>Sensitivity (%)</th>
<th>F1 score (%)</th>
<th>Accuracy (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Glioma</td>
<td>99.88</td>
<td>99.18</td>
<td>99.53</td>
<td>99.18</td>
</tr>
<tr>
<td>Meningioma</td>
<td>98.59</td>
<td>99.06</td>
<td>98.82</td>
<td>99.06</td>
</tr>
<tr>
<td>Pituitart</td>
<td>99.11</td>
<td>99.82</td>
<td>99.46</td>
<td>99.82</td>
</tr>
<tr>
<td>Overall</td>
<td>99.19</td>
<td>99.35</td>
<td>99.27</td>
<td>99.35</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4_4">
<label>4.4.4</label>
<title>Comparison with Existing Brain Tumor Classification Methods</title>
<p>To validate the performance of the proposed ParMamba, this section compares it with existing brain tumor classification methods, including those based on CNN and attention mechanisms. These methods were all conducted on either dataset 1 or dataset 2, as specifically shown in <xref ref-type="table" rid="table-6">Table 6</xref>. ParMamba outperforms Dense CNN Architecture [<xref ref-type="bibr" rid="ref-36">36</xref>], SSBTCNet [<xref ref-type="bibr" rid="ref-17">17</xref>], InceptionV3 [<xref ref-type="bibr" rid="ref-37">37</xref>], radimagenet pre-trained CNN [<xref ref-type="bibr" rid="ref-38">38</xref>], and DCST &#x002B; SVM [<xref ref-type="bibr" rid="ref-39">39</xref>] by 4.62%, 3.12%, 2.49%, 1.91%, and 1.91% in accuracy on dataset 1. Similarly, on dataset 2, it outperforms Deep CNN [<xref ref-type="bibr" rid="ref-16">16</xref>], BTSCNet [<xref ref-type="bibr" rid="ref-40">40</xref>], AP-CNN [<xref ref-type="bibr" rid="ref-41">41</xref>], MEEDNets [<xref ref-type="bibr" rid="ref-19">19</xref>], and RanMerFormer [<xref ref-type="bibr" rid="ref-27">27</xref>] by 4.61%, 2.68%, 1.93%, 0.72%, and 0.49% in accuracy, and also surpasses these brain tumor classification models in other metrics. Hese excellent results are closely related to the collaboration of ParMamba and CAPE which extracts local and global features of brain tumors, followed by the utilization of a channel enhancement module to enhance the merged channels. Additionally, the performance of the compared methods was directly obtained from their respective papers.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparison with existing brain tumor classification methods on dataset 1 and dataset 2</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Reference</th>
<th>Approach</th>
<th>Accuracy (%)</th>
<th>Precision (%)</th>
<th>Sensitivity (%)</th>
<th>F1 score (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Dataset 1</td>
<td>&#x00D6;zkaraca et al. [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>Dense CNN architecture</td>
<td>95.00</td>
<td>96.00</td>
<td>96.50</td>
<td>96.00</td>
</tr>
<tr>
<td></td>
<td>Atha et al. [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>SSBTCNet</td>
<td>96.50</td>
<td>92.50</td>
<td>91.80</td>
<td>92.00</td>
</tr>
<tr>
<td></td>
<td>Gomez-Guzman et al. [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>InceptionV3</td>
<td>97.13</td>
<td>97.97</td>
<td>96.59</td>
<td>97.26</td>
</tr>
<tr>
<td></td>
<td>Remzan et al. [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>Radimagenet pre-trained CNN</td>
<td>97.71</td>
<td>97.71</td>
<td>97.71</td>
<td>97.71</td>
</tr>
<tr>
<td></td>
<td>Raouf et al. [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>DCST &#x002B; SVM</td>
<td>97.71</td>
<td>97.80</td>
<td>96.62</td>
<td>97.21</td>
</tr>
<tr>
<td></td>
<td><bold>Ours</bold></td>
<td><bold>ParMamba</bold></td>
<td><bold>99.62</bold></td>
<td><bold>99.58</bold></td>
<td><bold>99.59</bold></td>
<td><bold>99.59</bold></td>
</tr>
<tr>
<td>Dataset 2</td>
<td>Ayadi et al. [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>Deep CNN</td>
<td>94.74</td>
<td>94.03</td>
<td>94.39</td>
<td>94.19</td>
</tr>
<tr>
<td></td>
<td>Chaki et al. [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>BTSCNet</td>
<td>96.67</td>
<td>93.00</td>
<td>95.03</td>
<td>94.00</td>
</tr>
<tr>
<td></td>
<td>Kakarla et al. [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>AP-CNN</td>
<td>97.42</td>
<td>97.41</td>
<td>97.42</td>
<td>97.41</td>
</tr>
<tr>
<td></td>
<td>Zhu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>MEEDNets</td>
<td>98.63</td>
<td>98.47</td>
<td>98.49</td>
<td>98.48</td>
</tr>
<tr>
<td></td>
<td>Wang et al. [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>RanMerFormer</td>
<td>98.86</td>
<td>98.87</td>
<td>98.46</td>
<td>99.39</td>
</tr>
<tr>
<td></td>
<td><bold>Ours</bold></td>
<td><bold>ParMamba</bold></td>
<td><bold>99.35</bold></td>
<td><bold>99.19</bold></td>
<td><bold>99.35</bold></td>
<td><bold>99.27</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To further validate ParMamba, this section conducts significance analyses of ParMamba and other models on dataset 1 and dataset 2. Nemenyi test is first conducted, with the results presented in <xref ref-type="table" rid="table-7">Tables 7</xref> and <xref ref-type="table" rid="table-8">8</xref>. On dataset 1, the <italic>p</italic>-values between ParMamba and both Dense CNN and SSBTCNet are less than 0.05, indicating significant differences. Similarly, on dataset 2, the <italic>p</italic>-values between ParMamba and both DeepCNN and BTSCNets are also less than 0.05, confirming significant differences. In addition, <italic>t</italic>-test are performed to explore the significance of the differences between ParMamba and other models further, as shown in <xref ref-type="table" rid="table-9">Tables 9</xref> and <xref ref-type="table" rid="table-10">10</xref>. For both dataset 1 and dataset 2, the normality <italic>p</italic>-values for the differences between ParMamba and other models, except for SSBTCNet, are greater than 0.05, suggesting that these differences follow a normal distribution. The <italic>t</italic>-test <italic>p</italic>-values are all less than 0.05, demonstrating that ParMamba exhibits statistically significant differences compared to other models. These results indicate that ParMamba significantly outperforms other models across both datasets.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Nemenyi test on dataset 1</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Models</th>
<th>ParMamba</th>
<th>Dense CNN</th>
<th>SSBTCNet</th>
<th>InceptionV3</th>
<th>Pre-trained CNN</th>
<th>DCST &#x002B; SVM</th>
</tr>
</thead>
<tbody>
<tr>
<td>ParMamba</td>
<td>1.0000</td>
<td>0.0166</td>
<td>0.0045</td>
<td>0.5266</td>
<td>0.7995</td>
<td>0.5812</td>
</tr>
<tr>
<td>Dense CNN</td>
<td>0.0166</td>
<td>1.0000</td>
<td>0.9000</td>
<td>0.6358</td>
<td>0.3517</td>
<td>0.5812</td>
</tr>
<tr>
<td>SSBTCNet</td>
<td>0.0045</td>
<td>0.9000</td>
<td>1.0000</td>
<td>0.4102</td>
<td>0.1698</td>
<td>0.3517</td>
</tr>
<tr>
<td>InceptionV3</td>
<td>0.5266</td>
<td>0.6358</td>
<td>0.4102</td>
<td>1.0000</td>
<td>0.9000</td>
<td>0.9000</td>
</tr>
<tr>
<td>Pre-trained CNN</td>
<td>0.7995</td>
<td>0.3517</td>
<td>0.1698</td>
<td>0.9000</td>
<td>1.0000</td>
<td>0.9000</td>
</tr>
<tr>
<td>DCST &#x002B; SVM</td>
<td>0.5812</td>
<td>0.5812</td>
<td>0.3517</td>
<td>0.9000</td>
<td>0.9000</td>
<td>1.0000</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Nemenyi test on dataset 2</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Models</th>
<th>ParMamba</th>
<th>DeepCNN</th>
<th>BTSCNet</th>
<th>AP-CNN</th>
<th>MEEDNets</th>
<th>RanMerFormer</th>
</tr>
</thead>
<tbody>
<tr>
<td>ParMamba</td>
<td>1.0000</td>
<td>0.0166</td>
<td>0.0166</td>
<td>0.2984</td>
<td>0.8541</td>
<td>0.9000</td>
</tr>
<tr>
<td>DeepCNN</td>
<td>0.0166</td>
<td>1.0000</td>
<td>0.9000</td>
<td>0.8541</td>
<td>0.2984</td>
<td>0.0868</td>
</tr>
<tr>
<td>BTSCNet</td>
<td>0.0166</td>
<td>0.9000</td>
<td>1.0000</td>
<td>0.8541</td>
<td>0.2984</td>
<td>0.0868</td>
</tr>
<tr>
<td>AP-CNN</td>
<td>0.2984</td>
<td>0.8541</td>
<td>0.8541</td>
<td>1.0000</td>
<td>0.9000</td>
<td>0.6358</td>
</tr>
<tr>
<td>MEEDNets</td>
<td>0.8541</td>
<td>0.2984</td>
<td>0.2984</td>
<td>0.9000</td>
<td>1.0000</td>
<td>0.9000</td>
</tr>
<tr>
<td>RanMerFormer</td>
<td>0.9000</td>
<td>0.0868</td>
<td>0.0868</td>
<td>0.6358</td>
<td>0.9000</td>
<td>1.0000</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title><italic>t</italic>-test on dataset 1</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Ours</th>
<th>Others</th>
<th>Normality <italic>p</italic>-value</th>
<th><italic>t</italic>-test <italic>p</italic>-value</th>
</tr>
</thead>
<tbody>
<tr>
<td>ParMamba</td>
<td>Dense CNN</td>
<td>0.3927</td>
<td>1.391e&#x2212;03</td>
</tr>
<tr>
<td></td>
<td>SSBTCNet</td>
<td>0.0346</td>
<td>&#x2013;</td>
</tr>
<tr>
<td></td>
<td>InceptionV3</td>
<td>0.8661</td>
<td>3.784e&#x2212;03</td>
</tr>
<tr>
<td></td>
<td>Pre-trained CNN</td>
<td>0.1945</td>
<td>2.138e&#x2212;07</td>
</tr>
<tr>
<td></td>
<td>DCST &#x002B; SVM</td>
<td>0.5445</td>
<td>3.555e&#x2212;03</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title><italic>t</italic>-test on dataset 2</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Ours</th>
<th>Others</th>
<th>Normality <italic>p</italic>-value</th>
<th><italic>t</italic>-test <italic>p</italic>-value</th>
</tr>
</thead>
<tbody>
<tr>
<td>ParMamba</td>
<td>DeepCNN</td>
<td>0.4133</td>
<td>3.2e&#x2212;5</td>
</tr>
<tr>
<td></td>
<td>BTSCNet</td>
<td>0.9024</td>
<td>8.62e&#x2212;3</td>
</tr>
<tr>
<td></td>
<td>AP-CNN</td>
<td>0.2616</td>
<td>1.5e&#x2212;5</td>
</tr>
<tr>
<td></td>
<td>MEEDNets</td>
<td>0.2725</td>
<td>1.79e&#x2212;4</td>
</tr>
<tr>
<td></td>
<td>RanMerFormer</td>
<td>0.9781</td>
<td>1.553e&#x2212;1</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>This paper proposes a novel parallel architecture named ParMamba for brain tumor classification. The ParMamba consists of four components: ConvBlock, MambaBlock, channel enhancement module, and CAPE. Among them, ConvBlock and MambaBlock extract local and global features respectively, channel enhancement module implements cross-channel communication, and CAPE is used for downsampling, effectively extracting fine-grained features of brain tumors from local and global contexts. Extensive experiments are conducted on two publicly available brain tumor datasets. Dataset 1 contains four categories of images: glioma, meningioma, pituitary, and non-tumor, while dataset 2 comprises three categories: glioma, meningioma, and pituitary. The experimental results demonstrate that the proposed ParMamba achieves outstanding classification performance, with accuracies of 99.62% and 99.35% on dataset 1 and dataset 2, respectively. This indicates that the model is capable of achieving excellent classification performance on both datasets. Compared to existing brain tumor classification methods, ParMamba marks a significant advancement in the field. Moreover, the accuracy of Mamba increased by 1.37% compared with ViT, while FLOPs decreased by 55%, and throughput rate also increased by more than 30% compared with ViT. This indicates that the proposed ParMamba is not only faster but also more accurate than attention mechanisms. Furthermore, ablation studies are conducted to validate the effectiveness of CAPE, MambaBlock, and channel enhancement module individually. The results reveal that the combination of CAPE, MambaBlock, and channel enhancement module effectively extracts brain tumor features, leading to excellent classification results.</p>
<p>While the proposed ParMamba model demonstrates strong overall performance, its recognition of meningioma remains relatively limited. Therefore, future efforts could focus on refining the model to better capture meningioma features. Moreover, future research could extend the application scope of ParMamba by testing its generalization performance on large-scale brain tumor datasets. Further potential directions include developing lightweight network designs for efficient deployment on resource-constrained devices.</p>
</sec>
</body>
<back>
<ack>
<p>The authors wish to express their appreciation to the reviewers for their helpful suggestions which greatly improved the presentation of this paper.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the Outstanding Youth Science and Technology Innovation Team Project of Colleges and Universities in Hubei Province (Grant no. T201923), Key Science and Technology Project of Jingmen (Grant nos. 2021ZDYF024, 2022ZDYF019), and Cultivation Project of Jingchu University of Technology (Grant no. PY201904).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Study conception and design: Gaoshuai Su, Hongyang Li; data collection: Gaoshuai Su; analysis and interpretation of results: Gaoshuai Su, Hongyang Li; draft manuscript preparation: Gaoshuai Su, Huafeng Chen. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are openly available at <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset">https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset</ext-link>, <ext-link ext-link-type="uri" xlink:href="https://figshare.com/articles/dataset/brain_tumor_dataset/1512427/5">https://figshare.com/articles/dataset/brain_tumor_dataset/1512427/5</ext-link> (accessed on 16 November 2024).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Muhammad</surname> <given-names>K</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ser</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Albuquerque</surname> <given-names>VH</given-names></string-name></person-group>. <article-title>Deep learning for multigrade brain tumor classification in smart healthcare systems: a prospective survey</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2020</year>;<volume>32</volume>(<issue>2</issue>):<fpage>507</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2020.2995800</pub-id>; <pub-id pub-id-type="pmid">32603291</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tiwari</surname> <given-names>A</given-names></string-name>, <string-name><surname>Srivastava</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pant</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Brain tumor segmentation and classification from magnetic resonance images: review of selected methods from 2014 to 2019</article-title>. <source>Pattern Recognit Lett</source>. <year>2020</year>;<volume>131</volume>(<issue>9</issue>):<fpage>244</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patrec.2019.11.020</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rajesh</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mani Malar</surname> <given-names>RS</given-names></string-name>, <string-name><surname>Geetha</surname> <given-names>MR</given-names></string-name></person-group>. <article-title>Brain tumor detection using optimization classification based on rough set theory</article-title>. <source>Cluster Comput</source>. <year>2019</year>;<volume>22</volume>(<issue>Suppl 6</issue>):<fpage>13853</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10586-018-2111-5</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Saurav</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>A</given-names></string-name>, <string-name><surname>Saini</surname> <given-names>R</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>S</given-names></string-name></person-group>. <article-title>An attention-guided convolutional neural network for auto-mated classification of brain tumor from MRI</article-title>. <source>Neural Comput Appl</source>. <year>2023</year>;<volume>35</volume>(<issue>3</issue>):<fpage>2541</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-022-07742-z</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shaik</surname> <given-names>NS</given-names></string-name>, <string-name><surname>Cherukuri</surname> <given-names>TK</given-names></string-name></person-group>. <article-title>Multi-level attention network: application to brain tumor classification</article-title>. <source>Signal Image Video Process</source>. <year>2022</year>;<volume>16</volume>(<issue>3</issue>):<fpage>817</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11760-021-02022-0</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Apostolopoulos</surname> <given-names>ID</given-names></string-name>, <string-name><surname>Aznaouridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tzani</surname> <given-names>M</given-names></string-name></person-group>. <article-title>An attention-based deep convolutional neural network for brain tumor and disorder classification and grading in magnetic resonance imaging</article-title>. <source>Information</source>. <year>2023</year>;<volume>14</volume>(<issue>3</issue>):<fpage>174</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info14030174</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Beyer</surname> <given-names>L</given-names></string-name>, <string-name><surname>Kolesnikov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Weissenborn</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>X</given-names></string-name></person-group>. <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>. In: <conf-name>ICLR 2021 Conference</conf-name>; <year>2020</year>. p. <fpage>1</fpage>&#x2013;<lpage>18</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dao</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Mamba: Linear-time sequence modeling with selective state spaces</article-title>. <comment>arXiv:2312.00752. 2023</comment>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Vision mamba: efficient visual representation learning with bidirectional state space model</article-title>. <comment>arXiv:2401.09417. 2024</comment>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Vmamba: visual state space model</article-title>. <comment>arXiv:2401.10166. 2024</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>El-Dahshan</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Hosny</surname> <given-names>T</given-names></string-name>, <string-name><surname>Salem</surname> <given-names>ABM</given-names></string-name></person-group>. <article-title>Hybrid intelligent techniques for MRI brain images classification</article-title>. <source>Digit Signal Process</source>. <year>2010</year>;<volume>20</volume>(<issue>2</issue>):<fpage>433</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dsp.2009.07.002</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shim</surname> <given-names>VB</given-names></string-name>, <string-name><surname>Holdsworth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Champagne</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Coverdale</surname> <given-names>NS</given-names></string-name>, <string-name><surname>Cook</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>TR</given-names></string-name></person-group>. <article-title>Rapid prediction of brain injury pattern in mTBI by combining FE analysis with a machine-learning based approach</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>179457</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2020.3026350</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Das</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chowdhury</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kundu</surname> <given-names>MK</given-names></string-name></person-group>. <article-title>Brain MR image classification using multiscale geometric analysis of ripplet</article-title>. <source>Prog Electromagn Res</source>. <year>2013</year>;<volume>137</volume>:<fpage>1</fpage>&#x2013;<lpage>17</lpage>. doi:<pub-id pub-id-type="doi">10.2528/PIER13010105</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>YD</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Magnetic resonance brain image classification by an improved artificial bee colony algorithm</article-title>. <source>Prog Electromagn Res</source>. <year>2011</year>;<volume>116</volume>:<fpage>65</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.2528/PIER11031709</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shim</surname> <given-names>V</given-names></string-name>, <string-name><surname>Tayebi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kwon</surname> <given-names>E</given-names></string-name>, <string-name><surname>Guild</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>Scadeng</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dubowitz</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Combining advanced magnetic resonance imaging (MRI) with finite element (FE) analysis for characterising subject-specific injury patterns in the brain after traumatic brain injury</article-title>. <source>Eng Comput</source>. <year>2022</year>;<volume>38</volume>(<issue>5</issue>):<fpage>3925</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00366-022-01697-4</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ayadi</surname> <given-names>W</given-names></string-name>, <string-name><surname>Elhamzi</surname> <given-names>W</given-names></string-name>, <string-name><surname>Charfi</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Deep CNN for brain tumor classification</article-title>. <source>Neural Process Lett</source>. <year>2021</year>;<volume>53</volume>(<issue>1</issue>):<fpage>671</fpage>&#x2013;<lpage>700</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11063-020-10398-2</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Atha</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chaki</surname> <given-names>J</given-names></string-name></person-group>. <article-title>SSBTCNet: semi-supervised brain tumor classification network</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>141485</fpage>&#x2013;<lpage>99</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2023.3343126</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rizwan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Javed</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Shabbir</surname> <given-names>A</given-names></string-name>, <string-name><surname>Baker</surname> <given-names>T</given-names></string-name>, <string-name><surname>Obe</surname> <given-names>DA</given-names></string-name></person-group>. <article-title>Brain tumor and glioma grade classification using Gaussian convolutional neural network</article-title>. <source>IEEE Access</source>. <year>2022</year>;<volume>10</volume>:<fpage>29731</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2022.3153108</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ulidowski</surname> <given-names>I</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MEEDNets: medical image classification via ensemble bio-inspired evolutionary DenseNets</article-title>. <source>Knowl-Based Syst</source>. <year>2023</year>;<volume>280</volume>(<issue>6</issue>):<fpage>111035</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2023.111035</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aamir</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Dayo</surname> <given-names>ZA</given-names></string-name>, <string-name><surname>Abro</surname> <given-names>WA</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>MI</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>I</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A deep learning approach for brain tumor classification using MRI images</article-title>. <source>Comput Electr Eng</source>. <year>2022</year>;<volume>101</volume>(<issue>5</issue>):<fpage>108105</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compeleceng.2022.108105</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>RL</given-names></string-name>, <string-name><surname>Kakarla</surname> <given-names>J</given-names></string-name>, <string-name><surname>Isunuri</surname> <given-names>BV</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Multi-class brain tumor classification using residual network and global average pooling</article-title>. <source>Multimed Tools Appl</source>. <year>2021</year>;<volume>80</volume>(<issue>9</issue>):<fpage>13429</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-020-10335-4</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Grsoy</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kaya</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Brain-GCN-Net: graph-convolutional neural network for brain tumor identification</article-title>. <source>Comput Biol Med</source>. <year>2024</year>;<volume>180</volume>(<issue>1</issue>):<fpage>108971</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiomed.2024.108971</pub-id>; <pub-id pub-id-type="pmid">39106672</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yaqub</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jinchao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mehmood</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chuhan</surname> <given-names>IS</given-names></string-name>, <string-name><surname>Manan</surname> <given-names>MA</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DeepLabV3, IBCO-based ALCResNet: a fully automated classification, and grading system for brain tumor</article-title>. <source>Alexandria Eng J</source>. <year>2023</year>;<volume>76</volume>(<issue>6</issue>):<fpage>609</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aej.2023.06.062</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ghassemi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Shoeibi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rouhani</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deep neural network with generative adversarial networks pre-training for brain tumor classification based on MR images</article-title>. <source>Biomed Signal Process Control</source>. <year>2020</year>;<volume>57</volume>(<issue>2</issue>):<fpage>101678</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2019.101678</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Su</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A multi-category brain tumor classification method based on improved ResNet50</article-title>. <source>Comput Mater Contin</source>. <year>2021</year>;<volume>69</volume>(<issue>2</issue>):<fpage>2355</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2021.019409</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dutta</surname> <given-names>TK</given-names></string-name>, <string-name><surname>Nayak</surname> <given-names>DR</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>YD</given-names></string-name></person-group>. <article-title>ARM-Net: attention-guided residual multiscale CNN for multiclass brain tumor classification using MR images</article-title>. <source>Biomed Signal Process Control</source>. <year>2024</year>;<volume>87</volume>(<issue>2</issue>):<fpage>105421</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2023.105421</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>SY</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>YD</given-names></string-name></person-group>. <article-title>RanMerFormer: randomized vision transformer with token merging for brain tumor classification</article-title>. <source>Neurocomputing</source>. <year>2024</year>;<volume>573</volume>(<issue>20</issue>):<fpage>127216</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2023.127216</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Isunuri</surname> <given-names>BV</given-names></string-name>, <string-name><surname>Kakarla</surname> <given-names>J</given-names></string-name></person-group>. <article-title>EfficientNet and multi-path convolution with multi-head attention network for brain tumor grade classification</article-title>. <source>Comput Electr Eng</source>. <year>2023</year>;<volume>108</volume>(<issue>11</issue>):<fpage>108700</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compeleceng.2023.108700</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yue</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Medmamba: vision mamba for medical image classification</article-title>. <comment>arXiv:2403.03849. 2024</comment>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>U-mamba: enhancing long-range dependency for biomedical image segmentation</article-title>. <comment>arXiv:2401.04722. 2024</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ruan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Vm-unet: vision mamba unet for medical image segmentation</article-title>. <comment>arXiv:2402.02491. 2024</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Setyawan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kurniawan</surname> <given-names>GW</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Hsieh</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Su</surname> <given-names>HK</given-names></string-name>, <string-name><surname>Kuo</surname> <given-names>WK</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>ParFormer: vision Transformer baseline with parallel local global token mixer and convolution attention patch embedding</article-title>. <comment>arXiv:2403.15004. 2024</comment>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Feichtenhofer</surname> <given-names>C</given-names></string-name>, <string-name><surname>Darrell</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>A convnet for the 2020s</article-title>. <source>Proc IEEE/CVF Conf Comput Vis Pattern Recognit</source>, <volume>2022</volume>:<fpage>11976</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01167</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Howard</surname> <given-names>AG</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kalenichenko</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Weyand</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Mobilenets: efficient convolutional neural networks for mobile vision applications</article-title>. <comment>arXiv:1704.04861. 2017</comment>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goel</surname> <given-names>K</given-names></string-name>, <string-name><surname>R.</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Efficiently modeling long sequences with structured state spaces</article-title>. <comment>arXiv:2111.00396. 2021</comment>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>&#x00D6;zkaraca</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ba&#x011F;r&#x0131;a&#x00E7;&#x0131;k</surname> <given-names>O</given-names></string-name>, <string-name><surname>G&#x00FC;r&#x00FC;ler</surname> <given-names>H</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>F</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>J</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Multiple brain tumor classification with dense CNN architecture using brain MRI images</article-title>. <source>Life</source>. <year>2023</year>;<volume>13</volume>(<issue>2</issue>):<fpage>349</fpage>. doi:<pub-id pub-id-type="doi">10.3390/life13020349</pub-id>; <pub-id pub-id-type="pmid">36836705</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>G&#x00F3;mez-Guzm&#x00E1;n</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Jim&#x00E9;nez-Berista&#x00ED;n</surname> <given-names>L</given-names></string-name>, <string-name><surname>Garc&#x00ED;a-Guerrero</surname> <given-names>EE</given-names></string-name>, <string-name><surname>L&#x00F3;pez-Bonilla</surname> <given-names>OR</given-names></string-name>, <string-name><surname>Tamayo-Perez</surname> <given-names>UJ</given-names></string-name>, <string-name><surname>Esqueda-Elizondo</surname> <given-names>JJ</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Classifying brain tumors on magnetic resonance imaging by using convolutional neural networks</article-title>. <source>Electronics</source>. <year>2023</year>;<volume>12</volume>(<issue>4</issue>):<fpage>955</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics12040955</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Remzan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Tahiry</surname> <given-names>K</given-names></string-name>, <string-name><surname>Farchi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Advancing brain tumor classification accuracy through deep learning: harnessing radimagenet pre-trained convolutional neural networks, ensemble learning, and machine learning classifiers on MRI brain images</article-title>. <source>Multimed Tools Appl</source>. <year>2024</year>;<fpage>1</fpage>&#x2013;<lpage>29</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-024-18780-1</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Raouf</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Fallah</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rashidi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Use of discrete cosine-based stockwell transform in the binary classification of magnetic resonance images of brain tumor</article-title>. In: <conf-name>2022 29th National and 7th International Iranian Conference on Biomedical Engineering (ICBME)</conf-name>; <year>2022 Dec</year>; <publisher-loc>Tehran, Iran</publisher-loc>: <publisher-name>IEEE</publisher-name>; p. <fpage>293&#x2013;8</fpage>, <comment>2022</comment></mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chaki</surname> <given-names>J</given-names></string-name>, <string-name><surname>Woniak</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A deep learning based four-fold approach to classify brain MRI: bTSCNet</article-title>. <source>Biomed Signal Process Control</source>. <year>2023</year>;<volume>85</volume>(<issue>2</issue>):<fpage>104902</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2023.104902</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kakarla</surname> <given-names>J</given-names></string-name>, <string-name><surname>Isunuri</surname> <given-names>BV</given-names></string-name>, <string-name><surname>Doppalapudi</surname> <given-names>KS</given-names></string-name>, <string-name><surname>Bylapudi</surname> <given-names>KSR</given-names></string-name></person-group>. <article-title>Three-class classification of brain magnetic resonance images using average-pooling convolutional neural network</article-title>. <source>Int J Imaging Syst Technol</source>. <year>2021</year>;<volume>31</volume>(<issue>3</issue>):<fpage>1731</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1002/ima.22554</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>