<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">80864</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.080864</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>GreenShield: A Lightweight and Robust Vision Transformer Framework in Retinal Disease Classification</article-title>
<alt-title alt-title-type="left-running-head">GreenShield: A Lightweight and Robust Vision Transformer Framework in Retinal Disease Classification</alt-title>
<alt-title alt-title-type="right-running-head">GreenShield: A Lightweight and Robust Vision Transformer Framework in Retinal Disease Classification</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Qasaimeh</surname><given-names>Munthir</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Ali</surname><given-names>Mostafa</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Abu Al-Haija</surname><given-names>Qasem</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>qsabuhaija@just.edu.jo</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Information Systems, Jordan University of Science and Technology</institution>, <addr-line>Irbid</addr-line>, <country>Jordan</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Cybersecurity, Jordan University of Science and Technology</institution>, <addr-line>Irbid</addr-line>, <country>Jordan</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Qasem Abu Al-Haija. Email: <email>qsabuhaija@just.edu.jo</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>5</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>2</issue>
<elocation-id>42</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>02</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_80864.pdf"></self-uri>
<abstract>
<p>Vision Transformers (ViTs) have recently achieved high performance in retinal Optical Coherence Tomography (OCT) classification studies. However, ViT models continue to face significant challenges, including high computational cost, vulnerability to adversarial attacks, and pronounced sensitivity to preprocessing techniques. This study introduces GreenShield, a unified framework designed to produce an efficient and robust ViT model, referred to as GreenShield-ViT, which outperforms existing lightweight ViT variants in terms of adversarial robustness for retinal OCT classification. The framework integrates a gradient-based block-importance pruning strategy to compress the ViT/B-16 architecture, and adversarial training with proper ImageNet normalization and anti-saturation techniques. The robustness was evaluated using FGSM, PGD, PGD-R3, Transfer-PGD, BIM, and the proposed hybrid attack (FGSM-PGD). The proposed approach achieves an approximately 50% reduction in Floating-Point Operations (FLOPs), inference time, and carbon footprint emissions, while preserving diagnostic accuracy. Experiments conducted using GPU P100 on the OCT-c8, OCTID, and UCSD-3 datasets achieved clean accuracies of 92.5%, 94.78%, and 89.20%, respectively, alongside a significant reduction in attack success rates and improved model calibration. GreenShield-ViT outperformed lightweight ViT variants (Mobile-ViT, ViT-Tiny, ViT-Small) in terms of robustness while offering competitive efficiency. These results suggest its applicability to similar ViT-based medical tasks.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Retinal OCT classification</kwd>
<kwd>green AI</kwd>
<kwd>adversarial training</kwd>
<kwd>vision transformers</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Retinal diseases are a major cause of blindness. Many patients lose vision early because of their retinal conditions. Optical Coherence Tomography (OCT) was performed in 1991 [<xref ref-type="bibr" rid="ref-1">1</xref>]. This provides the capability to capture retinal layers and deeper details with high-resolution imaging in a noninvasive manner. Despite the capabilities of photography, ophthalmologists still face challenges in their diagnosis. For example, when facing an overlapping disease pattern, manual interpretation is time-consuming and susceptible to human error. All of these challenges were addressed at the beginning of the Artificial Intelligence (AI) era. Deep learning models such as Convolutional Neural Networks (CNNs), which enable rapid and accurate analysis of large OCT volumes, support ophthalmologists in enhancing diagnostic confidence.</p>
<p>In 2020, the Vision Transformer models (ViTs) were introduced as an alternative to CNNs [<xref ref-type="bibr" rid="ref-2">2</xref>] for use in image-processing tasks. Recent retinal OCT studies have demonstrated that ViTs outperform CNNs in terms of performance, while CNNs are more efficient than ViTs [<xref ref-type="bibr" rid="ref-3">3</xref>]. However, both models are vulnerable to adversarial attacks. AI models in retinal OCT applications face several challenges, including security concerns, sustainability, and reliable preprocessing pipelines [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. Security challenges include adversarial attacks, which can apply slight changes to an image to fool AI models [<xref ref-type="bibr" rid="ref-6">6</xref>]. This may cause misdiagnosis in the medical field, which could be a primary cause of disasters. By contrast, hospitals and clinics often lack access to high-performance servers that can deploy computationally intensive models. Moreover, inference time is a critical factor, especially in cases that must handle a large patient volume daily. Reducing model complexity decreases energy consumption (carbon emissions) and makes deployment more eco-friendly, contributing to the advancement of green AI. A list of lightweight ViT architectures was introduced in retinal OCT studies to address efficiency challenges without security enhancement considerations [<xref ref-type="bibr" rid="ref-7">7</xref>]. Moreover, prior OCT studies, which focused on security, ignored preprocessing challenges, such as softmax saturation and normalization space [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>ViT models require ImageNet-consistent normalization owing to their ImageNet pretraining [<xref ref-type="bibr" rid="ref-11">11</xref>]. The conflict between the adversarial attack space (0, 1-pixel space) and the normalization space (ImageNet normalization space) changes the attack&#x2019;s power and harms adversarial training. Additionally, in adversarial work, handling Softmax saturation is necessary [<xref ref-type="bibr" rid="ref-12">12</xref>]. The ViT model splits the images into patches. When the model focuses on one patch and ignores the other patches, one patch becomes dominant (the highest attention score), and SoftMax saturates. After softmax saturation, the model will be overconfident. Slight changes in the dominant patch lead to misclassification with a high confidence rate [<xref ref-type="bibr" rid="ref-12">12</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>]. All of these challenges motivated us to design a robust and lightweight ViT model with a reliable preprocessing pipeline.</p>
<p>Many OCT classification studies have been limited to binary classification or a limited number of classes. In adversarial work, the majority of studies have focused on CNNs (ignoring ViTs). Studies that enhanced ViT efficiency did not introduce a compression technique. Several studies on improving OCT security in the literature have ignored sustainability, while studies that focused on improving efficiency have ignored security challenges. Limited retinal OCT studies have addressed the impact of preprocessing challenges such as softmax saturation and normalization in adversarial training. However, this study aims to fill these research gaps by making the following contributions:<list list-type="bullet">
<list-item>
<p>This study proposed an empirically validated framework that integrates the gradient-based block-importance pruning with adversarial training for retinal OCT classification, jointly enhancing sustainability and robustness against attacks.</p></list-item>
<list-item>
<p>Builds a lightweight model that outperforms existing lightweight ViT variants (Mobile-ViT, ViT-Small, and ViT-Tiny) in robustness against White-box L-infinity adversarial attacks (FGSM, PGD, BIM, and a Hybrid attack).</p></list-item>
<list-item>
<p>Enhances the perturbation generation in adversarial training by addressing the normalization issue in the prior OCT work.</p></list-item>
<list-item>
<p>Improves the Model calibration (to avoid softmax saturation).</p></list-item>
<list-item>
<p>Conducts a comparative adversarial robustness analysis of existing lightweight ViT variants (Mobile-ViT, ViT-Small, and ViT-Tiny).</p></list-item>
</list></p>
<p>This paper is structured as follows: <xref ref-type="sec" rid="s2">Section 2</xref> provides a background that covers related concepts and findings from related work. <xref ref-type="sec" rid="s3">Section 3</xref> describes the proposed methodology. <xref ref-type="sec" rid="s4">Section 4</xref> presents experiments and results. <xref ref-type="sec" rid="s5">Section 5</xref> presents the discussion. <xref ref-type="sec" rid="s6">Section 6</xref> provides the limitations and future directions. Finally, <xref ref-type="sec" rid="s7">Section 7</xref> presents our conclusions.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>This section offers a comprehensive review of the relevant literature, emphasizing its findings, scholarly contributions, and identified research gaps. Related work is systematically organized into three sections: CNN-based, transformer-based, and hybrid CNN-transformer-based approaches.</p>
<sec id="s2_1">
<label>2.1</label>
<title>CNN-Based Approaches</title>
<p>This section covers studies that have utilized CNNs to classify retinal images with different objectives. Numerous studies in existing literature have concentrated on enhancing robustness. For example, the work done in [<xref ref-type="bibr" rid="ref-15">15</xref>], the authors built a robust binary classification model to detect diabetic retinopathy (DR). The Indian IDRiD and OCT datasets were used. The study applied hold-out validation (15,400 training, 3720 validations, and 400 testing sets). The ResNet-50 model with pre-trained weights (ImageNet) was fine-tuned to the proposed datasets. Adversarial training and distillation were performed to address the FGSM and L-BFGS attacks. The distillation technique outperformed adversarial training, reducing the success rates of both attacks (FGSM and L-BFGS-B by 60% and 82%, respectively).</p>
<p>In contrast, the authors of [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed a light, robust CNN model to address the sustainability and security of retinal OCT classification. The study used Kermany OCT, COVID, and Kidney Stone CT datasets. An under-sampling technique was used to balance the data. Hold-out validation was applied (80% training and 20% for testing). Cross-validation with 10 folds was applied to the training set. The study designed a custom lightweight CNN model. No defense mechanisms were applied. The model was evaluated using four adversarial FGSM examples. Achieving 94% clean accuracy on the Kermany dataset.</p>
<p>Furthermore, some retinal OCT studies have utilized CNN with a focus on accuracy enhancements. For example, in [<xref ref-type="bibr" rid="ref-17">17</xref>], the study aimed to develop a system for detecting retinal diseases using OCT images. This study proposes two models for feature extraction (ResNet-50 and AlexNet) and uses the Kermany dataset. The proposed CNN network is used in parallel to extract the features. Two feature selection techniques were applied in parallel (PCA and Entropy-based Ant Colony System). The selected features were combined into an optimized vector and then fed into Machine Learning models, such as K-Nearest Neighbors(kNN), Linear Support Vector Machine (LSVM), Linear Discriminant (LD), and Decision trees. The LD classifier achieved high accuracy (99.9%).</p>
<p>Similarly, the authors in [<xref ref-type="bibr" rid="ref-18">18</xref>] addressed the challenge of multi-class accuracy (retinal disease classification) under a limited and imbalanced dataset, proposing a modified VGG16-based CNN with data augmentation, class weighting, and explainability techniques (Grad-CAM) using a small Mendeley fundus dataset of 302 images. The model achieved 93.42% training accuracy and 77.5% validation accuracy, with improved interpretability and class balance.</p>
<p>All these studies were limited to CNNs and did not address the improvement in the robustness and efficiency of ViT models.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Transformer-Based Approaches</title>
<p>This section covers studies in the existing literature that utilize Vision Transformers as the main model. Most studies that employed Transformers in Retinal OCT classification focused on efficiency. For example, the authors in [<xref ref-type="bibr" rid="ref-10">10</xref>] designed a light approach for Diabetic Retinopathy Classification. This study proposed resizing the images to 896 &#x00D7; 896 pixels and splitting them into 16 patches (224 &#x00D7; 224 pixels). This study proposes a ViT-small model as a feature extractor. The Classification tokens (CLS) were combined into a matrix (16 &#x00D7; 384) and fed into the Global Instance Computing Block (GICB) layer to compute attention scores, then fed to a Multi-Layer Perceptron (MLP) layer with a residual. The proposed approach was applied to APTOS 2019 and Messidor-1 datasets. The inference time is reduced by 62%, and the accuracy is improved by 2.1% for the APTOS and 12.1% for the Messidor dataset.</p>
<p>Extending the focus to efficiency, the authors in [<xref ref-type="bibr" rid="ref-19">19</xref>] evaluated a list of ViT models (ViT, T2T ViT, and Mobile-ViT) for retinal OCT classification. This study proposed mobile-ViT as the main approach. The public OCT and Mendeley datasets were used. Hold-out validation was applied to propose 10% as the validation dataset. Mobile-ViT outperformed all the models by achieving 99.1% accuracy with better efficiency.</p>
<p>Some studies have employed ViTs without considering sustainability, such as the authors in [<xref ref-type="bibr" rid="ref-9">9</xref>], who compared two Vision Transformer variants (ViT-16 and ViT-32) in the binary classification of retinal OCT images. The ViT models were trained to detect Diabetic Macular Edema (DME), using three optimizers: Adam, Stochastic Gradient Descent (SGD), and Root Mean Square Propagation (RMSProp). Two balanced datasets were proposed: KMC and Mendeley Datasets. A custom classification head was added. The GRAD-CAM was used for visualization. ViT-16 outperformed all the baselines across both datasets. The Adam optimizer achieved the best results, with 100% recall and stable training.</p>
<p>However, a limited number of OCT studies focused on ViT robustness against adversarial attacks, such as in [<xref ref-type="bibr" rid="ref-8">8</xref>]. This study proposed adversarial training as a defense mechanism. Exploring the performance using a binary labeled list of medical datasets: Fundoscopy, Chest X-ray, ISIC2019, Breast Histology. Images were normalized using ImageNet normalization (during the preprocessing phase). The proposed model is a pretrained ViT&#x2013;Tiny model. The study employed hold-out validation (85% of the training set and 15% of the testing set). All models were evaluated against multiple attacks (FGSM, BIM, PGD, and CW) using these hyperparameters (&#x03B5; varied between 0.25/255 and 5/255 for FGSM, PGD-20 steps, and BIM-40 steps). The ViT model outperformed all CNNs on balanced datasets in terms of both clean accuracy and adversarial robustness.</p>
<p>The authors of [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>] proposed lightweight models without evaluating their robustness to attacks. In [<xref ref-type="bibr" rid="ref-9">9</xref>], ViT variants were compared based on performance, without considering robustness against attacks. The study in [<xref ref-type="bibr" rid="ref-8">8</xref>] focused on robustness enhancement but utilized an imbalanced dataset and applied adversarial training without considering proper ImageNet normalization.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Hybrid CNN-Transformer-Based Approaches</title>
<p>Many studies have employed a hybrid framework combining a CNN with a Transformer to classify retinal OCT images in different research directions. For example, the authors of [<xref ref-type="bibr" rid="ref-20">20</xref>] aimed to enhance efficiency by proposing a hybrid CNN-Vision Transformer architecture on OCT2017, OCT-C8, and an external validation dataset. The model architecture comprises a list of blocks. The Depthwise Convolution DW block uses depth convolution for local lesion scanning. The MobileViT block (MV) is based on Mobile-ViT for learning global patterns. The Convolutional Block Attention Module (CBAM) block focuses on the relevant regions using average pooling, with a final layer for classification. The proposed approach achieved a high performance, but without enhancing the resistance of the model to attacks.</p>
<p>Similarly, the authors of [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed an MGR-GAN to classify retinal OCT images. This hybrid model architecture combines a Transformer with a CNN. The UCSD OCT dataset was used in this study. The study designed a generator (Transformer-based) to generate images, and a discriminator (ResNet) to distinguish between generated and original images. The discriminator is also used for feature extraction and classification with three convolutional layers (F1, F2, and F3) and a softmax layer. The model achieved a 99% accuracy rate. This study did not focus on the security and sustainability of the model.</p>
<p>Similarly, the authors [<xref ref-type="bibr" rid="ref-22">22</xref>] worked to address many problems in the related work, including the need for large, balanced, diverse datasets, and the reduction of health risks in Fluorescein Angiography (FA) imaging that requires dye injection. The study used a large ultra-widefield retinal imaging dataset with 1198 patients, along with Messidor-2 for external evaluation, and proposed a hybrid approach combining a pix2pixHD-based GAN for multi-phase FA image synthesis and a Swin Transformer for diabetic retinopathy classification. The results demonstrated that the generator achieved high realism (SSIM &#x003D; 0.82&#x2013;0.85) with an AUC of 0.910, and an accuracy of 82.9%.</p>
<p>Similarly, the authors of [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed an RAD-IoMT defense mechanism. It is a transformer-based detector that detects adversarial attacks and then passes only clean attacks to the CNN classifier. This study used three datasets (retinal OCT, skin cancer, and chest X-ray). The VGG-16 (standard 138 million parameters) model was trained as a classifier. The transformer was trained using binary classification to detect attacks. The study performed four attacks (white box: PGD and FGSM, and black box: AGN and AUN). The detector achieved an accuracy between 92% and 95%, whereas the classifier achieved an accuracy between 86% and 97%. However, deploying two complex models (detector and classifier) incurs an extra cost (heavy deployment). Additionally, if the attacker fools the filter (detector), the classifier crashes (a Single Point of Failure). <xref ref-type="table" rid="table-1">Table 1</xref> summarizes the results of the previous studies.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of related work.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Study</th>
<th>Objective</th>
<th>Dataset</th>
<th>Methodology</th>
<th>Model</th>
<th>Contributions and Results</th>
<th>Limitation</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>Enhance the Model Robustness</td>
<td>IDRiD Indian &#x0026; OCT dataset on Kaggle</td>
<td>Adversarial Training &#x0026; Distillation</td>
<td>ResNet-50</td>
<td><list list-type="bullet">
<list-item>
<p>Reducing Attack Success Rate by 60%&#x2013;82%</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Limited to Binary Classification</p></list-item>
<list-item>
<p>Limited to CNNs (NO ViTS)</p></list-item>
<list-item>
<p>No efficiency enhancements</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>Enhance Robustness and Efficiency</td>
<td>Kermany OCT dataset</td>
<td>Under-sampling &#x0026; Clean Training (Cross-Validation)</td>
<td>CNN Lightweight Model</td>
<td><list list-type="bullet">
<list-item>
<p>Achieved a high performance</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>No defense Mechanism</p></list-item>
<list-item>
<p>No Model compression mechanism</p></list-item>
<list-item>
<p>Improper evaluation</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>Performance enhancement</td>
<td>Kermany OCT dataset</td>
<td>Parallel feature extraction (Two CNNs), Parallel feature selection, with ML models</td>
<td>ResNet with modified AlexNet for feature extraction and ML classifier (LD)</td>
<td><list list-type="bullet">
<list-item>
<p>Achieved a high performance</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Limited to CNNs</p></list-item>
<list-item>
<p>No efficiency enhancement</p></list-item>
<list-item>
<p>No robustness enhancement</p></list-item>
<list-item>
<p>Did not handle the imbalance in class distribution</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Address data limitations and Imbalanced distribution</td>
<td>Small Mendeley fundus dataset</td>
<td>Augmentation, Class weighting, Grad-CAM, and SHAP</td>
<td>VGG16</td>
<td><list list-type="bullet">
<list-item>
<p>Achieved a high accuracy and addressed the data limitations and challenges</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Did not focus on efficiency or robustness enhancements against attacks</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>Efficiency Enhancement &#x0026; propose a new resizing approach</td>
<td>APTOS 2019 and Messidor-1</td>
<td>Resizing (896 &#x00D7; 896) &#x0026; Data augmentation</td>
<td>ViT-Small</td>
<td><list list-type="bullet">
<list-item>
<p>Built a resizing approach and enhanced the resolution</p></list-item>
<list-item>
<p>Proposed a lightweight model.</p></list-item>
<list-item>
<p>Reduced the inference time by 62%</p></list-item>
<list-item>
<p>Improved the accuracy by 2.1% on APTOS and 12.1% on the Messidor dataset</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>No defense mechanism</p></list-item>
<list-item>
<p>Didn&#x2019;t evaluate the light model robustness.</p></list-item>
<list-item>
<p>No compression mechanism</p></list-item>
<list-item>
<p>Softmax saturation mitigation and Normalization were not discussed</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Efficiency Enhancement</td>
<td>OCT images selected from the Mendeley dataset</td>
<td>Manual Balancing, resizing, and Hold-Out Validation</td>
<td>ViT models (ViT, T2T ViT, Mobile-ViT)</td>
<td><list list-type="bullet">
<list-item>
<p>The proposed Light model outperformed all baselines</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Limited number of classes</p></list-item>
<list-item>
<p>No model compression mechanism</p></list-item>
<list-item>
<p>No robustness evaluation</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td>Compare between ViT-16 and ViT-32</td>
<td>The KMC dataset, and Mendeley Data</td>
<td>Resizing, Cross-Validation</td>
<td>ViT-16 and ViT-32</td>
<td><list list-type="bullet">
<list-item>
<p>ViT-16 outperformed ViT-32 and all CNN baselines across both datasets</p></list-item>
<list-item>
<p>Adam optimizer achieved 100% recall</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>No Robustness comparison or evaluation</p></list-item>
<list-item>
<p>No compression mechanisms</p></list-item>
<list-item>
<p>Saturation and Normalization were not discussed</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>Enhance Robustness against attacks</td>
<td>Fundoscopy, Chest X-ray, ISIC2019, Breast Histology</td>
<td>Normalization and Hold-out Validation</td>
<td>ViT-Tiny</td>
<td><list list-type="bullet">
<list-item>
<p>The ViT model outperformed all CNN base-lines</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Improper Normalization</p></list-item>
<list-item>
<p>Saturation was not discussed.</p></list-item>
<list-item>
<p>Imbalanced Dataset</p></list-item>
<list-item>
<p>Limited to binary classification</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>Efficiency enhancement</td>
<td>OCT2017 (imbalanced), OCT-C8, and an external validation dataset</td>
<td>Weighted loss, Data augmentation</td>
<td>Hybrid CNN-Transformer</td>
<td><list list-type="bullet">
<list-item>
<p>Achieved an extremely high performance with a lower complexity, 99.7% accuracy on OCT2017, 98% on OCT-C8, and 99.1% on the external dataset</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>No architecture reduction techniques</p></list-item>
<list-item>
<p>No robustness evaluation or improvement</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>Performance enhancement</td>
<td>UCSD OCT dataset</td>
<td>Generator (Transformer) and a discriminator (ResNet)</td>
<td>Hybrid CNN-Transformer</td>
<td><list list-type="bullet">
<list-item>
<p>The model achieved 99%, outperforming all baselines</p></list-item>
<list-item>
<p>Introduced a novel hybrid architecture</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Limited Number of Classes</p></list-item>
<list-item>
<p>The study did not work on the model&#x2019;s security and sustainability</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>Enhance diabetic retinopathy classification and FA images generation.</td>
<td>Ultra-widefield retinal imaging dataset with messifor-2</td>
<td>pix2pixHD generator and Swin Transformer classifier</td>
<td>Hybrid GAN-Transformer</td>
<td><list list-type="bullet">
<list-item>
<p>Achieved high realism (SSIM &#x003D; 0.82&#x2013;85) with 82.9% accuracy</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Limited Number of Classes</p></list-item>
<list-item>
<p>Did not enhance efficiency or Robustness against attacks</p></list-item>
</list></td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Propose a Robust framework against adversarial attacks</td>
<td>Retinal OCT, Skin Cancer, and Chest X-ray</td>
<td>ViT attacks detector and CNN classifier</td>
<td>RAD-IoMT (hybrid framework(Transformer and CNN).</td>
<td><list list-type="bullet">
<list-item>
<p>The detector achieved between 92% and 95% accuracy, and the classifier achieved between 86% and 97% accuracy</p></list-item>
<list-item>
<p>Introducing a novel framework</p></list-item>
</list></td>
<td><list list-type="bullet">
<list-item>
<p>Heavy Deployment</p></list-item>
<list-item>
<p>Single-Point of Failure</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
 
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Materials and Methodology</title>
<p>This section presents the proposed GreenShield framework, detailing all phases involved in building and evaluating the GreenShield-ViT model. The process begins with the data collection phase, in which the target dataset is selected. This is followed by a model selection phase, in which an appropriate architecture is identified. A clean training phase was then conducted to train the model on the clean samples (unperturbed samples). Next, a model-pruning phase was applied to compress the model architecture and fine-tune the resulting reduced model. Subsequently, an adversarial training phase is performed to enhance the robustness of the pruned model against adversarial perturbations. Finally, the evaluation setup phase specified the evaluation criteria, established baselines, and defined the metrics used to assess the performance of the proposed approach through comprehensive experiments. The proposed methodology is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed methodology phases.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Data Collection Phase</title>
<p>This study adopts the Retinal OCT classification C-8 dataset [<xref ref-type="bibr" rid="ref-24">24</xref>] because this dataset was designed for multi-classification projects, including eight classes. This dataset is publicly available on Kaggle, labeled, and categorized with a balanced class distribution, including 24,000 images. This dataset contains seven retinal OCT diseases: Age-related Macular Degeneration (AMD), Choroidal Neovascularization (CNV), Central Serous Retinopathy (CSR), Diabetic Macular Edema (DME), Diabetic Retinopathy (DR), Drusen, Macular Hole, and Normal Retina. This makes the dataset ideal for building predictive models. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows an example of each class. In addition to the C-8 dataset, this study incorporates the OCTID dataset, which is available on Kaggle and consists of five classes: Age-related Macular Degeneration (AMD), Central Serous Retinopathy (CSR), Diabetic Retinopathy (DR), Macular Hole (MH), and Normal retina. The dataset contains around 588 OCT images. OCTID is widely used to evaluate model generalization under limited data scenarios. In this study, it is used to address domain shift through a small fine-tuning subphase [<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Retinal OCT C-8 dataset classes [<xref ref-type="bibr" rid="ref-24">24</xref>].</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-2.tif"/>
</fig>
<p>Furthermore, the UCSD-3 dataset (derived from the UCSD OCT dataset) is included for additional evaluation. This dataset is also available on Kaggle and consists of three classes: Choroidal Neovascularization (CNV), Diabetic Macular Edema (DME), and Normal retina. In this study, it is used as an external evaluation dataset. The test set includes 750 images, which are used for performance assessment [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Model Selection Phase</title>
<p>This study proposes training a Vision Transformer model to classify retinal OCT diseases. The selected model backbone was ViT/B-16. This model serves as the standard and original baseline for Vision Transformer models. This architecture has demonstrated a robust performance in retinal OCT studies. However, further improvements in efficiency and robustness are required. This architecture is based on the transformer model [<xref ref-type="bibr" rid="ref-2">2</xref>]. It splits images into patches, each of which is treated as a token. It was pre-trained on ImageNet1k. Patches are processed through a linear layer to form a vector, which is referred to as patch embedding. Positional embeddings were added to preserve the spatial ordering. In addition, a special classification token (CLS) is appended to summarize all patches and capture all the relationships and information needed to help in the prediction (holding the final decision). The vectors were then processed using blocks. Each block contained a multihead self-attention (MHSA) and an MLP layer. Through the self-attention mechanism, the image patches attend to each other to model long-range dependencies. MHSA computes the attention scores using a dot product. ViT splits embeddings into heads. For each head, the Core Attention <xref ref-type="disp-formula" rid="eqn-1">Formula (1)</xref> was applied [<xref ref-type="bibr" rid="ref-2">2</xref>].
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:msup><mml:mi>K</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>s</mml:mi><mml:mi>q</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>The Softmax is applied to assign attention weights to each patch, and the CLS token is then processed through a linear layer and softmax to generate probabilities for each class. After MHSA, the embeddings were processed through an MLP layer (with an activation function) to capture the structures and features inside each image. The backbone architecture consisted of 12 blocks, 12 attention heads, 64 head dimensions, a 4.0 MLP ratio with GELU activation, 196 patches including the CLS patch, 86 million parameters, and a classification head. It is the base ViT model version, with a 16 &#x00D7; 16 patch size, 224 &#x00D7; 224 image size, and 768 embedding lengths. The attention mechanism enables the ViT model to capture the relationships between patches and comprehend the differences between various regions within the images. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the architecture of the vision transformer [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Vision transformer architecture [<xref ref-type="bibr" rid="ref-2">2</xref>].</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-3.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Clean Training Phase</title>
<p>In this phase, the selected ViT model is trained. The proposed training pipeline starts with hold-out validation, 70% training, 15% validation, and 15% testing. This study proposed the application of clean training in two rounds. First, the model is trained to build a reliable classifier before pruning. Second, we fine-tune the pruned model to help it adapt to the new architecture. Two rounds were conducted using the same setup. Only a different number of epochs was used (first round 10 epochs, second round 2 epochs). An NVIDIA Tesla GPU P100 was utilized to train and evaluate the model (provided by Kaggle), 32 batch size and four CPU workers to prepare batches during training.</p>
<p>The training set was utilized for building the model, and the validation set for model selection, hyperparameter tuning, and pruning. The test set was used in the evaluation phase. The pipeline started resizing and then cropping to 224 &#x00D7; 224. Augmentation (random zoom, rotation (&#x002B;&#x2212;20), brightness (0.3), contrast adjustment (0.3), flipping (<italic>p</italic> &#x003D; 0.5), and shift (&#x002B;&#x2212;10)). Subsequently, ImageNet normalization was applied. In clean training, ImageNet normalization can be applied directly during the preprocessing phase without any issues. However, in the adversarial training phase, this study proposed a custom pipeline. The training setup includes cross-entropy as a loss function (difference between actual and predicted values) and an AdamW optimizer to change the model parameters according to gradients to reduce the loss. Weight decay (L2-regularization to prevent parameters from having large values) with the Warmup scheduler (to control the LR values during training), Automatic Mixed Precision (AMP) (to reduce training cost), and LR cosine. A list of anti-softmax saturation techniques was proposed: Temperature scaling 1.5, Post-Attention Layer Norm after attention residual, Label Smoothing 0.1, and Gradient Clipping 1.0.</p>
<p>In each epoch, the model switched from the training model to the validation model. This assists in monitoring the performance of the model during the training. This study proposes a light-validation setup that includes six batches (192 images). Early stopping is triggered to stop the training after three rounds without any enhancements. In the validation and evaluation, the image was resized and cropped to 224 &#x00D7; 244 pixels. However, without augmentation. In addition, light hyperparameter tuning was proposed to select the optimal values for two parameters: the Learning Rate and weight decay. The hyperparameter tuning mechanism was based on five training epochs (8% sample from the training set), with a small evaluation (while training) by a light validation of three batches. This selects the optimal values through the proposed search space for each parameter, as follows:<list list-type="bullet">
<list-item>
<p>Learning rate search space: {5 &#x00D7; 10<sup>5</sup>, 1 &#x00D7; 10<sup>4</sup>, 2 &#x00D7; 10<sup>4</sup>, 3 &#x00D7; 10<sup>4</sup>}</p></list-item>
<list-item>
<p>Weight decay search space: {0.005, 0.01, 0.02, 0.05}</p></list-item>
</list></p>
<p>The values selected based on hyperparameter tuning are (LR &#x003D; 5 &#x00D7; 10<sup>5</sup> and Weight decay &#x003D; 0.01). <xref ref-type="fig" rid="fig-4">Fig. 4</xref> summarizes the proposed clean training pipeline.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The proposed clean training pipeline, including all subphases.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-4.tif"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Model Pruning Phase</title>
<p>This phase prunes the proposed trained model after clean training. This study proposes a pruning mechanism based on a validation set. An importance score was assigned to each block of the ViT model.</p>
<p>In the ViT structure, a single gradient is specified for each parameter within a single neuron. The gradient during training represents two values: magnitude and direction. The magnitude represents the impact of this parameter change on the loss. The direction illustrates how the parameter should be adjusted to minimize loss. The direction was ignored by considering only the magnitude. This determines the importance of the parameter. A neuron with high-magnitude parameters (weights) is an especially important neuron, while a block that holds a particularly important list of neurons is also particularly important. The absolute gradient value represents the magnitude without direction (i.e., it does not include positive or negative signs in the absolute values). The average absolute value for all gradients belonging to a single block represents the importance of the block (how strongly this block impacts).</p>
<p>The proposed mechanism computes the average values of the absolute gradients within each block. Starting with 30 batches of images from the validation set. With forward and backward passes (backpropagation), gradients are calculated for each block. The forward pass is used to compute the loss, and the backward pass computes the gradients. The average of the absolute gradients can be computed as (sum of all gradients divided by the number of gradients). Then, a ranking phase was conducted to identify high- or low-impact blocks (ranked by importance). A low average has a minimal impact. The pruned architecture was fine-tuned using clean training within 2 epochs. This allows the model to adapt to the new architecture and achieve high performance. The proposed ViT backbone included 12 blocks. This mechanism proposes determining the search space of architectures to determine the optimal one. The proposed search space consisted of 12 blocks, resulting in 12/2 &#x003D; 6. Then, 6 &#x002B; 2 and 6 &#x2212; 2. The proposed search space consists of (4, 6, 8) blocks. In the experimental phase, we conducted a series of evaluations to assess each option within the search space. Starting from removing eight blocks and keeping four, pruning six, or pruning four. The study selects the most efficient architecture that only drops the accuracy by less than 2% to find a sustainable model without degrading the performance. The search space represents approximately 30%, 50%, and 70% reductions in model depth. This selection provides well-separated and meaningful compression regions. Where 8 blocks approximate the upper range (layers 7&#x2013;11), 6 blocks capture the central 50% compression point, and 4 blocks represent the lower range (layers 1&#x2013;5). This search space provides a representative summary of performance-efficiency trade-off across the full depth (1&#x2013;12 blocks), enabling effective analysis without exhaustively evaluating all configurations. The next phase involves the proposed adversarial training pipeline. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows the pruning approach.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Proposed pruning approach.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Adversarial Training Phase</title>
<p>The adversarial training phase is an extended training phase after cleaning to make the model robust against adversarial attacks. Adversarial training shows the model slightly modified images during training, which forces the model to learn a feature while staying stable even when the input changes. The model learned to focus on global patterns instead of intricate details, making its layers and tokens more robust. This phase applies adversarial training to the pruned model by applying the same anti-saturation techniques applied to the clean training pipeline (temperature scaling 1.5, post-attention layer norm, label smoothing 0.1, and Gradient Clipping 1.0). The hold-out validation includes a 70% training set, 15% testing set, 15% validation set. The validation set consisted of a stratified sample of 125 images belonging to each class. Validation with only the forward pass (without gradients). The model switches from training mode to validation mode after the end of each training epoch. Validation of 2800 images to maintain clean generalization. Validating clean performance is important to ensure that adversarial training does not harm the performance against clean data. Adversarial training is more complex than clean training and requires more resources. This study proposes the application of AMP precision, while hyperparameter tuning and augmentation were ignored in this phase to avoid additional costs. The same GPU engine (p100) was used in the training and evaluation phases.</p>
<p>The images were resized to 224 &#x00D7; 224 pixels. ImageNet normalization is then applied. ImageNet normalization is important for pretrained models (such as ViTs) to ensure consistency in the input distribution. Adversarial attacks apply perturbations to pixel space. The difference between spaces can change the determined epsilon parameter, which harms adversarial training. ImageNet normalization is represented by <xref ref-type="disp-formula" rid="eqn-2">formula (2)</xref> below [<xref ref-type="bibr" rid="ref-27">27</xref>]:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>here, <italic>c</italic> &#x2208; {R, G, B} denotes to the RGB channel, <italic>x</italic><sub><italic>c</italic></sub> &#x2208; [0, 1] is the pixel space, and <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the channel-wise mean and standard deviation computed from the ImageNet dataset. Adversarial perturbations are defined in the pixel space as <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x03B5;</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mo>&#x03B5;</mml:mo></mml:math></inline-formula> is the perturbation magnitude. The transformation to ImageNet space is computed as follows.</p>
<table-wrap id="table-18">
<table>
<colgroup>
<col align="center"/>
<col align="center"/> </colgroup>
<tbody>
<tr>
<td align="center" colspan="2"><bold>Impact of ImageNet normalization on adversarial training:</bold> Transformation from pixel-space to ImageNet space</td>
</tr>
<tr>
<td><bold>1:</bold></td>
<td><bold>Normalization:</bold> <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0.6</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mn>0.485</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>0.229</mml:mn><mml:mo>&#x2248;</mml:mo><mml:mn>0.502</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>2:</bold></td>
<td><bold>Perturbation in normalized space:</bold> <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mn>0.502</mml:mn><mml:mo>+</mml:mo><mml:mn>0.3</mml:mn><mml:mo>=</mml:mo><mml:mn>0.802</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>3:</bold></td>
<td><bold>De-normalization:</bold> <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mtext>&#xA0;*&#xA0;</mml:mtext></mml:mrow><mml:mn>0.229</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mn>0.485</mml:mn><mml:mo>&#x2248;</mml:mo><mml:mn>0.669</mml:mn></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Thus, the effective perturbation in pixel space becomes &#x0394;&#x00B7;<inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x003D; <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo>&#x03B5;</mml:mo></mml:math></inline-formula>&#x00B7;<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The effective perturbation is approximately 0.069, which is smaller than <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mo>&#x03B5;</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula>. This demonstrates that applying perturbations in normalized space leads to weaker adversarial perturbations during training and evaluation. This study proposes wrapping ImageNet normalization into the model. This adds the attack&#x2019;s perturbations on pixels directly (within the pixel space 0, 1) and then performs normalization.</p>
<table-wrap id="table-19">
<table>
<colgroup>
<col align="center"/>
<col align="center"/> </colgroup>
<tbody>
<tr>
<td align="center" colspan="2"><bold>Impact of Wrapping Normalization:</bold> Transformation from pixel-space to ImageNet space</td>
</tr>
<tr>
<td><bold>1:</bold></td>
<td><bold>Perturbation in pixel space:</bold> <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mn>0.6</mml:mn><mml:mo>+</mml:mo><mml:mn>0.3</mml:mn><mml:mo>=</mml:mo><mml:mn>0.9</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>2:</bold></td>
<td><bold>Normalization:</bold> <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0.9</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mn>0.485</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>0.229</mml:mn><mml:mo>&#x2248;</mml:mo><mml:mn>1.814</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><bold>3:</bold></td>
<td><bold>De-normalization:</bold> <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mn>0.229</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mn>0.485</mml:mn><mml:mo>&#x2248;</mml:mo><mml:mn>0.9</mml:mn></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The value 0.9 corresponds to the correctly perturbed pixel value in the original space. Wrapping normalization in a model is the correct normalization mechanism for adversarial training. Adversarial training generates attack samples by identifying spots to fool the model based on the gradient direction, thereby increasing loss. The model was then trained on the generated adversarial samples to become robust against these spots, thereby reducing the loss. The gradient direction guides adversarial sample generation and guides model training on these samples. <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> represents the adversarial training equation [<xref ref-type="bibr" rid="ref-28">28</xref>].
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>E</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>f</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B4;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03B4;</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow></mml:math></disp-formula></p>
<p>The proposed model was trained using a combination of clean and adversarial examples. Clean samples were injected with PGD and BIM samples, 32 batch size with (four CPU workers) for all the samples (1:1:1 Clean: PGD: BIM). The batch size consisted of 96 images. The gradients of the loss were computed by automatic differentiation in Pytorch and used to construct the PGD and BIM perturbations under the L-infinity constraint. These perturbations were dynamically updated at each iteration using the current model parameters. The training hyperparameters were assigned based on the default setup and the most common practice (&#x03B5; &#x003D; 8/255, &#x03B1; &#x003D; 2/255, PGD-7 steps, BIM-10), including six training epochs and early stopping triggered by the validation loss (two epochs without improvements, the model will stop training), and a cross-entropy loss function (with label smoothing 0.1). AdamW optimizer (Learning Rate: 1e&#x2212;4 and weight decay: 0.01 assigned manually without tuning). In this phase, the PGD and BIM generated adversarial samples retained the original ground-truth labels of their corresponding clean images, and the model was trained jointly on clean and perturbed samples using these labels. <xref ref-type="disp-formula" rid="eqn-4">Formula (4)</xref> represents the PGD attack, and <xref ref-type="disp-formula" rid="eqn-5">Formula (5)</xref> represents the BIM attack. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> shows the adversarial training pipeline [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>]. After the adversarial training, the model was fine-stuned for two epochs using clean training to adapt to domain shift on a mixed dataset composed of OCT-C8 and OCTID samples. During this stage, most model parameters were frozen, and updates were applied only to the classification head, the last transformer block, and the final normalization layer. The fine-tuning was performed on 5000 images and included only the shared classes between the two datasets.<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>adv&#xA0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>&#xA0;Proj</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>&#xA0;sign</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>J</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:msup><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mo>&#x2227;</mml:mo></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>&#xA0;Clip</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="normal">&#x03B5;</mml:mi></mml:mrow><mml:mo>&#x2227;</mml:mo></mml:msup><mml:mi mathvariant="normal">&#x221E;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>&#xA0;sign</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>J</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Adversarial training phase, including all subphases.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-6.tif"/>
</fig>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Evaluation Setup</title>
<p>The evaluation plan included several experiments to evaluate the proposed model (GreenShield-ViT). The study plan explored the pruning search space to select the optimal model. This comparison was used to investigate sustainability enhancements. t-Distributed Stochastic Neighbor Embedding (t-SNE) plots and attention rollout-visualization are investigated to ensure that clusters remain clearly separated and the attention is in the same relevant regions. Following adversarial training, the GreenShield-ViT model&#x2019;s performance and calibration were assessed against adversarial samples. Five repeated stratified hold-out experiments were conducted to ensure statistical robustness. The experiments were performed using different random seeds (42, 123, 2024, 3407, and 9999). In each trial, the dataset was randomly split into 70% training, 15% validation, and 15% testing while preserving class distribution. The GreenShield-ViT is benchmarked against Mobile-ViT Extra Extra Small (XXS), ViT-Small (ViT-Tiny/16), and ViT-Tiny (ViT-Tiny/16). These baselines were selected from related work sections [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]. We explored their robustness against attacks after applying the proposed clean and adversarial training pipelines to all models. Investigating their robustness against attacks. The robustness evaluation is based on a small set of white-box attacks, including (FGSM, PGD, BIM, and the proposed hybrid model) [<xref ref-type="bibr" rid="ref-6">6</xref>]. The proposed hybrid attack is a combination of FGSM and PGD attacks. The hybrid attack begins with an FGSM step (based on the epsilon range and gradient direction) to make a significant jump in reaching the bound (the allowed value based on the epsilon range). Subsequently, if the gradient direction changes, it includes the PGD steps. If the direction remains the same, the steps are projected to prevent pixels from moving farther than &#x00B1;&#x03B5;. This attack is designed to initiate from the boundary rather than starting from the pixel value. This attack is introduced as an additional evaluation scenario to assess the robustness of the proposed model under mixed iterative attack strategies. This setting reflects a potential adversarial scenario where multiple attack strategies are combined to fool the model. This evaluation also assesses the robustness against unexpected attack strategies that were not explicitly used during training. The same evaluation hyperparameters were applied to all attacks. Default hyperparameters were selected (following widespread practice). The hyperparameters of the proposed evaluation attack are listed in <xref ref-type="table" rid="table-2">Table 2</xref>. Additionally <xref ref-type="fig" rid="fig-7">Fig. 7</xref> shows how the adopted parameters can generate a visually imperceptible perturbation. The per-pixel difference was computed, and matplotlib was utilized to highlight the changed pixels (the highest perturbation magnitude).</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Evaluation hyperparameters.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>Epsilon</td>
<td>8/255</td>
<td>Perturbation Bound</td>
</tr>
<tr>
<td>Step Size</td>
<td>2/255</td>
<td>Step Size Per Iteration</td>
</tr>
<tr>
<td>FGSM Steps</td>
<td>1 step</td>
<td>Standard FGSM</td>
</tr>
<tr>
<td>PGD Steps</td>
<td>7 steps</td>
<td>PGD-7</td>
</tr>
<tr>
<td>BIM Steps</td>
<td>10 steps</td>
<td>BIM-10</td>
</tr>
<tr>
<td>Hybrid FGSM-PGD</td>
<td>1 FGSM &#x002B; 5 PGD</td>
<td>A Hybrid Proposed Attack</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Adversarial perturbations on the proposed hyperparameters.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-7.tif"/>
</fig>
<p>After demonstrating that the proposed model is more robust than the baselines, the model was fine-tuned using a combined dataset from OCT-C8 and OCTID to mitigate domain shift and improve generalization. It was subsequently evaluated on three datasets (OCT-C8, UCSD-3, and OCTID), where UCSD-3 represents an unseen dataset for assessing generalization, and OCTID provides an additional evaluation under a different domain shift [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>]. The UCSD-3 dataset was proposed solely for evaluation to ensure that the model can generalize well against external samples beyond the main dataset. This dataset is available on Kaggle with three labels (CNV, DRUSEEN, Normal). The test set contained 750 images. All the experiments were conducted using PyTorch and Torchvision. Finally, the model is evaluated under different epsilon levels (2/255, 4/255, 8/255, 16/255) to ensure that the model is robust beyond the common evaluation practice. <xref ref-type="fig" rid="fig-8">Fig. 8</xref> shows the attacks on the different epsilon levels. In addition, the model is evaluated against PGD-7 with three random restarts and Transfer-PGD attack, which was generated on a surrogate original ViT model before pruning (12-blocks).</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Adversarial perturbations at different epsilon levels.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-8.tif"/>
</fig>
<p>A set of evaluation metrics was adapted to interpret the findings and results. This study evaluates the model performance based on the Accuracy, Precision, Recall, F1-score, and average of max softmax probabilities to check the softmax saturation. Prior studies have demonstrated that the average max softmax probability (Avg MPS) is a valid metric to check softmax saturation [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>]. When the maximum softmax probability ranges between (0.97&#x2013;1.00), then the softmax is saturated. In addition, the Expected Calibration Error (ECE) measures the gap between the confidence and accuracy of the model (higher ECE indicates a higher gap) [<xref ref-type="bibr" rid="ref-31">31</xref>]. The efficiency of the model is evaluated using Floating-point operations per second (FLOPs), which is computed using fvcore and ptflops (in Python), Number of Model parameters, Model Size (in MB), Inference Time [<xref ref-type="bibr" rid="ref-32">32</xref>], Inference Memory usage [<xref ref-type="bibr" rid="ref-33">33</xref>], and Inference Carbon Footprint (CO<sub>2</sub>). The carbon footprint refers to the greenhouse gas emissions generated by inference measured in CO<sub>2</sub>-equivalent units. Carbon emissions were computed using the Green Algorithm estimation tool [<xref ref-type="bibr" rid="ref-34">34</xref>]. Pytorch Compute Unified Device Architecture (CUDA) was used for timing and memory measurements. The robustness of the Model was evaluated using the model&#x2019;s accuracy against each attack [<xref ref-type="bibr" rid="ref-35">35</xref>], Attack Success Rate, the robustness gap between clean and adversarial accuracies, and the improvement rate before and after applying adversarial training [<xref ref-type="bibr" rid="ref-36">36</xref>]. The formula below represents the carbon footprint equation [<xref ref-type="bibr" rid="ref-34">34</xref>].
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>Carbon Footprint</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Energy Needed</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>*</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>Carbon Intensity&#xA0;</mml:mtext></mml:mrow></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments and Results</title>
<p>This section presents all the experiments and evaluations in this study, including the evaluation of the ViT model, exploration of the pruning search space, highlighting efficiency enhancements, and investigation of robustness against attacks.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Experiment One</title>
<p>In this experiment, the lightest architecture was selected without degrading accuracy by more than 2%. The results demonstrated that the 6-block model (pruning 6) was the optimal model, balancing performance and efficiency. The study selected the 6-block as the GreenShield-ViT model architecture. The 4-block model is the lightest, but the performance is reduced from 96.2% to 89.9% (6.3% reduction). While the 6-block model reduced from 96.2% to 95.6% (0.6% reduction), which is less than 2%. The 8-block model achieved an impressive performance, but the 6-block model achieved a competitive performance with significant enhancement of efficiency. The stability of block importance reveals that Transformer layers have unequal contributions, where lower-ranked blocks provide limited information and can be removed. Pruning up to 6 blocks maintains performance, whereas further pruning eliminates critical feature representation and results in performance degradation. All models were evaluated under identical experimental conditions. All experiments were conducted in a Kaggle notebook using an NVIDIA Tesla P100 GPU, with the same software framework (PyTorch and TIMM), identical input resolution (224 &#x00D7; 224), and the same test dataset. Under these standardized settings, all measurements were made to provide unbiased comparisons across models. The pruning was conducted based on 30 batches from the validation set. The results demonstrate that the top-rank blocks remain consistent across different batch sizes (30 and 40 batches), indicating that importance estimation is stable, and based on this observation, 30 batches were selected. <xref ref-type="table" rid="table-3">Table 3</xref> demonstrated stability in ranking. <xref ref-type="table" rid="table-4">Table 4</xref> presents the results of the pruning search space, and <xref ref-type="fig" rid="fig-9">Fig. 9</xref> shows the efficiency enhancement. <xref ref-type="fig" rid="fig-10">Fig. 10</xref> shows the ROC curve. The t-SNE plot (<xref ref-type="fig" rid="fig-11">Fig. 11</xref>) shows that class separability is preserved after pruning, as the clusters remain clearly separated. In addition, the attention rollout visualization (<xref ref-type="fig" rid="fig-12">Fig. 12</xref>) confirms that the model focuses on similar clinically relevant regions before and after pruning. <xref ref-type="table" rid="table-5">Tables 5</xref> and <xref ref-type="table" rid="table-6">6</xref> demonstrate consistent performance across five trials.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Pruning batches analysis&#x2014;Experiment 1.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Batches</th>
<th>Top-4 Layers</th>
<th>Top-6 Layers</th>
<th>Top-8 Layers</th>
</tr>
</thead>
<tbody>
<tr>
<td>30</td>
<td>10.9.8.5</td>
<td>10.9.8.5.6.4</td>
<td>10.9.8.5.6.4.3.7</td>
</tr>
<tr>
<td>40</td>
<td>10.9.8.5</td>
<td>10.9.8.5.6.4</td>
<td>10.9.8.5.6.4.3.7</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Pruning search space results&#x2014;Experiment 1.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Standard ViT (ViT/B-16)</th>
<th>4-Block</th>
<th>6-Block</th>
<th>8-Block</th>
</tr>
</thead>
<tbody>
<tr>
<td>Accuracy</td>
<td>96.2%</td>
<td>89.9%</td>
<td>95.6%</td>
<td>96.3%</td>
</tr>
<tr>
<td>Precision</td>
<td>96.3%</td>
<td>90.8%</td>
<td>95.6%</td>
<td>96.3%</td>
</tr>
<tr>
<td>Recall</td>
<td>96.2%</td>
<td>89.9%</td>
<td>95.6%</td>
<td>96.3%</td>
</tr>
<tr>
<td>F1-score</td>
<td>96.2%</td>
<td>89.9%</td>
<td>95.6%</td>
<td>96.8%</td>
</tr>
<tr>
<td>FLOPs</td>
<td>33.76</td>
<td>11.42</td>
<td>17.00</td>
<td>23.40</td>
</tr>
<tr>
<td>Parameters (M)</td>
<td>85.82</td>
<td>29.11</td>
<td>43.29</td>
<td>57.00</td>
</tr>
<tr>
<td>Fine-Tune Time (s)</td>
<td>3077.43</td>
<td>213.55</td>
<td>311</td>
<td>600</td>
</tr>
<tr>
<td>Inference Time (s)</td>
<td>23.85</td>
<td>12.71</td>
<td>12.70</td>
<td>16.27</td>
</tr>
<tr>
<td>Inference Memory Usage (MB)</td>
<td>2222.16</td>
<td>1225</td>
<td>1471</td>
<td>1726</td>
</tr>
<tr>
<td>Model Size (MB)</td>
<td>335</td>
<td>113</td>
<td>170</td>
<td>224</td>
</tr>
<tr>
<td>Inference CO<sub>2</sub> (kg)</td>
<td>0.00034</td>
<td>0.00016</td>
<td>0.00018</td>
<td>0.00021</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Comparison of computational complexity between GreenShield-ViT and ViT/B-16.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-9.tif"/>
</fig><fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>ROC curve for the proposed GreenShield-ViT model.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-10.tif"/>
</fig><fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>T-SNE plot to ensure that the clusters remain clearly separated after pruning.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-11.tif"/>
</fig><fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Attention-rollout visualization to show attention regions on both GreenShield-ViT and ViT/B-16 models.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-12.tif"/>
</fig><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Proposed model (6-blocks) performance across five trials&#x2014;Experiment 1.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Trial</th>
<th>Seed</th>
<th>Clean (%)</th>
<th>FGSM (%)</th>
<th>PGD (%)</th>
<th>BIM (%)</th>
<th>Hybrid (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Trial 1</td>
<td>42</td>
<td>95.6</td>
<td>95.6</td>
<td>95.6</td>
<td>95.6</td>
<td>42</td>
</tr>
<tr>
<td>Trial 2</td>
<td>123</td>
<td>95.47</td>
<td>95.45</td>
<td>95.47</td>
<td>95.45</td>
<td>123</td>
</tr>
<tr>
<td>Trial 3</td>
<td>2024</td>
<td>95.39</td>
<td>95.42</td>
<td>95.39</td>
<td>95.37</td>
<td>2024</td>
</tr>
<tr>
<td>Trial 4</td>
<td>3407</td>
<td>94.58</td>
<td>94.82</td>
<td>94.58</td>
<td>94.59</td>
<td>3407</td>
</tr>
<tr>
<td>Trial 5</td>
<td>9999</td>
<td>95.67</td>
<td>95.68</td>
<td>95.67</td>
<td>95.65</td>
<td>9999</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Statistical summary of proposed model (6-blocks) performance across five trials&#x2014;Experiment 1.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Mean (%)</th>
<th>Std. (%)</th>
<th>95% CI (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Accuracy</td>
<td>95.34</td>
<td>0.42</td>
<td>[94.97, 95.71]</td>
</tr>
<tr>
<td>Precision</td>
<td>95.39</td>
<td>0.31</td>
<td>[95.12, 95.66]</td>
</tr>
<tr>
<td>Recall</td>
<td>95.34</td>
<td>0.42</td>
<td>[94.97, 95.71]</td>
</tr>
<tr>
<td>F1-score</td>
<td>95.33</td>
<td>0.41</td>
<td>[94.97, 95.69]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The results indicate that reducing the number of transformer blocks does not significantly degrade performance beyond a certain depth. This is because earlier transformer layers capture the most critical structural and textural features of retinal OCT images, while deeper layers provide only marginal refinements. Given the structured nature of OCT data, redundant deeper representations can be removed without sacrificing discriminative capability. This explains why the 6-block configuration achieves a near-optimal balance between efficiency and accuracy, while the 8-block configuration offers minimal additional performance gains at a higher computational cost.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experiment Two</title>
<p>This experiment investigated the robustness of the selected pruned model (6-Block) against adversarial attacks. Adversarial training was applied to the pruned model (GreenShield-ViT). This model represents the proposed approach, including all phases in this experiment. This experiment included the results before and after the proposed adversarial training to demonstrate the enhancement. The results indicated that the proposed adversarial training enhanced the robustness and calibration of the model. Before applying adversarial training, the model was unstable, achieving low accuracy with a high confidence rate. <xref ref-type="table" rid="table-7">Table 7</xref> shows the enhancement in the robustness. The calibration enhancements are demonstrated in <xref ref-type="table" rid="table-8">Table 8</xref> using the average of the maximum softmax probability (Avg. MSP) and expected calibration error (ECE). <xref ref-type="fig" rid="fig-13">Fig. 13</xref> illustrates the alignment between the accuracy and average MSP before and after adversarial training. As shown in <xref ref-type="table" rid="table-9">Tables 9</xref> and <xref ref-type="table" rid="table-10">10</xref>, the model demonstrates consistent performance across all trials, with minimal variation in both clean and adversarial accuracy, confirming its stability.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Performance of the proposed model under adversarial attacks&#x2014;Experiment 2.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Before Adversarial Training</th>
<th>After Adversarial Training</th>
<th>Improvement Rate</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clean Accuracy</td>
<td>95.6%</td>
<td>90%</td>
<td>&#x2212;5.6% (Worse)</td>
</tr>
<tr>
<td>FGSM Accuracy</td>
<td>0.47%</td>
<td>76.31%</td>
<td>&#x002B;75.84% (Better)</td>
</tr>
<tr>
<td>PGD Accuracy</td>
<td>0.00%</td>
<td>73.03%</td>
<td>&#x002B;73.03% (Better)</td>
</tr>
<tr>
<td>BIM Accuracy</td>
<td>0.00%</td>
<td>70.72%</td>
<td>&#x002B;70.72% (Better)</td>
</tr>
<tr>
<td>Hybrid Attack Accuracy</td>
<td>0.00%</td>
<td>72.97%</td>
<td>&#x002B;72.97% (Better)</td>
</tr>
<tr>
<td>FGSM Success Rate</td>
<td>99.53%</td>
<td>23.69%</td>
<td>&#x2212;75.84% (Better)</td>
</tr>
<tr>
<td>PGD Success Rate</td>
<td>100%</td>
<td>26.97%</td>
<td>&#x2212;73.03% (Better)</td>
</tr>
<tr>
<td>BIM Success Rate</td>
<td>100%</td>
<td>29.28%</td>
<td>&#x2212;70.72% (Better)</td>
</tr>
<tr>
<td>Hybrid Attack Success Rate</td>
<td>100%</td>
<td>27.03%</td>
<td>&#x2212;72.97% (Better)</td>
</tr>
<tr>
<td>Robustness Gap: FGSM and Clean</td>
<td>95.13%</td>
<td>13.69%</td>
<td>&#x2212;81.44% (Better)</td>
</tr>
<tr>
<td>Robustness Gap: Clean and PGD</td>
<td>95.60%</td>
<td>16.97%</td>
<td>&#x2212;78.63% (Better)</td>
</tr>
<tr>
<td>Robustness Gap: BIM and Clean</td>
<td>95.60%</td>
<td>19.28%</td>
<td>&#x2212;76.32% (Better)</td>
</tr>
<tr>
<td>Robustness Gap: Hybrid attack and Clean</td>
<td>95.60%</td>
<td>17.03%</td>
<td>&#x2212;78.57% (Better)</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Statistical robustness analysis of model performance across five independent runs (mean &#x00B1; standard deviation and 95% confidence intervals).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Condition</th>
<th>Avg. MSP before Adversarial Training</th>
<th>Avg. MSP after Adversarial Training</th>
<th>ECE before Adversarial Training</th>
<th>ECE after Adversarial Training</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clean</td>
<td>0.867</td>
<td>0.78</td>
<td>0.089</td>
<td>0.117</td>
</tr>
<tr>
<td>FGSM</td>
<td>0.585</td>
<td>0.725</td>
<td>0.589</td>
<td>0.049</td>
</tr>
<tr>
<td>PGD</td>
<td>0.986</td>
<td>0.713</td>
<td>0.988</td>
<td>0.039</td>
</tr>
<tr>
<td>BIM</td>
<td>0.994</td>
<td>0.705</td>
<td>0.994</td>
<td>0.044</td>
</tr>
<tr>
<td>Hybrid Attack</td>
<td>0.969</td>
<td>0.711</td>
<td>0.969</td>
<td>0.040</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>This figure demonstrates the alignment between Model accuracy and confidence (AVG MSP). (<bold>a</bold>) represents the results before applying the adversarial training; (<bold>b</bold>) presents the strong alignment after applying the adversarial training.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-13.tif"/>
</fig><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Model performance (GreenShield-ViT) across five trials&#x2014;Experiment 2.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Trial</th>
<th>Seed</th>
<th>Clean (%)</th>
<th>FGSM (%)</th>
<th>PGD (%)</th>
<th>BIM (%)</th>
<th>Hybrid (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Trial 1</td>
<td>42</td>
<td>91.5</td>
<td>78.5</td>
<td>74.5</td>
<td>72</td>
<td>74.5</td>
</tr>
<tr>
<td>Trial 2</td>
<td>123</td>
<td>91.14</td>
<td>78.81</td>
<td>74.78</td>
<td>72.08</td>
<td>74.83</td>
</tr>
<tr>
<td>Trial 3</td>
<td>2024</td>
<td>90.78</td>
<td>76</td>
<td>72.47</td>
<td>70.47</td>
<td>72.64</td>
</tr>
<tr>
<td>Trial 4</td>
<td>3407</td>
<td>89.75</td>
<td>76.97</td>
<td>73.92</td>
<td>72.11</td>
<td>73.78</td>
</tr>
<tr>
<td>Trial 5</td>
<td>9999</td>
<td>90.17</td>
<td>75.03</td>
<td>71.97</td>
<td>70.36</td>
<td>72.22</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Statistical summary of model performance (GreenShield-ViT) across five trials&#x2014;Experiment 2.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Mean &#x00B1; Std.</th>
<th>95% Confidence Interval</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clean</td>
<td>90.67&#x00B1; 0.67</td>
<td>[89.84, 91.50]</td>
</tr>
<tr>
<td>FGSM</td>
<td>76.86 &#x00B1; 1.55</td>
<td>[74.93, 78.79]</td>
</tr>
<tr>
<td>PGD</td>
<td>73.53 &#x00B1; 1.15</td>
<td>[72.10, 74.96]</td>
</tr>
<tr>
<td>BIM</td>
<td>71.40 &#x00B1; 0.95</td>
<td>[70.22, 72.58]</td>
</tr>
<tr>
<td>Hybrid</td>
<td>73.59 &#x00B1; 1.03</td>
<td>[72.31, 74.87]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To address potential concerns regarding overfitting and result saturation, we conducted five independent experimental trials using different random seeds. The results, summarized in <xref ref-type="table" rid="table-7">Tables 7</xref> and <xref ref-type="table" rid="table-8">8</xref>, report mean, standard deviation, and 95% confidence intervals for all evaluation metrics. The observed low variance across trials confirms the stability and generalization capability of the proposed GreenShield-ViT model, even under adversarial conditions.</p>

</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experiment Three</title>
<p>This experiment aims to evaluate the robustness of the baselines (Mobile-ViT, ViT-Tiny, and ViT-Small) against adversarial attacks (FGSM, PGD, BIM, and the hybrid attack). The robustness performance is evaluated to compare baselines with the proposed model and identify the most robust, lightweight model. The same proposed pipeline was applied to all models. Exploring the robustness of these lightweight models addresses a research gap in prior OCT work. To ensure a fair comparison, all models were trained using identical data splits, preprocessing steps, hyperparameter settings, and adversarial training configuration, and evaluated under the same attack conditions. The results demonstrated that the models achieved impressive performance against clean samples, with a significant robustness gap. Furthermore, larger models are not necessarily more robust, as ViT-Small achieved an impressive performance against PGD and BIM samples, which were included in training, but performed poorly against FGSM and Hybrid attack, which indicates attack-specific overfitting and limited generalization. <xref ref-type="table" rid="table-11">Table 11</xref> presents the results of the baseline models, and <xref ref-type="fig" rid="fig-14">Fig. 14</xref> illustrates how the proposed model is significantly more robust than the baseline models, based on the success rate of each attack.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Performance of baseline models under adversarial attacks&#x2014;Experiment 3.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Mobile-ViT</th>
<th>ViT-Tiny</th>
<th>ViT-Small</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clean Accuracy</td>
<td>96.17%</td>
<td>94.72%</td>
<td>95.44%</td>
</tr>
<tr>
<td>FGSM Accuracy</td>
<td>40.11%</td>
<td>39.06</td>
<td>33.39%</td>
</tr>
<tr>
<td>PGD Accuracy</td>
<td>35.19%</td>
<td>39.22%</td>
<td>63.33%</td>
</tr>
<tr>
<td>BIM Accuracy</td>
<td>34.56%</td>
<td>38.06%</td>
<td>65.03%</td>
</tr>
<tr>
<td>Hybrid Attack Accuracy</td>
<td>36.92%</td>
<td>38.72%</td>
<td>55.14%</td>
</tr>
<tr>
<td>FGSM Success Rate</td>
<td>59.89%</td>
<td>60.94%</td>
<td>66.61%</td>
</tr>
<tr>
<td>PGD Success Rate</td>
<td>64.81%</td>
<td>60.78%</td>
<td>36.67%</td>
</tr>
<tr>
<td>BIM Success Rate</td>
<td>65.44%</td>
<td>61.94%</td>
<td>34.97%</td>
</tr>
<tr>
<td>Hybrid Attack Success Rate</td>
<td>63.08%</td>
<td>61.28%</td>
<td>44.86%</td>
</tr>
<tr>
<td>Robustness Gap: FGSM and Clean</td>
<td>56.06%</td>
<td>55.66%</td>
<td>62.05%</td>
</tr>
<tr>
<td>Robustness Gap: Clean and PGD</td>
<td>60.97%</td>
<td>55.50%</td>
<td>32.11%</td>
</tr>
<tr>
<td>Robustness Gap: BIM and Clean</td>
<td>61.61%</td>
<td>56.66%</td>
<td>30.41%</td>
</tr>
<tr>
<td>Robustness Gap: Hybrid attack and Clean</td>
<td>59.25%</td>
<td>56%</td>
<td>40.30%</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>Comparison of robustness between GreenShield-ViT and baselines.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-14.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experiment Four</title>
<p>This experiment aims to fine-tune the model on a mix of OCTID and OCT-C8 datasets. This fine-tuning helps the model to adapt to domain shift. After fine-tuning, the model was evaluated against three datasets: OCT-C8 (15% test set images), OCTID (115 images), and UCSD-3 (735 images). The results indicate that the proposed model achieves a high clean accuracy. The results against the OCTID dataset indicate that the model adapted to domain shift, while the results against UCSD-3 indicate that the model can generalize against unseen data. The significant improvement in the robustness after adversarial training demonstrates that the proposed pipeline builds a robust model. In addition, the findings illustrate a strong alignment between the accuracy and confidence levels after adversarial training. <xref ref-type="table" rid="table-12">Table 12</xref> lists the experimental results, while <xref ref-type="table" rid="table-13">Tables 13</xref>&#x2013;<xref ref-type="table" rid="table-15">15</xref> represent the confusion matrix across different datasets.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Performance of the proposed model on OCT-C8, OCTID, and UCSD-3 datasets&#x2014;Experiment 4.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>OCT-C8</th>
<th>OCTID</th>
<th>UCSD-3</th>
</tr>
</thead>
<tbody>
<tr>
<td>Clean Accuracy</td>
<td>92.54%</td>
<td>94.78%</td>
<td>89.20%</td>
</tr>
<tr>
<td>FGSM Accuracy</td>
<td>80.29%</td>
<td>74.78%</td>
<td>66%</td>
</tr>
<tr>
<td>PGD Accuracy</td>
<td>77.36%</td>
<td>71.30%</td>
<td>61.07%</td>
</tr>
<tr>
<td>BIM Accuracy</td>
<td>75.43%</td>
<td>66.09%</td>
<td>58.67%</td>
</tr>
<tr>
<td>Hybrid Attack Accuracy</td>
<td>77.64%</td>
<td>70.43%</td>
<td>61.33%</td>
</tr>
<tr>
<td>AVG Max of Softmax Probability (Clean)</td>
<td>84.34%</td>
<td>88.32%</td>
<td>79.17%</td>
</tr>
<tr>
<td>AVG Max of Softmax Probability (FGSM)</td>
<td>78.67%</td>
<td>79.40%</td>
<td>72.78%</td>
</tr>
<tr>
<td>AVG Max of Softmax Probability (PGD)</td>
<td>77.54%</td>
<td>78.54%</td>
<td>71.95%</td>
</tr>
<tr>
<td>AVG Max of Softmax Probability (BIM)</td>
<td>76.65%</td>
<td>77.51%</td>
<td>71.57%</td>
</tr>
<tr>
<td>AVG Max of Softmax Probability (Hybrid Attack)</td>
<td>77.42%</td>
<td>78.49%</td>
<td>71.94%</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Confusion matrix for clean samples (OCT-C8 dataset)&#x2014;Experiment 4.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th></th>
<th>AMD</th>
<th>CNV</th>
<th>CSR</th>
<th>DME</th>
<th>DR</th>
<th>DRUSEN</th>
<th>MH</th>
<th>NORMAL</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>AMD</bold></td>
<td>350</td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>0</td>
</tr>
<tr>
<td><bold>CNV</bold></td>
<td>0</td>
<td>324</td>
<td>0</td>
<td>7</td>
<td>0</td>
<td>18</td>
<td>0</td>
<td>1</td>
</tr>
<tr>
<td><bold>CSR</bold></td>
<td>0</td>
<td>0</td>
<td>343</td>
<td>0</td>
<td>5</td>
<td>0</td>
<td>0</td>
<td>2</td>
</tr>
<tr>
<td><bold>DME</bold></td>
<td>0</td>
<td>8</td>
<td>0</td>
<td>290</td>
<td>0</td>
<td>2</td>
<td>0</td>
<td>50</td>
</tr>
<tr>
<td><bold>DR</bold></td>
<td>0</td>
<td>0</td>
<td>1</td>
<td>0</td>
<td>343</td>
<td>0</td>
<td>1</td>
<td>5</td>
</tr>
<tr>
<td><bold>DRUSEN</bold></td>
<td>0</td>
<td>17</td>
<td>0</td>
<td>5</td>
<td>0</td>
<td>257</td>
<td>0</td>
<td>71</td>
</tr>
<tr>
<td><bold>MH</bold></td>
<td>0</td>
<td>0</td>
<td>3</td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>347</td>
<td>0</td>
</tr>
<tr>
<td><bold>NORMAL</bold></td>
<td>0</td>
<td>2</td>
<td>0</td>
<td>9</td>
<td>0</td>
<td>2</td>
<td>0</td>
<td>337</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Confusion matrix for clean samples (OCTID dataset)&#x2014;Experiment 4.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th></th>
<th>AMD</th>
<th>CSR</th>
<th>DR</th>
<th>MH</th>
<th>NORMAL</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>AMD</bold></td>
<td>5</td>
<td>4</td>
<td>2</td>
<td>0</td>
<td>0</td>
</tr>
<tr>
<td><bold>CSR</bold></td>
<td>0</td>
<td>21</td>
<td>0</td>
<td>0</td>
<td>0</td>
</tr>
<tr>
<td><bold>DR</bold></td>
<td>0</td>
<td>0</td>
<td>22</td>
<td>0</td>
<td>0</td>
</tr>
<tr>
<td><bold>MH</bold></td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>20</td>
<td>0</td>
</tr>
<tr>
<td><bold>NORMAL</bold></td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>0</td>
<td>41</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Confusion matrix for clean samples (UCD-3 dataset)&#x2014;Experiment 4.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th></th>
<th>CNV</th>
<th>DRUSEN</th>
<th>NORMAL</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>CNV</bold></td>
<td>243</td>
<td>2</td>
<td>0</td>
<td>245</td>
</tr>
<tr>
<td><bold>DRUSEN</bold></td>
<td>38</td>
<td>186</td>
<td>25</td>
<td>249</td>
</tr>
<tr>
<td><bold>NORMAL</bold></td>
<td>0</td>
<td>1</td>
<td>240</td>
<td>241</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Experiment Five</title>
<p>This Experiment aims to evaluate the proposed model against varying epsilon levels to ensure that the model can generalize beyond the specific default parameters proposed for evaluation (default in ImageNet 8/255). The results demonstrate that the model can achieve competitive results under standard adversarial benchmark settings (2/255, 4/255, and 8/255), while maintaining acceptable robustness at 16/255, which exceeds commonly adopted research benchmarks. The model was evaluated against PGD-R3(3 random restarts) and Transfer-PGD scenarios. The experimental results are listed in <xref ref-type="table" rid="table-16">Table 16</xref>. <xref ref-type="fig" rid="fig-15">Fig. 15</xref> represents the loss level against different epsilon values. The results indicate that the model can achieve a satisfactory performance against different scenarios, while the model does not exhibit gradient masking, as increasing the perturbation magnitude consistently leads to a higher loss and lower accuracy.</p>
<table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Performance of the proposed model at different epsilon levels&#x2014;Experiment 5.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Epsilon (&#x03B5;)</th>
<th>FGSM</th>
<th>PGD</th>
<th>BIM</th>
<th>Hybrid Attack</th>
<th>PGD (R3)</th>
<th>Transfer-PGD</th>
</tr>
</thead>
<tbody>
<tr>
<td>2/255</td>
<td>90.34%</td>
<td>90.36%</td>
<td>90.39%</td>
<td>90.36%</td>
<td>90.36%</td>
<td>92.21%</td>
</tr>
<tr>
<td>4/255</td>
<td>87.36%</td>
<td>86.89%</td>
<td>86.86%</td>
<td>86.86%</td>
<td>86.68%</td>
<td>91.82%</td>
</tr>
<tr>
<td>8/255</td>
<td>80.29%</td>
<td>77.43%</td>
<td>75.43%</td>
<td>77.64%</td>
<td>77.32%</td>
<td>91.54%</td>
</tr>
<tr>
<td>16/255</td>
<td>64.26%</td>
<td>65.75%</td>
<td>41.11%</td>
<td>52.29%</td>
<td>64.89%</td>
<td>91.25%</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>GreenShield-ViT loss levels against different epsilon values.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80864-fig-15.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion</title>
<p>This study investigated the robustness of the lightweight VIT architecture. In this study, the proposed GreenShield framework enhanced model efficiency through a gradient-based block-importance pruning, while robustness against white-box L-infinity attacks was enhanced through an integrated adversarial defense stage. The study contributions and findings are as follows:<list list-type="bullet">
<list-item>
<p>GreenShield is introduced as a novel empirically validated framework that integrates gradient-based block-importance pruning with adversarial training in retinal OCT. Achieving a significant sustainability enhancement, including FLOPs reduction by 49.6%, the model&#x2019;s parameters were reduced by 49.6%. Inference memory usage was reduced by 33.8% in MB, faster inference in seconds by 46.8%, model size reduction by 49.3% in MB, and 47.1% lower emissions (CO<sub>2</sub>). The overall success rate for all attacks reduced across three datasets (OCT-c8, UCSD-3, and OCTID).</p></list-item>
<list-item>
<p>GreenShield-ViT outperformed existing lightweight ViT variants (Mobile-ViT, ViT-Small, and ViT-Tiny) in terms of robustness against all attacks (FGSM, PGD, BIM, and a Hybrid Attack). Presenting reliable clinical decision-making and efficient integration into hospital systems.</p></list-item>
<list-item>
<p>The experimental results indicate that adversarial robustness depends more on architectural design than model size. The ViT-small, the largest baseline model, suffers from overfitting, while the proposed model, with even greater capacity, achieves better generalization and maintains stable performance.</p></list-item>
<list-item>
<p>Adversarial training was enhanced by wrapping normalization into the model, which improved perturbation generation. This issue has rarely been examined in prior OCT studies.</p></list-item>
<list-item>
<p>Model calibration was significantly improved, demonstrating a strong alignment between the accuracy and confidence levels under attack conditions (stable confidence levels).</p></list-item>
<list-item>
<p>The proposed framework jointly improves both robustness and efficiency, addressing a research gap in retinal OCT studies where these objectives have rarely been investigated together.</p></list-item>
</list></p>
<p>The proposed adversarial training enhanced the overall robustness of all the models included in this study. This study focuses on retinal OCT disease classification and fills a list of research gaps. <xref ref-type="table" rid="table-17">Table 17</xref> presents the gaps identified in previous studies and illustrates how they were addressed in the present study.</p>
<table-wrap id="table-17">
<label>Table 17</label>
<caption>
<title>The Prior work limitations that are addressed in this study.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Studies</th>
<th>Research Gap or Limitation</th>
<th>Addressing the Gap or Limitation in This Study</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Lack of sustainable or efficient models</td>
<td>The proposed approach reduced the ViT model size by 49.3%.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Insufficient security mechanism or adversarial Robustness</td>
<td>The proposed adversarial training enhanced the model&#x2019;s robustness against attacks.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Weak or Inconsistent Evaluation</td>
<td>The evaluation included more than 4 attacks, with a large sample.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Few or Limited Diagnoses Classes</td>
<td>The proposed methodology achieved high performance against 8 classes.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Absence of Vision Transformers Compression</td>
<td>The proposed pruning mechanism is applied to the ViT model.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Absence of Vision Transformers Adversarial Defense</td>
<td>The proposed security mechanism is applied to the Pruned ViT model.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Did not handle the Imbalanced Datasets</td>
<td>The proposed dataset is already balanced.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>Sub-optimal or Incorrect pre-processing practices</td>
<td>The proposed approach applied a correct normalization with anti-saturation techniques.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6">
<label>6</label>
<title>Limitations and Future Work</title>
<p>The proposed framework is limited to a specific type of attack, namely FGSM, PGD, and BIM (white-box attacks). This study can be extended to incorporate a larger set of attacks, such as L2 attacks (DeepFool, AutoAttack, and CW).</p>
<p>We consider hardware as a limitation of this study, where we utilized a GPU p100 (with a limited speed and budget). The available hardware does not allow us to investigate additional attacks or pruning techniques. This may open up more future paths, such as injecting iterative attacks (L2), while considering more powerful GPUs, such as NVIDIA RTX 4090. The significant improvement against L-infinity attacks (results before and after adversarial training) illustrates that the attacks included are important and can significantly harm the model performance, indicating that these attacks should be considered. Future workers could explore other pruning methods (such as magnitude pruning, Taylor pruning, and Soft pruning) to enhance efficiency and produce more robust architectures, while exploring more adversarial attacks and scenarios.</p>
</sec>
<sec id="s7">
<label>7</label>
<title>Conclusion</title>
<p>This study introduced GreenShield, a novel framework that produces a lightweight Vision Transformer with competitive adversarial robustness compared to existing lightweight ViT variants (Mobile-ViT, ViT-Tiny, and ViT-Small). Major gaps in prior OCT studies are addressed, including sustainability, robustness against white-box L-infinity attacks, and the development of reliable preprocessing pipelines. The proposed approach demonstrated that the ViT model can be lightweight and robust without degrading generalization performance. The proposed framework integrates gradient-based importance pruning for model compression and a robust adversarial training pipeline incorporating correct ImageNet normalization and anti-saturation techniques to mitigate softmax saturation. This approach is supported by a strong evaluation plan, including FGSM, PGD, BIM, and the proposed hybrid FGSM-PGD attack. The findings demonstrated significant efficiency enhancements, reducing FLOPs, parameters, carbon footprint emissions, model size, inference time, and memory usage by around 45%&#x2013;50% across multiple efficiency metrics, while preserving generalization performance. In addition, the robustness of the GreenShield-ViT model significantly outperformed other lightweight ViT variants across all adversarial scenarios. A significant reduction in the attack&#x2019;s success rate across three datasets (OCT-c8, OCTID, and UCSD-3), maintaining strong accuracy under multiple adversarial attacks while preserving generalization, and achieving dramatic improvements in calibration. This transforms the model from an overconfident system (with low accuracy) into a well-calibrated predictor. This paper presents a robust and lightweight model that can be extended to other medical imaging tasks. Future researchers can explore more pruning techniques to generate more efficient ViT architecture. In addition, this work can be extended to incorporate additional adversarial attacks, such as L2 black-box or white-box attacks.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to acknowledge Kaggle for providing the computational environment (a GPU P100) used for training and evaluation, and the providers of the OCT datasets.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Munthir Qasaimeh, Mostafa Ali and Qasem Abu Al-Haija; Methodology, Munthir Qasaimeh, Mostafa Ali and Qasem Abu Al-Haija; Experiments and formal analysis, Munthir Qasaimeh; Writing&#x2014;original draft, Munthir Qasaimeh; Writing&#x2014;review &#x0026; editing, Munthir Qasaimeh, Mostafa Ali and Qasem Abu Al-Haija; Visualization, Munthir Qasaimeh; Supervision, Mostafa Ali and Qasem Abu Al-Haija. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data supporting the findings of this study are publicly available in the Kaggle repository. The Retinal OCT C-8 dataset is available at: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/obulisainaren/retinal-oct-c8.">https://www.kaggle.com/datasets/obulisainaren/retinal-oct-c8</ext-link>. The UCSD-3 OCT dataset is available at: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/mmazizi/ucsd-3-class-labeled-retinal-oct-images.">https://www.kaggle.com/datasets/mmazizi/ucsd-3-class-labeled-retinal-oct-images</ext-link>. Source code is available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/muntherqasaimeh/VTwork.">https://github.com/muntherqasaimeh/VTwork</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable. This study used a publicly available dataset and did not involve data collection from human participants or animals.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Swanson</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>CP</given-names></string-name>, <string-name><surname>Schuman</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Stinson</surname> <given-names>WG</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Optical coherence tomography</article-title>. <source>Science</source>. <year>1991</year>;<volume>254</volume>(<issue>5035</issue>):<fpage>1178</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1126/science.1957169</pub-id>; <pub-id pub-id-type="pmid">1957169</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Beyer</surname> <given-names>L</given-names></string-name>, <string-name><surname>Kolesnikov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Weissenborn</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Unterthiner</surname> <given-names>T</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>An image is worth 16 x 16 words: transformers for image recognition at scale</article-title>. In: <conf-name>Proceedings of the International Conference on Learning Representations; 2021 May 3&#x2013;7</conf-name>; <publisher-loc>Virtual</publisher-loc>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Takahashi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sakaguchi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Kouno</surname> <given-names>N</given-names></string-name>, <string-name><surname>Takasawa</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ishizu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Akagi</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Comparison of vision transformers and convolutional neural networks in medical image analysis: a systematic review</article-title>. <source>J Med Syst</source>. <year>2024</year>;<volume>48</volume>(<issue>1</issue>):<fpage>84</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10916-024-02105-8</pub-id>; <pub-id pub-id-type="pmid">39264388</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Gupta</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kapoor</surname> <given-names>M</given-names></string-name>, <string-name><surname>Debnath</surname> <given-names>SK</given-names></string-name></person-group>. <chapter-title>Challenges and risks of AI-enabled healthcare security</chapter-title>. In: <source>Artificial intelligence-enabled security for healthcare systems: safeguarding patient data and improving services</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2025</year>. p. <fpage>101</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-82810-2_6</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ahmed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Al Arafat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Najafi</surname> <given-names>D</given-names></string-name>, <string-name><surname>Mahmood</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rizve</surname> <given-names>MN</given-names></string-name>, <string-name><surname>Al Nahian</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>DeepCompress-ViT: rethinking model compression to enhance efficiency of vision transformers at the edge</article-title>. In: <conf-name>Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2025 Jun 10&#x2013;17</conf-name>; <publisher-loc>Nashville, TN, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52734.2025.02806</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Shlens</surname> <given-names>J</given-names></string-name>, <string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Explaining and harnessing adversarial examples</article-title>. In: <conf-name>Proceedings of the 3rd International Conference on Learning Representations; 2015 May 7&#x2013;9</conf-name>; <publisher-loc>San Diego, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liew</surname> <given-names>A</given-names></string-name>, <string-name><surname>Agaian</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Comprehensive survey of OCT-based disorders diagnosis: from feature extraction methods to robust security frameworks</article-title>. <source>Bioengineering</source>. <year>2025</year>;<volume>12</volume>(<issue>9</issue>):<fpage>914</fpage>. doi:<pub-id pub-id-type="doi">10.3390/bioengineering12090914</pub-id>; <pub-id pub-id-type="pmid">41007159</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kanca</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ayas</surname> <given-names>S</given-names></string-name>, <string-name><surname>Baykal Kablan</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ekinci</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Evaluating and enhancing the robustness of vision transformers against adversarial attacks in medical imaging</article-title>. <source>Med Biol Eng Comput</source>. <year>2025</year>;<volume>63</volume>(<issue>3</issue>):<fpage>673</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11517-024-03226-5</pub-id>; <pub-id pub-id-type="pmid">39453557</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pavithra</surname> <given-names>KC</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>P</given-names></string-name>, <string-name><surname>Geetha</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bhandary</surname> <given-names>SV</given-names></string-name>, <string-name><surname>Ajitha Shenoy</surname> <given-names>KB</given-names></string-name>, <string-name><surname>Rao</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Transformer-based DME classification using retinal OCT images without data augmentation: an evaluation of ViT-B16 and ViT-B32 with optimizer impact</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>180781</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2025.3620945</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>P</given-names></string-name></person-group>. <article-title>A novel transformer model with multiple instance learning for diabetic retinopathy classification</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>6768</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3351473</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Krizhevsky</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sutskever</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name></person-group>. <chapter-title>ImageNet classification with deep convolutional neural networks</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Pereira</surname> <given-names>F</given-names></string-name>, <string-name><surname>Burges</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Bottou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Weinberger</surname> <given-names>KQ</given-names></string-name></person-group>, editors. <source>Communications of the ACM</source>. <publisher-loc>Santa Clara, CA, USA</publisher-loc>: <publisher-name>Curran Associates, Inc</publisher-name>.; <year>2017</year>. p. <fpage>84</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Du</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Noisy softmax: improving the generalization ability of DCNN via postponing the early softmax saturation</article-title>. In: <conf-name>Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21&#x2013;26</conf-name>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2017.428</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pearce</surname> <given-names>T</given-names></string-name>, <string-name><surname>Brintrup</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Understanding softmax confidence and uncertainty</article-title>. <comment>arXiv:2106.04972. 2021</comment>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Likhomanenko</surname> <given-names>T</given-names></string-name>, <string-name><surname>Littwin</surname> <given-names>E</given-names></string-name>, <string-name><surname>Busbridge</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ramapuram</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Stabilizing transformer training by preventing attention entropy collapse</article-title>. In: <conf-name>Proceedings of the 40th International Conference on Machine Learning; 2023 Jul 23&#x2013;29</conf-name>; <publisher-loc>Honolulu, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Bharath Kumar</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Dunston</surname> <given-names>SD</given-names></string-name>, <string-name><surname>Rajam</surname> <given-names>VMA</given-names></string-name></person-group>. <chapter-title>Analysis of the impact of white box adversarial attacks in ResNet while classifying retinal fundus images</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Owoc</surname> <given-names>ML</given-names></string-name>, <string-name><surname>Sicily</surname> <given-names>FEV</given-names></string-name>, <string-name><surname>Rajaram</surname> <given-names>K</given-names></string-name>, <string-name><surname>Balasundaram</surname> <given-names>P</given-names></string-name></person-group>, editors. <source>Computational intelligence in data science</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2022</year>. p. <fpage>162</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-16364-7_13</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhandari</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shahi</surname> <given-names>TB</given-names></string-name>, <string-name><surname>Neupane</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Evaluating retinal disease diagnosis with an interpretable lightweight CNN model resistant to adversarial attacks</article-title>. <source>J Imag</source>. <year>2023</year>;<volume>9</volume>(<issue>10</issue>):<fpage>219</fpage>. doi:<pub-id pub-id-type="doi">10.3390/jimaging9100219</pub-id>; <pub-id pub-id-type="pmid">37888326</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Umer</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Sharif</surname> <given-names>M</given-names></string-name>, <string-name><surname>Raza</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kadry</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A deep feature fusion and selection-based retinal eye disease detection from OCT images</article-title>. <source>Expert Syst</source>. <year>2023</year>;<volume>40</volume>(<issue>6</issue>):<fpage>e13232</fpage>. doi:<pub-id pub-id-type="doi">10.1111/exsy.13232</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Anvesh</surname> <given-names>K</given-names></string-name>, <string-name><surname>Reshmi</surname> <given-names>BM</given-names></string-name>, <string-name><surname>Hariharan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Reddy</surname> <given-names>HV</given-names></string-name>, <string-name><surname>Krishnamoorthy</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kukreja</surname> <given-names>V</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A novel approach deep learning framework for automatic detection of diseases in retinal fundus images</article-title>. <source>Comput Model Eng Sci</source>. <year>2025</year>;<volume>143</volume>(<issue>2</issue>):<fpage>1485</fpage>&#x2013;<lpage>517</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2025.063239</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ak&#x00E7;a</surname> <given-names>S</given-names></string-name>, <string-name><surname>Garip</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ekinci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Atban</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Automated classification of choroidal neovascularization, diabetic macular edema, and drusenfrom retinal OCTimages using vision transformers: a comparative study</article-title>. <source>Lasers Med Sci</source>. <year>2024</year>;<volume>39</volume>(<issue>1</issue>):<fpage>140</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10103-024-04089-w</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Miao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>A lightweight model for the retinal disease classification using optical coherence tomography</article-title>. <source>Biomed Signal Process Control</source>. <year>2025</year>;<volume>101</volume>:<fpage>107146</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2024.107146</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>K</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Retinal OCT image classification based on MGR-GAN</article-title>. <source>Med Biol Eng Comput</source>. <year>2025</year>;<volume>63</volume>(<issue>6</issue>):<fpage>1749</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11517-025-03286-1</pub-id>; <pub-id pub-id-type="pmid">39862318</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Palaniappan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tak</surname> <given-names>TK</given-names></string-name>, <string-name><surname>Vijayan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Maram</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kshirsagar</surname> <given-names>PR</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Enhancement of medical imaging technique for diabetic retinopathy: realistic synthetic image generation using GenAI</article-title>. <source>Comput Model Eng Sci</source>. <year>2025</year>;<volume>145</volume>(<issue>3</issue>):<fpage>4107</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2025.073387</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rahman</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fallah</surname> <given-names>A</given-names></string-name>, <string-name><surname>Doss</surname> <given-names>R</given-names></string-name>, <string-name><surname>Karmakar</surname> <given-names>C</given-names></string-name></person-group>. <article-title>RAD-IoMT: robust adversarial defence mechanisms for IoMT medical image analysis</article-title>. <source>Ad Hoc Netw</source>. <year>2025</year>;<volume>178</volume>:<fpage>103935</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.adhoc.2025.103935</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Naren</surname> <given-names>OS</given-names></string-name></person-group>. <article-title>Retinal OCT image classification&#x2014;C8 [Internet]</article-title>. <comment>2024 [cited 2026 Apr 6]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/dsv/9595300">https://www.kaggle.com/dsv/9595300</ext-link>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gholami</surname> <given-names>P</given-names></string-name>, <string-name><surname>Roy</surname> <given-names>P</given-names></string-name>, <string-name><surname>Parthasarathy</surname> <given-names>MK</given-names></string-name>, <string-name><surname>Lakshminarayanan</surname> <given-names>V</given-names></string-name></person-group>. <article-title>OCTID: optical coherence tomography image database</article-title>. <source>Comput Electr Eng</source>. <year>2020</year>;<volume>81</volume>:<fpage>106532</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compeleceng.2019.106532</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Goldbaum</surname> <given-names>DKKZM</given-names></string-name></person-group>. <article-title>Large dataset of labeled optical coherence tomography (OCT) and chest X-ray images</article-title>. <source>Mendeley Data</source>. <year>2018</year>. doi:<pub-id pub-id-type="doi">10.17632/rscbjbr9sj.3</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>W</given-names></string-name>, <string-name><surname>Manoranjan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <etal>et al.</etal></person-group> <source>Image and video technology</source>. <edition>1st ed</edition>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-26431-3</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Benz</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ham</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Karjauv</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kweon</surname> <given-names>IS</given-names></string-name></person-group>. <article-title>Adversarial robustness comparison of vision transformer and MLP-mixer to CNNs</article-title>. <comment>arXiv:2110.02797. 2021</comment>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Madry</surname> <given-names>A</given-names></string-name>, <string-name><surname>Makelov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tsipras</surname> <given-names>D</given-names></string-name>, <string-name><surname>Vladu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Towards deep learning models resistant to adversarial attacks</article-title>. In: <conf-name>Proceedings of the 6th International Conference on Learning Representations 2018; 2018 Apr 30&#x2013;May 3</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kurakin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Adversarial examples in the physical world</article-title>. In: <conf-name>Proceedings of the 5th International Conference on Learning Representations 2017; 2017 Apr 24&#x2013;26</conf-name>; <publisher-loc>Toulon, France</publisher-loc>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Pleiss</surname> <given-names>G</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Weinberger</surname> <given-names>KQ</given-names></string-name></person-group>. <article-title>On calibration of modern neural networks</article-title>. In: <conf-name>Proceedings of the 34th International Conference on Machine Learning (ICML 2017); 2017 Aug 6&#x2013;11</conf-name>; <publisher-loc>Sydney, NSW, Australia</publisher-loc>. p. <fpage>1321</fpage>&#x2013;<lpage>30</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Training and inference time efficiency assessment framework for machine learning algorithms: a case study for hyperspectral image classification</article-title>. <source>Int J Appl Earth Obs Geoinf</source>. <year>2025</year>;<volume>141</volume>:<fpage>104591</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jag.2025.104591</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Jian</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>HMO: host memory optimization for model inference acceleration on edge devices</article-title>. In: <conf-name>Proceedings of the 2024 IEEE International Conference on Systems, Man, and Cybernetics (SMC); 2024 Oct 6&#x2013;10</conf-name>; <publisher-loc>Kuching, Malaysia</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/SMC54092.2024.10831215</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lannelongue</surname> <given-names>L</given-names></string-name>, <string-name><surname>Grealey</surname> <given-names>J</given-names></string-name>, <string-name><surname>Inouye</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Green algorithms: quantifying the carbon footprint of computation</article-title>. <source>Adv Sci</source>. <year>2021</year>;<volume>8</volume>(<issue>12</issue>):<fpage>2100707</fpage>. doi:<pub-id pub-id-type="doi">10.1002/advs.202100707</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lou</surname> <given-names>T</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Z</given-names></string-name></person-group>. <chapter-title>Enhancing adversarial robustness via anomaly-aware adversarial training</chapter-title>. In: <source>Knowledge science, engineering and management</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>328</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-40283-8_28</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>He</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Song</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Spatially transformed adversarial examples</article-title>. In: <conf-name>Proceedings of the 6th International Conference on Learning Representations (ICLR 2018); 2018 Apr 30&#x2013;May 3</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>.</mixed-citation></ref>
</ref-list>
</back></article>