<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">76888</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.076888</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>DRIVE: Diagnostic Report Integration via VLM and LLM Explanations for Explainable Vehicle Engine Fault Diagnosis</article-title>
<alt-title alt-title-type="left-running-head">DRIVE: Diagnostic Report Integration via VLM and LLM Explanations for Explainable Vehicle Engine Fault Diagnosis</alt-title>
<alt-title alt-title-type="right-running-head">DRIVE: Diagnostic Report Integration via VLM and LLM Explanations for Explainable Vehicle Engine Fault Diagnosis</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Lee</surname><given-names>Jaeseung</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Rew</surname><given-names>Jehyeok</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>jhrew@duksung.ac.kr</email></contrib>
<aff id="aff-1"><label>1</label><institution>School of Electrical Engineering, Korea University</institution>, <addr-line>Seoul</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Data Science, Duksung Women&#x2019;s University</institution>, <addr-line>Seoul</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Jehyeok Rew. Email: <email>jhrew@duksung.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>4</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>1</issue>
<elocation-id>22</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>11</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>01</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_76888.pdf"></self-uri>
<abstract>
<p>The engine serves as the primary component that generates power and drives vehicle movement. Given its critical role, accurately diagnosing engine faults is essential for ensuring vehicle safety and reliability. Recent advances in machine learning (ML) have enabled the development of artificial intelligence (AI)-based diagnostic models with strong predictive performance. However, the lack of transparency in these models constrains user confidence in their diagnostic outcomes. While explainable AI (XAI) methods such as local interpretable model-agnostic explanations (LIME) and Shapley additive explanations (SHAP) have been introduced to improve interpretability, their reliance on visual outputs requires manual interpretation, which can be inefficient and prone to subjectivity. To address this limitation, we propose DRIVE, a novel method for explainable vehicle engine fault diagnosis. In DRIVE, LIME and SHAP are applied to an ML-based diagnostic model, and their visual outputs are translated into textual explanations using the vision-language models (VLMs). These complementary explanations are then synthesized by a large language model (LLM) into a unified diagnostic report, providing a coherent narrative of the model&#x2019;s reasoning and emphasizing abnormal input features. Experiments conducted on a publicly available vehicle engine fault dataset demonstrate that DRIVE not only produces accurate and transparent diagnostic rationales but also generates structured reports that enhance usability for domain experts. By integrating multiple XAI methods with multimodal LLMs, DRIVE advances the transparency, trustworthiness, and practicality of AI-driven vehicle engine fault diagnosis.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Vehicle engine</kwd>
<kwd>fault diagnosis</kwd>
<kwd>vision-language model</kwd>
<kwd>large language model</kwd>
<kwd>explainable artificial intelligence</kwd>
<kwd>local interpretable model-agnostic explanations</kwd>
<kwd>Shapley additive explanations</kwd>
<kwd>energy</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Research Foundation of Korea (NRF) grant funded by the Korea government</funding-source>
<award-id>RS-2025-00516023</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The internal combustion engine is widely used in modern vehicles, serving as a core component that governs vehicle performance, reliability, and operational efficiency [<xref ref-type="bibr" rid="ref-1">1</xref>]. Despite its central role, the engine remains vulnerable to various types of operational failures, including fuel delivery imbalances, ignition system malfunctions, and intake airflow disturbances [<xref ref-type="bibr" rid="ref-2">2</xref>]. These faults emerge subtly in their early stages, making them difficult to detect without intelligent monitoring systems. As a result, such malfunctions can compromise the safety of drivers and passengers, elevate maintenance costs, and cause financial losses due to unexpected downtime [<xref ref-type="bibr" rid="ref-3">3</xref>]. Moreover, undetected or unresolved faults result in excessive emissions, leading to environmental degradation and non-compliance with regulatory standards. In this context, the development of advanced diagnostic methods for early fault detection and predictive maintenance has emerged as a key area for both vehicle manufacturers and owners [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. Beyond vehicle-level diagnostics, accurate engine fault diagnosis plays an important role in energy systems by improving fuel efficiency, reducing unnecessary energy consumption, and mitigating excessive emissions [<xref ref-type="bibr" rid="ref-6">6</xref>]. From a broader energy-system perspective, data-driven diagnostic frameworks contribute to more efficient utilization of energy resources and support sustainable operation of transportation-related energy infrastructures [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Early approaches to vehicle engine fault diagnosis primarily relied on rule-based systems derived from domain-specific expert knowledge [<xref ref-type="bibr" rid="ref-8">8</xref>]. While these methods were effective in identifying predefined fault patterns, they struggled to detect novel or previously unseen anomalies. Furthermore, they required continuous expert involvement for maintenance and updates, which limited their scalability and increased operational costs. As a result, there has been a growing transition toward data-driven diagnostic frameworks that leverage advanced computational intelligence technologies to improve fault detection and decision-making.</p>
<p>Recent advancements in machine learning (ML) have significantly enhanced vehicle engine fault diagnosis by enabling accurate and efficient data-driven fault classification [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. In parallel, more advanced deep learning-based diagnostic frameworks have been proposed to address challenges such as data scarcity, and multi-sensor heterogeneity in complex mechanical systems [<xref ref-type="bibr" rid="ref-10">10</xref>]. Several studies have explored few-shot and multimodal large model approaches to improve fault diagnosis performance under limited labeled data, as well as graph-based and channel-adaptive feature methods to robustly integrate multi-sensor signals [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Despite their performance gains, most ML-based diagnostic models are composed of complex internal structures, making it difficult to interpret how individual input features influence diagnostic outcomes. As a result, they are frequently characterized as &#x2018;black box&#x2019; models [<xref ref-type="bibr" rid="ref-13">13</xref>]. Moreover, their reliance on high-dimensional sensor data poses additional challenges for effective visualization and intuitive understanding, thereby hindering root-cause analysis and model refinement when incorrect predictions occur [<xref ref-type="bibr" rid="ref-14">14</xref>]. These limitations highlight the growing need for interpretable and trustworthy diagnostic reasoning that can explain model decisions beyond predictive accuracy alone.</p>
<p>Explainable artificial intelligence (XAI) has recently emerged as a key approach for improving the transparency and trustworthiness of ML models [<xref ref-type="bibr" rid="ref-15">15</xref>]. XAI refers to a set of techniques that aim to interpret how an ML model generates its outputs and to present these interpretations in a human-understandable form. By elucidating the underlying mechanisms behind model predictions, XAI facilitates greater interpretability and fosters user trust in automated decision systems [<xref ref-type="bibr" rid="ref-16">16</xref>]. Extensive research has been conducted to enhance the explainability of ML models, including Local Interpretable Model-agnostic Explanation (LIME) [<xref ref-type="bibr" rid="ref-17">17</xref>], Shapley Additive Explanations (SHAP) [<xref ref-type="bibr" rid="ref-18">18</xref>], and other post-hoc interpretability methods [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
<p>However, existing XAI methods have primarily focused on visualizing model explanations through graphical representations, which require users to manually interpret visual outputs to understand the model&#x2019;s reasoning. Given that users tend to summarize these numerical data and graphs in textual form, it is essential to present model explanations in a more accessible and interpretable format [<xref ref-type="bibr" rid="ref-21">21</xref>]. Enhancing the explainability of XAI-derived insights through concrete textual descriptions can significantly support users in understanding and utilizing the model&#x2019;s predictions. Furthermore, it is essential to interpret and compare model behaviors using multiple XAI methods, such as LIME and SHAP [<xref ref-type="bibr" rid="ref-22">22</xref>&#x2013;<xref ref-type="bibr" rid="ref-24">24</xref>]. LIME provides localized and instance-level explanations, while SHAP offers consistent feature attributions grounded in cooperative game theory. Combining these complementary methods enables a more comprehensive understanding of the model&#x2019;s decision making and ensures that insights are not biased toward a single method. This comparative perspective ultimately strengthens the reliability and practical applicability of XAI in safety-critical domains such as vehicle engine fault diagnosis.</p>
<p>With recent advancements in multimodal artificial intelligence, vision-language models (VLMs) have attracted growing attention [<xref ref-type="bibr" rid="ref-25">25</xref>]. These models are trained on large-scale datasets that integrate visual and textual information, enabling them to capture cross-modal relationships and semantic alignments between images and language. Beyond simple pattern recognition, VLMs exhibit strong reasoning capabilities across modalities, allowing them to generate natural language descriptions, provide context-aware interpretations, and support decision-making tasks [<xref ref-type="bibr" rid="ref-26">26</xref>]. These characteristics make them particularly suitable for applications that require both visual and textual understanding, such as interpreting model outputs and conveying diagnostic insights. In the context of vehicle engine fault diagnosis, these capabilities allow VLMs to translate visual outputs from XAI methods, such as LIME or SHAP, into concrete textual narratives that highlight abnormal input features and clarify diagnostic classifications.</p>
<p>Complementing these multimodal models, large language models (LLMs) have recently emerged as powerful tools for advanced textual reasoning and knowledge integration [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]. Trained on massive corpora of text, LLMs can generate coherent narratives, contextualizing multimodal evidence, and consolidating diverse diagnostic information into structured and human-interpretable explanations [<xref ref-type="bibr" rid="ref-29">29</xref>]. In particular, LLMs are well-suited to transform fragmented insights from XAI methods and VLM outputs into comprehensive diagnostic reports that enhance interpretability, coherence, and practical decision support.</p>
<p>In this paper, we propose DRIVE, a novel method for enhancing the explainability of vehicle engine fault diagnosis. The method begins with the development of an ML-based diagnostic model using extreme gradient boosting (XGBoost) [<xref ref-type="bibr" rid="ref-30">30</xref>] to predict potential engine faults. To interpret the model&#x2019;s diagnostic decisions, DRIVE integrates two complementary XAI methods: LIME and SHAP. The visual explanations generated by these methods are processed by a VLM, which translates the graphical outputs into coherent, human-readable textual descriptions. These descriptions improve interpretability by explicitly clarifying how abnormal input features contribute to the diagnostic predictions. Finally, DRIVE employs an LLM-based unified report generator to integrate the prediction results, VLM-derived explanations, and cross-method insights into a single coherent diagnostic report. This unified report not only elucidates the reasoning behind each diagnosis but also provides a comprehensive summary that facilitates decision-making in vehicle engine fault management.</p>
<p>The main contributions of this paper are summarized as follows:</p>
<p>We propose DRIVE, a novel method for vehicle engine fault diagnosis that integrates VLM-based interpretation with LLM-based report generation to enhance the transparency and usability of diagnostic models.</p>
<p>We design a structured prompt template that enables VLMs to transform LIME and SHAP visual outputs into consistent and informative textual explanations. In addition, we develop a complementary prompt template for LLMs, allowing them to synthesize VLM-derived explanations with model predictions into a coherent, unified diagnostic report.</p>
<p>The proposed DRIVE method is evaluated on a publicly available vehicle engine fault diagnosis dataset, demonstrating its effectiveness in improving interpretability, generating coherent diagnostic reports, and supporting informed decision-making.</p>
<p>The rest of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> reviews related works on vehicle engine fault diagnosis, XAI, VLM, and LLM. <xref ref-type="sec" rid="s3">Section 3</xref> presents the proposed DRIVE method in detail. <xref ref-type="sec" rid="s4">Section 4</xref> describes the dataset used in the experiments along with the experimental settings. <xref ref-type="sec" rid="s5">Section 5</xref> reports the experimental results and provides an in-depth analysis of the findings. <xref ref-type="sec" rid="s6">Section 6</xref> provides a discussion of the main findings and outlines directions for future improvements. Finally, <xref ref-type="sec" rid="s7">Section 7</xref> concludes the paper by summarizing the main contributions and discussing potential directions for future research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<sec id="s2_1">
<label>2.1</label>
<title>Vehicle Engine Fault Diagnosis Using Machine Learning</title>
<p>ML models have demonstrated outstanding performance in diagnosing vehicle engine faults, primarily due to their ability to learn complex patterns from large-scale sensor data. This capability enables more accurate and automated fault detection compared with traditional rule-based methods.</p>
<p>Du and Wei [<xref ref-type="bibr" rid="ref-4">4</xref>] proposed a method for detecting vehicle engine faults using analytic hierarchy process (AHP) and neural network. They first used AHP to determine the optimal input variables for vehicle engine fault diagnosis. Then, they maximized the performance of the model by tuning the hyperparameters of the neural network through a gradual growth method. Nixon et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] proposed an ML approach to diesel engine health prognostics using engine controller data. They employed a random forest-based fault diagnosis model and analyzed the correlations between the input variables. Li et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] proposed a fault diagnosis for vehicle engine using an ensemble learning-based ML model. They combined the probability values derived from random forest, k-nearest neighbor, and XGBoost, and classified the fault type based on the highest value among the average probabilities. Akbal&#x0131;k et al. [<xref ref-type="bibr" rid="ref-2">2</xref>] used sound signals for detecting vehicle engine faults. They preprocessed the sound data using discrete wavelet transform and classified the fault types using an extreme learning machine.</p>
<p>However, these ML models are inherently limited in their ability to provide transparent explanations for their diagnostic decisions. In the context of vehicle engine fault diagnosis, it is essential not only to detect the presence of a fault but also to identify the specific input variables that contributed to the prediction, as such information supports appropriate repair and maintenance strategies. Conventional ML models function as black boxes, lacking interpretability regarding their internal reasoning processes. As a result, they fail to reveal which features played a decisive role in the diagnosis, thereby constraining their practical applicability in real-world maintenance scenarios where understanding the root cause of failure is critical.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Fault Diagnosis Using Explainable Artificial Intelligence</title>
<p>With recent advancements in XAI, fault diagnosis models that incorporate XAI methods have gained increasing attention. These approaches aim to identify the key input variables that significantly contribute to fault occurrences. By providing such insights, they enhance the interpretability of the diagnostic process, enabling users to better understand and trust model predictions.</p>
<p>Hasan et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] developed an explainable fault diagnosis model for bearing systems. They first used Boruta and Spearman&#x2019;s rank correlation coefficient to select key input variables. Then, they constructed k-nearest neighbor-based diagnosis model and analyzed the feature contributions of the model using SHAP. Jang et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed an XAI method for fault diagnosis of industrial processes. They interpreted the diagnosis model based on adversarial autoencoder using SHAP. They performed hierarchical clustering based on the local SHAP values of faulty samples to intuitively examine the influence of input variables. Zereen et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed an XAI method for fault diagnosis using audio data which was collected from machine sensors. They extracted features of the audio data in terms of time and frequency. Then, they constructed a fault diagnosis model using logistic regression and analyzed its explainability using SHAP and LIME, comparing the two methods.</p>
<p>As research in this domain progresses, the focus has gradually shifted from solely improving diagnostic accuracy to enhancing model transparency and usability in real-world applications. The integration of XAI methods into fault diagnosis enables more informed decision-making by bridging the gap between model outputs and human understanding. Such developments not only support practical maintenance strategies but also lay the foundation for more interpretable diagnostic systems.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Vision-Language Model</title>
<p>VLMs have recently emerged as powerful tools capable of jointly processing and reasoning over both visual and textual information. Although they were initially developed for multimodal applications such as image captioning and visual question answering, their use has expanded to areas involving structured data analysis and decision-support tasks across diverse domains. These models exhibit strong capabilities in extracting contextual indicators, generating human-readable explanations, and enhancing interpretability within complex data-driven systems.</p>
<p>Roberts et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] proposed Image2Struct, a benchmark designed to evaluate VLMs in extracting structured representations from images. Their method enables automated and quantitative assessment by reconstructing image structures (e.g., LaTeX, HTML) and comparing them with the original data using multiple similarity metrics. Evaluations across diverse domains, including webpages, mathematical equations, and musical scores, revealed significant performance variations, demonstrating the benchmark&#x2019;s effectiveness in distinguishing VLM capabilities. Xia et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] evaluated VLMs on spreadsheet comprehension through self-supervised tasks involving optical character recognition (OCR), spatial reasoning, and format perception. While VLMs exhibited strong OCR performance, they struggled with issues such as cell omission, misalignment, and limited spatial understanding, indicating the need for further improvements. Chen et al. [<xref ref-type="bibr" rid="ref-36">36</xref>] proposed Ocean-OCR, a large-scale VLM that achieves outstanding performance in OCR tasks, including document understanding, scene text recognition, and handwriting interpretation. Leveraging a native resolution vision transformer and extensive OCR datasets, Ocean-OCR outperformed professional systems such as TextIn and PaddleOCR while maintaining robust general multimodal reasoning capabilities.</p>
<p>Although VLMs have demonstrated remarkable progress in analyzing both image and text data, their application to interpreting XAI outputs remains underexplored. In particular, few studies have investigated the use of VLMs to generate textual explanations for vehicle engine fault diagnosis models by integrating insights from multiple XAI methods. To bridge this gap, our method leverages VLMs to translate LIME and SHAP outputs from ML-based diagnostic models into coherent textual explanations, which are subsequently synthesized by an LLM into a unified diagnostic report. This translation process mitigates the cognitive burden of interpreting complex visual outputs and enhances user accessibility.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Large Language Model</title>
<p>LLMs are advanced neural architectures trained on massive textual corpora, enabling strong capabilities in language comprehension and generation. Beyond producing coherent natural language, LLMs excel at synthesizing information from diverse sources into structured and comprehensive narratives. These capabilities make LLMs particularly well-suited for applications that require the integration of heterogeneous outputs into unified reports, thereby enhancing interpretability, knowledge accessibility, and practical decision support across domains.</p>
<p>Xie et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] investigated whether LLMs can simulate human trust behavior. Their study revealed that LLMs generally exhibit trust behavior under the framework of Trust Games, which are widely recognized in behavioral economics. Tu et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] evaluated the evolution of ChatGPT across time. They constructed ChatLog, an ever-updating dataset with large-scale records of long-form ChatGPT responses across 21 natural language processing (NLP) benchmarks. Their comprehensive evaluation demonstrated that most capabilities of ChatGPT have progressively improved over time, exhibiting a stepwise evolutionary pattern. Sui et al. [<xref ref-type="bibr" rid="ref-39">39</xref>] developed a benchmark to assess the capability of LLMs to interpret tabular data. Their study included 7 tasks, such as cell lookup and row retrieval, to evaluate the structural understanding capabilities of LLMs. They observed that performance varied depending on various input factors, including table input format, content order, role prompting, and partition marks.</p>
<p>Despite the advancements in ML-based fault diagnosis, the integration of LLMs for interpreting model predictions remains limited. In particular, their potential to consolidate heterogeneous explanation outputs into a unified and accessible format has not been extensively explored. In this paper, we propose a novel VLM-LLM integrated method in which VLMs first generate textual interpretations of XAI outputs such as LIME and SHAP, and LLMs subsequently synthesize these complementary explanations into a coherent diagnostic report. By combining the visual interpretability of XAI with the integrative narrative capabilities of LLMs, the proposed method alleviates the cognitive burden of interpretation and enhances decision support in vehicle engine fault diagnosis.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Method</title>
<p>This section describes the overall structure of the proposed method, DRIVE. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> presents the overall structure of the proposed method.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of the proposed method, DRIVE.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_76888-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Vehicle Engine Fault Diagnosis Using Machine Learning</title>
<p>The ML model is formulated as a classification system to diagnose vehicle engine faults based on sensor data. Specifically, we employ the XGBoost classifier [<xref ref-type="bibr" rid="ref-30">30</xref>], a gradient boosting-based model known for its outstanding predictive performance in fault diagnosis tasks [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>]. XGBoost constructs an ensemble of decision trees in a sequential manner, where each subsequent tree focuses on minimizing the residual errors of the preceding ones, thereby enhancing the overall model accuracy.</p>
<p>A notable advantage of XGBoost is its ability to effectively handle sparse data through sparsity-aware split findings. This mechanism enables robust learning even when the dataset contains incomplete or irregular sensor readings, which are common in real-world vehicle diagnostics. In addition, XGBoost supports both first- and second-order gradient optimization, allowing the model to capture not only the direction but also the curvature of the loss function. This feature improves convergence speed and enhances optimization stability, which is useful in complex fault pattern classification tasks. Another important strength is the inclusion of L1 and L2 regularization terms in the objective function. This helps control model complexity and reduce overfitting, thereby improving generalization to unseen data. Furthermore, XGBoost employs a weighted quantile sketch algorithm to efficiently select split points in large datasets. This reduces both memory consumption and computational cost, making the model suitable for processing large-scale vehicle sensor data.</p>
<p>To enhance the performance of vehicle engine fault diagnosis, we fine-tune key hyperparameters, including learning rate, tree depth, and the number of estimators using grid search. This tuning process is guided by validation loss, and early stopping is employed to prevent overfitting, thereby achieving a well-balanced trade-off between model complexity and generalization capability. As a result, the optimized XGBoost effectively discriminates among different engine fault types while maintaining high predictive reliability under varying input conditions.</p>
<p>Previous studies have validated the effectiveness of XGBoost in vehicle engine fault diagnosis. For instance, Tao et al. [<xref ref-type="bibr" rid="ref-42">42</xref>] applied XGBoost to detect misfire conditions in diesel engines, achieving outstanding accuracy compared to traditional ML models. Similarly, Pratap et al. [<xref ref-type="bibr" rid="ref-43">43</xref>] developed an AI-based vehicle fault diagnosis system that employed XGBoost to analyze two-wheeler vehicle signals, demonstrating strong robustness in real-world applications. These findings support our choice of XGBoost as the core classifier in the proposed method.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Vision-Language Model-Based LIME Analysis</title>
<p>LIME analysis [<xref ref-type="bibr" rid="ref-17">17</xref>] is employed to improve the explainability of the ML-based vehicle engine fault diagnosis model. LIME approximates the local behavior of a complex predictive model by generating perturbed samples around a given input and observing the resulting changes in predictions. It then fits a simple and interpretable linear model to these samples to identify the features that most strongly influence the original output. For example, when a model diagnoses vehicle engine faults, LIME helps determine whether the classification was driven by variations in specific attribute values or by recurring patterns within the input data.</p>
<p>Existing studies have predominantly presented the outputs of LIME in visual formats such as bar charts and decision rule paths, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. However, these visualizations require manual interpretation and conversion into textual explanations by the user, which can be both time-consuming and prone to subjective bias. To overcome these limitations, we employ a VLM to automatically transform the results of LIME analysis into structured and coherent textual explanations. This approach clarifies the model&#x2019;s reasoning for each specific prediction by providing interpretable descriptions of how individual input features contribute to the diagnostic outcome.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Visualization of LIME analysis.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_76888-fig-2.tif"/>
</fig>
<p>In this context, the effectiveness of our method critically depends on the design of prompts used for the VLM [<xref ref-type="bibr" rid="ref-44">44</xref>]. Since the VLM generates responses by interpreting prompts through its pre-trained knowledge, the quality and relevance of the output can vary significantly depending on the structure, wording, and contextual detail of the input prompt. Therefore, the design of prompts with adequate contextual information is essential to ensure accurate, reliable, and informative textual explanations.</p>
<p>Due to the length, the structured prompt designed to guide the VLM in generating textual explanations for the LIME analysis is provided in <xref ref-type="app" rid="app-1">Appendix A</xref>. The prompt is composed of four main components: a role description, a dataset description, a task specification, and an interpretation guideline.</p>
<p>The role description defines the function of the VLM as a domain expert in vehicle engine diagnostics, ML, and XAI. It instructs the model to analyze LIME-based explanations produced by a vehicle engine fault diagnosis model. The dataset description outlines the input and output variables used in the model. It also specifies the normalization method applied to the data and summarizes key statistical properties including mean, standard deviation, and median for each variable. The output variable is a categorical label representing the engine&#x2019;s operational state, and the description includes the characteristics associated with each class. This information provides essential contextual grounding for the VLM&#x2019;s interpretive reasoning. The task specification introduces the LIME visualization and presents its components, including prediction probabilities, decision rules, and local feature contributions. Lastly, the interpretation guideline provides detailed instructions for how the VLM should analyze the LIME output. It requests a structured explanation that addresses the relative influence of each variable, the prioritization of variables, the interpretation of variables with minimal contributions, and the overall decision-making logic.</p>
<p>In the prompt design, several key elements are presented in bold black fonts to emphasize critical components. These include the statistical summaries of the input data, the employed prediction model, the predicted class, the value of the highest prediction probability, and the LIME feature conditions with their contribution values. These elements are dynamically populated within the prompt according to the characteristics of each test instance. Moreover, numbered formats are utilized to represent major analysis steps that require sequential reasoning, whereas dash markers are used to denote supporting details or explanatory instructions within each step.</p>
<p>This structure helps the VLM generate context-specific explanations grounded in model reasoning. Building upon this structure, the overall prompt design is formulated to guide the VLM in generating technically accurate and analytically well-founded explanations. This enhances the interpretability of the diagnostic process, supporting its applicability in practical deployment contexts.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Vision-Language Model-Based SHAP Analysis</title>
<p>Similar to LIME analysis in <xref ref-type="sec" rid="s3_2">Section 3.2</xref>, SHAP analysis [<xref ref-type="bibr" rid="ref-18">18</xref>] is employed to improve the explainability of the ML-based vehicle engine fault diagnosis model. Based on cooperative game theory, SHAP assigns each feature a Shapley value that quantifies its individual contribution to a specific prediction, thereby providing both local and global perspectives on the model&#x2019;s decision-making process. For instance, in the context of vehicle engine fault diagnosis, SHAP can highlight which sensor readings or operational parameters strongly drive the model&#x2019;s diagnostic outcome.</p>
<p>A widely used visualization for this purpose is SHAP waterfall plot, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. This plot illustrates how the model&#x2019;s output gradually evolves from a baseline value by sequentially adding the positive and negative contributions of individual features until the final output is reached. It provides an intuitive decomposition of feature-level effects for a given instance, clearly indicating which variables increase or decrease the model&#x2019;s predicted outcome. However, similar to other SHAP visualizations, it requires manual interpretation, which may introduce subjective bias. To address this limitation, we employ a VLM to automatically transform SHAP waterfall plots into structured and coherent textual explanations, thereby improving interpretability and enhancing accessibility of the diagnostic reasoning process.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Visualization of SHAP analysis.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_76888-fig-3.tif"/>
</fig>
<p>As with LIME, the effectiveness of employing a VLM to interpret SHAP outputs largely depends on the design of prompts. Since the VLM generates textual explanations based on its pre-trained knowledge, variations in the phrasing, structure, and contextual completeness of prompts can influence the clarity, faithfulness, and reliability of the generated output. Accordingly, providing well-structured prompts enriched with sufficient contextual information is essential to ensure accurate, consistent, and meaningful interpretations of SHAP results.</p>
<p>The structured prompt for VLM-based SHAP analysis was adapted from the corresponding prompt designed for LIME analysis, with targeted modifications to reflect the characteristics of SHAP interpretation. The complete prompt is provided in <xref ref-type="app" rid="app-1">Appendix A</xref> due to its length. Similar to the prompt for LIME analysis, it is composed of four main components: a role description, a dataset description, a task specification, and an interpretation guideline.</p>
<p>The role description was retained with minor terminology updates to reference SHAP, and the dataset description remained unchanged since both tasks rely on the same input and output variables. The task specification was revised to capture the distinctive characteristics of SHAP waterfall plots, highlighting the stepwise decomposition of the prediction from the baseline value to the final output through sequential positive and negative feature contributions, along with their magnitudes and color-coded representations. Finally, the interpretation guideline instructs the VLM to generate a structured and rigorous explanation by identifying key positive and negative features, describing the cumulative contribution process, discussing low-impact variables, and analyzing the decision logic underlying the model&#x2019;s prediction.</p>
<p>This design ensures that the generated explanations remain locally focused and resistant to overgeneralization, thereby producing clear, reliable, and contextually grounded explanations suitable for diagnostic reporting. In doing so, the SHAP prompt complements the LIME-based approach, providing a consistent yet distinct framework for generating VLM-driven explanations. Together, these prompt structures establish a robust foundation for integrating multiple XAI methods into a unified interpretability pipeline.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Large Language Model-Based Report Generation</title>
<p>Although VLM-based analyses of LIME and SHAP individually provide valuable textual explanations, their outputs often highlight different aspects of the model&#x2019;s decision-making process. LIME focuses on local surrogate approximations derived from feature perturbations, while SHAP provides game-theoretic attributions that quantify cumulative feature contributions. As a result, relying on a single method may constrain interpretability and introduce bias. To overcome this limitation, we employ an LLM to synthesize the complementary insights obtained from both VLM-based LIME and SHAP analyses into a unified and coherent diagnostic report. This integration enables users to leverage the strengths of each explanation method while mitigating their respective limitations.</p>
<p>To achieve this objective, we designed a structured prompt for unified report generation. The complete prompt is provided in <xref ref-type="app" rid="app-1">Appendix A</xref> due to its length. The prompt follows the same four-part structure as those used in the VLM-based analysis stages: a role description, a dataset description, a task specification, and an interpretation guideline.</p>
<p>The role description defines the objective of the LLM as generating a consolidated diagnostic report that integrates VLM-based LIME and SHAP explanations. The dataset description provides context regarding the engine behavior and emission-related variables, their normalization, and the diagnostic output classes. This ensures that the LLM remains grounded in the domain-specific characteristics of the data. The task specification instructs the LLM to merge the two explanations into a single coherent narrative by consolidating overlapping evidence, reconciling differences, and preserving technical rigor. This step encourages the model to produce a holistic interpretation of each prediction instance rather than a simple presentation of LIME and SHAP outputs. Lastly, the interpretation guideline describes the structure of the final report. It requires the LLM to summarize the predicted fault class, highlight consistent feature contributions, incorporate method-specific insights, and present an integrated decision logic. The guideline further constrains the explanation to local interpretation, thereby ensuring clarity and reliability.</p>
<p>By explicitly encoding these elements, the unified report generation prompt ensures that the resulting explanations are not only coherent and comprehensive but also precise and academically rigorous, thereby making them well-suited for diagnostic reporting.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Setup</title>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset</title>
<p>To validate the effectiveness of the proposed method, we utilize the publicly available EngineFaultDB dataset [<xref ref-type="bibr" rid="ref-45">45</xref>]. EngineFaultDB is a comprehensive dataset designed for vehicle engine fault classification, comprising 55,999 entries across 14 distinct variables. The dataset was constructed using real-world data collected from the C14NE spark ignition engine [<xref ref-type="bibr" rid="ref-46">46</xref>] during vehicle acceleration. It contains measurements from both normal operations and multiple fault scenarios.</p>
<p><xref ref-type="table" rid="table-1">Table 1</xref> presents the input variables of the EngineFaultDB dataset. These input variables are designed to capture both the dynamic behavior of the engine during vehicle acceleration and its corresponding emission characteristics. The dataset includes key engine parameters including manifold absolute pressure (MAP), throttle position sensor (TPS), force, power, revolutions per minute (RPM), fuel consumption (measured in l/h and l/100 km), and speed. In addition, it incorporates exhaust gas emission variables, including carbon monoxide (CO), hydrocarbons (HC), carbon dioxide (CO<sub>2</sub>), oxygen (O<sub>2</sub>), lambda, and air-fuel ratio (AFR). These variables were measured using the non-dispersive gas analyzer (NGA) 6000 to ensure high precision data acquisition. All input variables are represented as continuous numerical values. <xref ref-type="table" rid="table-2">Table 2</xref> further summarizes the distribution of these variables by reporting their mean, standard deviation and quartiles, thereby providing a comprehensive statistical context for the subsequent analysis.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Input variables of Enginefaultdb dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Input Variable</th>
<th>Description</th>
<th>Variable Type</th>
</tr>
</thead>
<tbody>
<tr>
<td>Manifold absolute pressure (MAP)</td>
<td>Pressure within the intake manifold</td>
<td>Continuous</td>
</tr>
<tr>
<td>Throttle position sensor (TPS)</td>
<td>Position of the throttle</td>
<td>Continuous</td>
</tr>
<tr>
<td>Force</td>
<td>Torque or rotational force of the engine</td>
<td>Continuous</td>
</tr>
<tr>
<td>Power</td>
<td>Rate of the energy that is transferred in the engine</td>
<td>Continuous</td>
</tr>
<tr>
<td>Revolutions per minute (RPM)</td>
<td>Speed of the engine</td>
<td>Continuous</td>
</tr>
<tr>
<td>Fuel consumption l/h</td>
<td>Fuel consumption rate of the engine</td>
<td>Continuous</td>
</tr>
<tr>
<td>Fuel consumption l/100 km</td>
<td>Fuel efficiency of the engine over a given distance</td>
<td>Continuous</td>
</tr>
<tr>
<td>Speed (km/h)</td>
<td>Travel speed of the vehicle</td>
<td>Continuous</td>
</tr>
<tr>
<td>Carbon monoxide (CO)</td>
<td>Concentration of CO in the exhaust gases</td>
<td>Continuous</td>
</tr>
<tr>
<td>Hydrocarbons (HC)</td>
<td>Concentration of unburnt hydrocarbons in the exhaust</td>
<td>Continuous</td>
</tr>
<tr>
<td>Carbon dioxide (CO<sub>2</sub>)</td>
<td>Concentration of CO<sub>2</sub> in the exhaust</td>
<td>Continuous</td>
</tr>
<tr>
<td>Oxygen (O<sub>2</sub>)</td>
<td>Amount of oxygen in the exhaust</td>
<td>Continuous</td>
</tr>
<tr>
<td>Lambda</td>
<td>Air-fuel equivalence ratio</td>
<td>Continuous</td>
</tr>
<tr>
<td>Air-fuel ratio (AFR)</td>
<td>Ratio of air to fuel in the combustion chambers</td>
<td>Continuous</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Summary statistics of the input variables of Enginefaultdb dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Input Variable</th>
<th>Mean</th>
<th>Standard Deviation</th>
<th>Minimum Value</th>
<th>25%</th>
<th>50%</th>
<th>75%</th>
<th>Maximum Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Manifold absolute pressure (MAP)</td>
<td>1.83</td>
<td>0.84</td>
<td>0.45</td>
<td>1.22</td>
<td>1.54</td>
<td>1.94</td>
<td>4.55</td>
</tr>
<tr>
<td>Throttle position sensor (TPS)</td>
<td>1.40</td>
<td>0.91</td>
<td>0.38</td>
<td>0.90</td>
<td>1.01</td>
<td>1.26</td>
<td>4.05</td>
</tr>
<tr>
<td>Force</td>
<td>286.69</td>
<td>378.77</td>
<td>2.58</td>
<td>76.85</td>
<td>92.50</td>
<td>257.99</td>
<td>1537.12</td>
</tr>
<tr>
<td>Power</td>
<td>5.66</td>
<td>7.68</td>
<td>0.47</td>
<td>0.99</td>
<td>2.40</td>
<td>4.70</td>
<td>33.95</td>
</tr>
<tr>
<td>Revolutions per minute (RPM)</td>
<td>2398.05</td>
<td>932.01</td>
<td>1066.45</td>
<td>1830.20</td>
<td>2105.59</td>
<td>2761.37</td>
<td>5013.40</td>
</tr>
<tr>
<td>Fuel consumption l/h</td>
<td>4.50</td>
<td>2.22</td>
<td>1.92</td>
<td>2.98</td>
<td>3.82</td>
<td>5.14</td>
<td>14.81</td>
</tr>
<tr>
<td>Fuel consumption l/100 km</td>
<td>8.94</td>
<td>3.15</td>
<td>5.19</td>
<td>6.57</td>
<td>8.07</td>
<td>9.92</td>
<td>20.04</td>
</tr>
<tr>
<td>Speed (km/h)</td>
<td>51.69</td>
<td>20.14</td>
<td>22.76</td>
<td>39.41</td>
<td>45.39</td>
<td>59.51</td>
<td>107.54</td>
</tr>
<tr>
<td>Carbon monoxide (CO)</td>
<td>1.93</td>
<td>1.99</td>
<td>0.42</td>
<td>0.64</td>
<td>1.13</td>
<td>2.46</td>
<td>10.13</td>
</tr>
<tr>
<td>Hydrocarbons (HC)</td>
<td>188.45</td>
<td>111.05</td>
<td>1.79</td>
<td>158.80</td>
<td>178.26</td>
<td>203.68</td>
<td>975.66</td>
</tr>
<tr>
<td>Carbon dioxide (CO<sub>2</sub>)</td>
<td>13.04</td>
<td>1.05</td>
<td>8.65</td>
<td>12.84</td>
<td>13.24</td>
<td>13.64</td>
<td>15.13</td>
</tr>
<tr>
<td>Oxygen (O<sub>2</sub>)</td>
<td>0.59</td>
<td>0.22</td>
<td>0.20</td>
<td>0.41</td>
<td>0.53</td>
<td>0.79</td>
<td>1,15</td>
</tr>
<tr>
<td>Lambda</td>
<td>0.96</td>
<td>0.07</td>
<td>0.69</td>
<td>0.94</td>
<td>0.98</td>
<td>1.01</td>
<td>1.15</td>
</tr>
<tr>
<td>Air-fuel ratio (AFR)</td>
<td>14.17</td>
<td>0.97</td>
<td>10.21</td>
<td>13.78</td>
<td>14.37</td>
<td>14.82</td>
<td>16.89</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The proposed method classifies vehicle engine fault into 4 categories: no fault, rich mixture, lean mixture, and low voltage, as detailed in <xref ref-type="table" rid="table-3">Table 3</xref>. The no fault class represents standard operating conditions without any detectable anomalies. The rich mixture class corresponds to an excessive fuel-to-air ratio, leading to increased carbon monoxide emissions due to incomplete combustion. The lean mixture class indicates insufficient fuel supply, which results in elevated oxygen levels and unstable engine performance. Lastly, the low voltage class is related to electrical irregularities in the ignition or fuel systems, causing reduced combustion efficiency and abnormal sensor readings. The dataset comprises 16,000 normal samples, 10,998 rich mixture samples, 15,000 lean mixture samples, and 14,001 low voltage samples. It provides a diverse and representative foundation for evaluating the performance of the proposed vehicle engine fault diagnosis method.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Output variables of EngineFaultDB dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Class</th>
<th>Number of Counts</th>
<th>Conditions</th>
</tr>
</thead>
<tbody>
<tr>
<td>No fault</td>
<td>16,000</td>
<td>Normal state</td>
</tr>
<tr>
<td rowspan="6">Rich mixture</td>
<td rowspan="6">10,998</td>
<td>Incorrect sensor performance</td>
</tr>
<tr><td>High fuel pressure</td></tr>
<tr><td>Defective injector</td></tr>
<tr><td>Fault pressure regulator</td></tr>
<tr><td>Clogged air filter</td></tr>
<tr><td>Clogged fuel return line</td>
</tr>
<tr>
<td rowspan="1">Lean mixture</td>
<td rowspan="1">15,000</td>
<td>Incorrect sensor performance<break/>
Low fuel pressure<break/>
Defective injector<break/>
Faulty pressure regulator</td>
</tr>
<tr>
<td rowspan="1">Low voltage</td>
<td rowspan="1">14,001</td>
<td>Worn spark plugs<break/>
Faulty ignition cables<break/>
Defective coil<break/>
Faulty sensor wiring</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Settings</title>
<p>The performance of the proposed method was evaluated through comparative experiments with multiple ML-based classification models. All input variables were normalized to the range [0, 1] using min-max normalization [<xref ref-type="bibr" rid="ref-47">47</xref>]. The dataset was partitioned into training, validation, and test sets in proportions of 70%, 10%, and 20%, respectively. A total of 13 ML models were employed for the experiments, including logistic regression [<xref ref-type="bibr" rid="ref-48">48</xref>], k-nearest neighbor [<xref ref-type="bibr" rid="ref-49">49</xref>], multi-layer perceptron (MLP) classifier [<xref ref-type="bibr" rid="ref-50">50</xref>], random forest [<xref ref-type="bibr" rid="ref-51">51</xref>], decision tree [<xref ref-type="bibr" rid="ref-52">52</xref>], extratree classifier [<xref ref-type="bibr" rid="ref-53">53</xref>], AdaBoost [<xref ref-type="bibr" rid="ref-54">54</xref>], linear discriminant analysis [<xref ref-type="bibr" rid="ref-55">55</xref>], quadratic discriminant analysis [<xref ref-type="bibr" rid="ref-56">56</xref>], gradient boosting [<xref ref-type="bibr" rid="ref-57">57</xref>], LightGBM [<xref ref-type="bibr" rid="ref-58">58</xref>], CatBoost [<xref ref-type="bibr" rid="ref-59">59</xref>], and XGBoost [<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
<p>In <xref ref-type="sec" rid="s3_2">Sections 3.2</xref>&#x2013;<xref ref-type="sec" rid="s3_4">3.4</xref>, this study employed OpenAI GPT-4o [<xref ref-type="bibr" rid="ref-60">60</xref>]. GPT-4o is a multimodal LLM that can process and generate outputs across text and image inputs within a unified architecture. It has demonstrated high performance in a wide range of tasks, including natural language understanding, complex reasoning, knowledge retrieval, and cross-modal information processing. The model has been successfully applied to various domains, including scientific research and real-time human-computer interaction. In this study, GPT-4o was utilized to generate interpretable textual explanations and to support diagnostic reasoning in vehicle engine fault classification.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation Metric</title>
<p>To quantitatively evaluate the performance of the vehicle engine fault diagnosis model, we employed 4 widely used classification metrics: precision, recall, f1-score, and accuracy [<xref ref-type="bibr" rid="ref-61">61</xref>]. Precision quantifies the proportion of true positive predictions among all positive predictions made by the model. Recall measures the ability of the model to correctly identify positive cases from all actual positives. The F1 score, which is the harmonic mean of precision and recall, provides a balanced evaluation metric when both precision and recall are critical. Accuracy represents the overall ratio of correctly classified instances to the total number of samples, serving as a general measure of classification performance. The mathematical definitions of these metrics are presented in <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-4">(4)</xref>.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>To evaluate the quality of the textual explanations generated in <xref ref-type="sec" rid="s3_2">Sections 3.2</xref>&#x2013;<xref ref-type="sec" rid="s3_4">3.4</xref>, we adopted a dual evaluation framework that combines VLM-as-a-Judge [<xref ref-type="bibr" rid="ref-62">62</xref>] and human expert assessment [<xref ref-type="bibr" rid="ref-21">21</xref>]. The VLM-as-a-Judge employs a VLM to automatically assess the quality of generated explanations according to predefined criteria. In this study, OpenAI GPT-4o was used as the evaluation model due to its strong multimodal reasoning capabilities, while LLaMA 3.2-Vision was additionally employed as an independent evaluator to mitigate potential model-specific bias [<xref ref-type="bibr" rid="ref-63">63</xref>]. By aggregating evaluation outcomes across these two VLMs, we aim to enhance the robustness of the automated assessment. This automated approach ensures scalability and reproducibility, whereas human expert assessment provides domain-grounded insights that capture practical interpretability in real diagnostic contexts. By employing both evaluation methods under the same conditions, we leverage the complementary strengths of automation and expert judgment, thereby improving the robustness and reliability of the overall evaluation.</p>
<p>Both evaluations were conducted using a 5-point Likert scale [<xref ref-type="bibr" rid="ref-64">64</xref>] to ensure comparability, where 1 indicates strongly disagree and 5 indicates strongly agree. The evaluation criteria employed in this study are summarized in <xref ref-type="table" rid="table-4">Table 4</xref>. The prompt for the VLM-as-a-Judge is presented in <xref ref-type="table" rid="table-17">Table A4</xref>. The comprehensibility assesses whether the explanation is expressed in a clear and easily interpretable manner, while the faithfulness evaluates the degree to which the explanation accurately reflects the underlying visualization and its diagnostic meaning. The diagnostic relevance examines the alignment of the explanation with established knowledge in vehicle engine fault diagnosis, and the explanatory value evaluates the usefulness of the explanation in enhancing user understanding beyond visual analysis. Finally, the reliability captures the perceived credibility and trustworthiness of the explanation from the perspective of domain experts. Collectively, these criteria enable a comprehensive evaluation of both the technical fidelity and the practical interpretability of the generated reports, ensuring a balanced assessment of explanation quality.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Evaluation criteria of VLM-generated explanations assessed on a 5-point Likert scale.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Criterion</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>Comprehensibility</td>
<td>The degree to which the explanation is articulated in a clear and easily interpretable manner.</td>
</tr>
<tr>
<td>Faithfulness</td>
<td>The extent to which the explanation correctly represents the plot and its underlying information.</td>
</tr>
<tr>
<td>Domain relevance</td>
<td>The extent to which the explanation corresponds to established knowledge in engine fault diagnosis.</td>
</tr>
<tr>
<td>Explanatory value</td>
<td>The usefulness of the explanation in improving user understanding compared to the visualization alone.</td>
</tr>
<tr>
<td>Reliability</td>
<td>The perceived credibility and trustworthiness of the explanation from a domain expert&#x2019;s perspective.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Result</title>
<sec id="s5_1">
<label>5.1</label>
<title>Evaluation of Vehicle Engine Fault Diagnosis</title>
<p>Based on the experimental settings described in <xref ref-type="sec" rid="s4_2">Section 4.2</xref>, we evaluated the diagnostic performance of 13 ML models for vehicle engine fault diagnosis. <xref ref-type="table" rid="table-5">Table 5</xref> presents the results in terms of precision, recall, f1-score, and accuracy. For each metric, the best performing model is highlighted in bold, and the second-best model is underlined.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Fault diagnosis performance of the vehicle engine fault diagnosis model.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Precision</th>
<th>Recall</th>
<th>F1 Score</th>
<th>Accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td>Logistic regression</td>
<td>0.5605</td>
<td>0.5624</td>
<td>0.5610</td>
<td>0.5626</td>
</tr>
<tr>
<td>K-nearest neighbor</td>
<td>0.7590</td>
<td>0.7590</td>
<td>0.7590</td>
<td>0.7518</td>
</tr>
<tr>
<td>MLP classifier</td>
<td>0.7610</td>
<td>0.7593</td>
<td>0.7475</td>
<td>0.7481</td>
</tr>
<tr>
<td>Random forest</td>
<td>0.7575</td>
<td>0.7575</td>
<td>0.7575</td>
<td>0.7501</td>
</tr>
<tr>
<td>Decision tree</td>
<td>0.7583</td>
<td>0.7583</td>
<td>0.7582</td>
<td>0.7508</td>
</tr>
<tr>
<td>Extratree classifier</td>
<td>0.7554</td>
<td>0.7554</td>
<td>0.7553</td>
<td>0.7483</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>0.6584</td>
<td>0.6637</td>
<td>0.6561</td>
<td>0.6528</td>
</tr>
<tr>
<td>Linear discriminant analysis</td>
<td>0.5567</td>
<td>0.5473</td>
<td>0.5502</td>
<td>0.5517</td>
</tr>
<tr>
<td>Quadratic discriminant analysis</td>
<td>0.6832</td>
<td>0.6883</td>
<td>0.6510</td>
<td>0.6679</td>
</tr>
<tr>
<td>Gradient boosting</td>
<td>0.7569</td>
<td>0.7570</td>
<td>0.7568</td>
<td>0.7493</td>
</tr>
<tr>
<td>CatBoost</td>
<td>0.7535</td>
<td>0.7535</td>
<td>0.7535</td>
<td>0.7461</td>
</tr>
<tr>
<td>LightGBM</td>
<td><underline>0.7618</underline></td>
<td><underline>0.7618</underline></td>
<td><underline>0.7618</underline></td>
<td><underline>0.7546</underline></td>
</tr>
<tr>
<td>XGBoost</td>
<td><bold>0.7627</bold></td>
<td><bold>0.7627</bold></td>
<td><bold>0.7627</bold></td>
<td><bold>0.7555</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Among all evaluated models, the XGBoost-based fault diagnosis model achieved the highest performance across all evaluation metrics, with an overall accuracy of 75.55%. LightGBM, another gradient boosting-based ensemble model, ranked second with comparably strong results. These findings highlight the effectiveness of boosting-based ensemble methods in capturing complex, non-linear patterns inherent in vehicle engine fault data. In contrast, traditional statistical models such as logistic regression and linear discriminant analysis exhibited substantially lower performance, reflecting their limited capacity to model high-dimensional and non-linear sensor relationships.</p>
<p>It is worth noting that the EngineFaultDB dataset represents a challenging four-class classification problem involving noisy sensor measurements and overlapping fault characteristics. Under such realistic conditions, moderate accuracy levels are commonly observed, and perfect classification performance is difficult to achieve. In this context, the proposed method does not aim to provide deterministic or definitive fault judgments. Instead, it focuses on enhancing transparency by explaining how the diagnostic model arrives at a given prediction.</p>
<p>Overall, these results indicate that modern ensemble models, particularly XGBoost, provide robust diagnostic performance, while the interpretability method proposed in this study complements the classifier by exposing its reasoning process. This combination supports informed decision-making by domain experts, especially in scenarios where model predictions may involve inherent uncertainty.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Vision-Language Model-Based LIME Analysis</title>
<p>We evaluated the effectiveness of a VLM in generating textual explanations for LIME analysis results obtained from the vehicle engine fault diagnosis model. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> presents the LIME visualization of the XGBoost classifier, and <xref ref-type="table" rid="table-6">Table 6</xref> provides the corresponding VLM-based textual interpretation. In this instance, the model classified the instance as a lean mixture fault with a prediction probability of 84%.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Plot of LIME analysis.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_76888-fig-4.tif"/>
</fig><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Interpretation of LIME analysis using the VLM.</title>
</caption>
<table>
<colgroup>
<col align="left"/> </colgroup>
<thead>
<tr>
<th>Result</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Local interpretation of engine fault classification: lean mixture via LIME</bold></td>
</tr>
<tr>
<td><bold>1. Key features influencing the prediction</bold></td>
</tr>
<tr>
<td>The classifier assigned this instance to the <bold>Lean Mixture</bold> category with high confidence (<bold>probability &#x003D; 0.84</bold>). The LIME explanation provides the following key feature-based conditions contributing to this local decision:</td>
</tr>
<tr>
<td><bold>&#x2022;&#x2002;Positive contributions (supporting the Lean Mixture classification):</bold></td>
</tr>
<tr>
<td>&#x2003;&#x2002;o CO &#x2264; 0.02: <bold>&#x002B;0.1305</bold>&#x2014;Extremely low carbon monoxide levels indicate highly complete combustion, a &#x2003;&#x2002;typical signature of air-dominated (lean) fuel mixtures.</td>
</tr>
<tr>
<td>&#x2003;&#x2002;o 0.09 &#x003C; Consumption L/100 km &#x2264; 0.19: <bold>&#x002B;0.0759</bold>&#x2014;Low normalized fuel consumption suggests reduced fuel injection, consistent with lean operation.</td>
</tr>
<tr>
<td>&#x2003;&#x2002;o Speed &#x003E; 0.43: <bold>&#x002B;0.0686</bold>&#x2014;Elevated vehicle speed, when accompanied by low fuel consumption, &#x2003;&#x2002;supports the interpretation of lean combustion under driving load.</td>
</tr>
<tr>
<td>&#x2003;&#x2002;o Consumption L/H &#x003E; 0.25: <bold>&#x002B;0.0515</bold>&#x2014;High instantaneous fuel use in the presence of low distance-normalized consumption may reflect rapid air intake and combustion cycles under high RPM, aligning with lean dynamics.</td>
</tr>
<tr>
<td>&#x2002;o 0.16 &#x003C; HC &#x2264; 0.18: <bold>&#x002B;0.0344</bold>&#x2014;Slightly elevated hydrocarbon emissions may reflect incomplete burn in certain lean zones.</td>
</tr>
<tr>
<td>&#x2002;o 0.65 &#x003C; CO<sub>2</sub> &#x2264; 0.71: <bold>&#x002B;0.0245</bold>&#x2014;Moderate CO<sub>2</sub> emissions suggest relatively complete combustion, reinforcing the lean mixture hypothesis.</td>
</tr>
<tr>
<td>&#x2002;o 0.06 &#x003C; Power &#x2264; 0.12: <bold>&#x002B;0.0220</bold>&#x2014;Moderately low power output aligns with reduced fuel delivery in lean conditions.</td>
</tr>
<tr>
<td>&#x2002;o RPM &#x003E; 0.43: <bold>&#x002B;0.0202</bold>&#x2014;High engine speed supports lean operation under dynamic driving conditions.</td>
</tr>
<tr>
<td><bold>&#x2022;&#x2002;Negative contributions (countering the Lean Mixture classification):</bold></td>
</tr>
<tr>
<td>&#x2002;o 0.06 &#x003C; Force &#x2264; 0.17: <bold>&#x2212;0.0627</bold>&#x2014;Moderate torque may suggest that fuel delivery is not overly reduced, somewhat opposing the lean hypothesis.</td>
</tr>
<tr>
<td>&#x2002;o 0.35 &#x003C; O<sub>2</sub> &#x2264; 0.63: <bold>&#x2212;0.0379</bold>&#x2014;This oxygen range is indicative of possible sensor stabilization rather than extreme lean conditions, slightly dampening the classification confidence.</td>
</tr>
<tr>
<td><bold>2. Feature prioritization</bold></td>
</tr>
<tr>
<td>The classifier heavily emphasizes <bold>combustion byproducts (CO, HC, CO</bold><sub><bold>2</bold></sub><bold>) and fuel delivery metrics (L/100 km, L/H)</bold> to determine the local fault type. The most influential variable, <bold>CO &#x2264; 0.02</bold>, is well below the dataset mean (0.155), strongly indicating oxygen-rich combustion and thereby substantiating the lean classification.</td>
</tr>
<tr>
<td><bold>Vehicle dynamics</bold>&#x2014;specifically <bold>speed, power,</bold> and <bold>RPM</bold>&#x2014;are integrated into the decision, suggesting that the model incorporates operational context to differentiate between lean and other high-efficiency or fault modes.</td>
</tr>
<tr>
<td>Notably, <bold>L/100 km</bold> is far below the dataset median (0.193), reinforcing the conclusion that the vehicle is operating with minimal fuel expenditure, typical of lean engine behavior.</td>
</tr>
<tr>
<td><bold>3. Features with negative or low impact</bold></td>
</tr>
<tr>
<td>Several features exhibited marginal influence in this instance:</td>
</tr>
<tr>
<td>&#x2002;&#x2022;&#x2002;Power and RPM, while contributing positively, have relatively low weights (&#x002B;0.0220 and &#x002B;0.0202), indicating they act as reinforcing but non-decisive features.</td>
</tr>
<tr>
<td>&#x2002;&#x2022;&#x2002;MAP and TPS do not appear in the LIME explanation, suggesting they fell within typical operational ranges and did not provide discriminative value for this particular prediction.</td>
</tr>
<tr>
<td>&#x2002;&#x2022;&#x2002;Lambda and AFR are similarly absent, likely because their values did not deviate enough from the normative range to influence the local classification meaningfully.</td>
</tr>
<tr>
<td><bold>4. Logical decision logic</bold></td>
</tr>
<tr>
<td>The model&#x2019;s decision process, as approximated by LIME, hinges on the convergence of <bold>low emissions, high mechanical output,</bold> and <bold>low fuel consumption</bold>:</td>
</tr>
<tr>
<td>&#x2002;&#x2022;&#x2002;The strong presence of <bold>complete combustion indicators</bold> (low CO, moderate CO<sub>2</sub>) combined with <bold>low fuel consumption</bold> per kilometer and moderate consumption per hour supports the idea of a <bold>lean air-fuel mixture</bold>.</td>
</tr>
<tr>
<td>&#x2002;&#x2022;&#x2002;High <bold>speed and RPM</bold> suggest the engine is under load, and the model appears to infer that efficient combustion under load with minimal fuel injection is a diagnostic hallmark of lean operation.</td>
</tr>
<tr>
<td>&#x2002;&#x2022;&#x2002;The model balances this against potential counter-signals such as moderate torque and oxygen range, showing that it accounts for partial ambiguity in real-world conditions.</td>
</tr>
<tr>
<td>This nuanced decision boundary reflects the model&#x2019;s robustness in discerning lean mixture faults using a multifaceted evaluation of sensor and performance metrics.</td>
</tr>
<tr>
<td><bold>Conclusion</bold></td>
</tr>
<tr>
<td>The LIME-based interpretation of this XGBoost prediction confirms that the model identifies the <italic>Lean Mixture</italic> fault class based on coherent and domain-relevant signals: low CO, minimal fuel consumption, and sustained engine performance. The inclusion of weakly negative features demonstrates the model&#x2019;s attempt to account for real-world variability and enhances the interpretability of the decision-making process. This analysis reinforces the role of interpretable AI in improving diagnostic transparency and trustworthiness in automotive fault detection systems.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The explanation identified extremely low CO emissions as the most influential variable in classifying the instance as a lean mixture condition. This observation is consistent with established principles in combustion theory, where excess oxygen leads to reduced CO output. The second most influential factor was fuel consumption l/100 km, supporting the model&#x2019;s interpretation of reduced fuel injection, which is characteristic of lean operation. Additional contributing features included elevated vehicle speed and HC, which the VLM interpreted as compensatory responses to limited fuel delivery.</p>
<p>The analysis identified features with marginal contribution, such as power output and RPM, which served as supplementary indicators of engine activity but were not decisive for the lean mixture classification. It also noted the absence of variables such as MAP, TPS, Lambda, and AFR, whose values likely remained within normal operating ranges and thus offered limited discriminatory power in this instance.</p>
<p>The overall explanation revealed a coherent inference of the model&#x2019;s decision-making process. It demonstrated that the classification was based on the convergence of low emissions, high mechanical output, and low fuel consumption. The VLM effectively captured this logic through an interpretable narrative, clarifying how each input variable contributed to the final prediction. This supports the feasibility of combining LIME with VLM-based interpretation to improve transparency in vehicle engine fault diagnosis and to facilitate the communication of model reasoning in a user-accessible format. Additional results of the VLM-based LIME analysis can be found in the Supplementary Materials.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Vision-Language Model-Based SHAP Analysis</title>
<p>We further employed a VLM to generate textual explanations for the SHAP waterfall plot results of the vehicle engine fault diagnosis model. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> illustrates the SHAP waterfall plot corresponding to the same instance analyzed in <xref ref-type="sec" rid="s5_2">Section 5.2</xref>, and <xref ref-type="table" rid="table-7">Table 7</xref> presents the VLM-based interpretation of this result.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Plot of SHAP analysis.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_76888-fig-5.tif"/>
</fig><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Interpretation of SHAP analysis using the VLM.</title>
</caption>
<table>
<colgroup>
<col align="left"/> </colgroup>
<thead>
<tr>
<th>Result</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Instance-level interpretation of engine fault diagnosis: lean mixture case explained by SHAP waterfall plot</bold></td>
</tr>
<tr>
<td><bold>1. Key features</bold></td>
</tr>
<tr>
<td>The most influential contributors in this instance are:</td>
</tr>
<tr>
<td><list list-type="bullet">
<list-item>
<p><bold>Positive contributions (toward Lean Mixture classification)</bold>
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>CO (&#x002B;0.64)</bold>: Carbon monoxide levels provided the strongest positive push toward the Lean Mixture class, suggesting abnormal gas emissions are a decisive indicator.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>RPM (&#x002B;0.26)</bold>: Elevated engine revolutions per minute significantly increased the likelihood of lean operation.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>Fuel consumption L/H (&#x002B;0.20)</bold>: Excessive hourly fuel consumption pushed the model toward the fault class.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>CO</bold><sub><bold>2</bold></sub> <bold>(&#x002B;0.11)</bold> and <bold>consumption L/100 km (&#x002B;0.09)</bold>: Both fuel-efficiency and emission parameters further strengthened the Lean Mixture assignment.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>Speed (&#x002B;0.07)</bold> and <bold>HC (&#x002B;0.06)</bold>: Moderate positive effects indicated abnormal driving dynamics and hydrocarbon emissions consistent with a disturbed air&#x2013;fuel balance.</p></list-item>
</list></p></list-item>
<list-item>
<p><bold>Negative contributions (counteracting the Lean Mixture classification)</bold>
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>Force (&#x2212;0.06)</bold> and <bold>O</bold><sub><bold>2</bold></sub> <bold>(&#x2212;0.06)</bold>: Reduced engine torque and oxygen levels exerted the strongest negative effects, mitigating the Lean Mixture prediction.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>TPS (&#x2212;0.04)</bold>, <bold>Power (&#x2212;0.02)</bold>, and <bold>Lambda (&#x2212;0.02)</bold>: Throttle position, power output, and lambda ratio slightly counteracted the positive drivers.</p></list-item>
</list></p></list-item>
</list></td>
</tr>
<tr>
<td><bold>2. Stepwise contribution process</bold></td>
</tr>
<tr>
<td><list list-type="bullet">
<list-item>
<p>The model&#x2019;s baseline (expected) output for this class was <bold>0.743</bold>.</p></list-item>
<list-item>
<p>Sequential feature additions progressively moved the prediction upward:
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>CO (&#x002B;0.64)</bold> drove the largest jump, immediately shifting the score above 1.3.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>RPM (&#x002B;0.26)</bold> and <bold>Consumption L/H (&#x002B;0.20)</bold> compounded this upward trajectory, elevating the score to &#x007E;1.8.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>CO</bold><sub><bold>2</bold></sub> <bold>(&#x002B;0.11)</bold> and <bold>Consumption L/100 km (&#x002B;0.09)</bold> contributed incremental increases, reinforcing the emission&#x2013;fuel linkage.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>Speed (&#x002B;0.07)</bold> and <bold>HC (&#x002B;0.06)</bold> provided final boosts, yielding a total of <bold>f(x) &#x003D; 2.025</bold>.</p></list-item>
</list></p></list-item>
<list-item>
<p>Negative contributions (e.g., Force, O<sub>2</sub>, TPS) slightly suppressed the upward trajectory but were insufficient to offset the dominant positive signals.</p></list-item>
</list></td>
</tr>
<tr>
<td><bold>3. Low-impact features</bold></td>
</tr>
<tr>
<td><list list-type="bullet">
<list-item>
<p><bold>MAP (&#x002B;0.03)</bold> and <bold>AFR (&#x002B;0.00)</bold> contributed negligibly.</p></list-item>
<list-item>
<p>Their limited influence suggests that, for this instance, manifold pressure and air&#x2013;fuel ratio readings were close to normal ranges, providing little discriminative information.</p></list-item>
<list-item>
<p>Similarly, <bold>Lambda (&#x2212;0.02)</bold> and <bold>Power (&#x2212;0.02)</bold> showed only marginal effects, implying that deviations in combustion stoichiometry and engine output were minimal compared to other features.</p></list-item>
</list></td>
</tr>
<tr>
<td><bold>4. Decision logic</bold></td>
</tr>
<tr>
<td>The SHAP decomposition reveals that the model&#x2019;s decision hinged on the interplay of <bold>emission indicators (CO, CO</bold><sub><bold>2</bold></sub><bold>, HC)</bold> and <bold>engine operational metrics (RPM, fuel consumption, speed)</bold>. The dominant signal was the <bold>abnormally high CO level</bold>, strongly characteristic of disrupted combustion. Elevated RPM and fuel consumption further reinforced the lean condition, consistent with scenarios where increased air inflow is not matched by proportional fuel injection. Although torque (Force) and O<sub>2</sub> suggested a partial correction toward normal operation, their magnitudes were insufficient to counterbalance the strong emission-driven evidence.</td>
</tr>
<tr>
<td>Thus, the local decision logic can be summarized as:</td>
</tr>
<tr>
<td><italic>Abnormal emissions combined with high RPM and fuel consumption strongly overrode minor stabilizing signals, leading the model to confidently assign the Lean Mixture fault class for this instance</italic>.</td>
</tr>
<tr>
<td><bold>Conclusion</bold></td>
</tr>
<tr>
<td>This SHAP-based local interpretation demonstrates that the model&#x2019;s Lean Mixture classification was primarily driven by abnormal emissions (CO, CO<sub>2</sub>, HC) and elevated operational parameters (RPM, fuel consumption), with only minor counteracting effects from torque and oxygen levels. The analysis confirms that the decision was not arbitrary but instead aligned with domain knowledge of combustion imbalance. This case study highlights the effectiveness of SHAP in elucidating fault-specific reasoning in vehicle engine diagnostics, thereby enhancing the transparency and trustworthiness of machine learning&#x2013;based diagnostic systems.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The explanation highlighted abnormal emission indicators as the dominant factors influencing the classification. In particular, CO emerged as the most decisive feature, reflecting a combustion imbalance characteristic of lean operation. Supporting contributions were observed from elevated engine revolutions, increased fuel consumption per hour, and moderate variations in CO<sub>2</sub> and HC levels, all reinforcing the inference of excess air relative to fuel injection. Vehicle speed was further identified as an operational variable that strengthened the model&#x2019;s confidence in diagnosing a lean mixture under load conditions.</p>
<p>At the same time, the VLM identified counter-signals such as torque and oxygen concentration, which partially mitigated the lean diagnosis but were ultimately insufficient to outweigh the stronger emission- and operation-driven evidence. Variables including MAP, AFR, and Lambda were characterized as having negligible impact, suggesting that their values remained within normal operating ranges for this instance.</p>
<p>Overall, the SHAP-based explanation revealed a stepwise contribution process in which dominant emission-related features progressively elevated the prediction toward the lean mixture class, while weaker counteracting signals exerted only limited influence. The VLM effectively captured this reasoning in narrative form, demonstrating the feasibility of converting SHAP attributions into coherent, human-interpretable explanations. This approach complements the LIME-based analysis by providing a magnitude-oriented perspective, thereby enhancing the transparency and trustworthiness of the diagnostic model. Additional results of the VLM-based SHAP analysis can be found in the Supplementary Materials.</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Large Language Model-Based Report Generator Analysis</title>
<p>To demonstrate the feasibility of integrating multiple interpretability methods, we employed an LLM as a report generator to consolidate VLM-based explanations derived from both LIME and SHAP analyses. While <xref ref-type="sec" rid="s5_2">Sections 5.2</xref> and <xref ref-type="sec" rid="s5_3">5.3</xref> presented instance-level narratives independently generated from LIME and SHAP visualizations, this section highlights how the LLM synthesizes these complementary perspectives into a single unified diagnostic report.</p>
<p>As shown in <xref ref-type="table" rid="table-8">Table 8</xref>, the generated report first confirmed a consistent classification outcome, identifying the instance as a lean mixture condition with high confidence across both interpretability methods. The LLM subsequently aligned overlapping evidence, emphasizing that emission-related features, particularly low carbon monoxide, moderate CO<sub>2</sub>, and slightly elevated HC, served as strong and coherent indicators of lean combustion. In parallel, operational metrics such as high RPM and vehicle speed, combined with the dual fuel consumption pattern, low distance-normalized and high hourly consumption, were identified as mutually reinforcing indicators.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Integration of LIME and SHAP analysis explanation using the LLM.</title>
</caption>
<table>
<colgroup>
<col align="left"/> </colgroup>
<thead>
<tr>
<th>Result</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Unified local diagnostic report&#x2014;lean mixture (VLM-LIME &#x00D7; SHAP)</bold></td>
</tr>
<tr>
<td><bold>(1) Summary of predicted class</bold></td>
</tr>
<tr>
<td>Both explanations attribute the instance to <bold>Lean Mixture</bold> with high confidence (LIME reports class probability 0.84; SHAP shows a strong positive deviation from the baseline toward this class). The local evidence coherently indicates <bold>air-dominant combustion under load</bold>, consistent with a lean condition.</td>
</tr>
<tr>
<td><bold>(2) Cross-method consistencies (strong, overlapping evidence)</bold></td>
</tr>
<tr>
<td><list list-type="bullet">
<list-item>
<p><bold>Combustion by-products</bold>
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>CO (low)</bold>: LIME identifies <bold>CO &#x2264; 0.02</bold> as the single most decisive pro-lean cue. SHAP assigns the <bold>largest positive contribution to CO</bold>. Interpreted together, this means <bold>CO below its typical level increases the Lean probability</bold> (see reconciliation note on directionality in &#x00A7;3).</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>CO</bold><sub><bold>2</bold></sub> <bold>(moderate)</bold> and <bold>HC (slightly elevated)</bold>: Both methods register <bold>moderate CO</bold><sub><bold>2</bold></sub> and <bold>mildly elevated HC</bold> as supporting signals of lean, reflecting relatively complete combustion with pockets of incomplete burn.</p></list-item>
</list></p></list-item>
<list-item>
<p><bold>Fuel&#x2013;efficiency pattern</bold>
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>L/100 km (low)</bold> and <bold>L/H (high)</bold>: LIME highlights <bold>low distance-normalized consumption</bold> with <bold>elevated hourly consumption</bold>; SHAP also gives <bold>L/H</bold> a sizable positive contribution and <bold>L/100 km</bold> a smaller, yet positive one. Together with speed/RPM (below), this pattern is consistent with <bold>sustained cruising/acceleration at load</bold> where fuel per hour rises with airflow/RPM, but <bold>fuel per distance remains low</bold> (efficient burn with air surplus).</p></list-item>
</list></p></list-item>
<list-item>
<p><bold>Operating context</bold>
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>RPM (high)</bold> and <bold>Speed (elevated)</bold>: Both methods show these as <bold>pro-lean</bold> contributors, indicating lean behavior <bold>under dynamic load</bold> rather than idle/steady conditions.</p></list-item>
</list></p></list-item>
<list-item>
<p><bold>Counter-signals (negative contributions)</bold>
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>Force (torque)</bold> and <bold>O</bold><sub><bold>2</bold></sub>: Both appear as <bold>negative pushes</bold> that partially temper the lean conclusion, but <bold>not enough</bold> to overturn the dominant emission-and-efficiency evidence.</p></list-item>
</list></p></list-item>
</list></td>
</tr>
<tr>
<td><bold>(3) Method-specific insights (and reconciliation of differences)</bold></td>
</tr>
<tr>
<td><bold>3.1 LIME (VLM-based narrative on rule-like cues)</bold></td>
</tr>
<tr>
<td><list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>Rule-style thresholds and context coupling</bold>: The explanation surfaces <bold>discrete, human-readable conditions</bold> (e.g., <italic>CO &#x2264; 0.02</italic>, <italic>Speed &#x003E; 0.43</italic>, <italic>0.09 &#x003C; L/100 km &#x2264; 0.19</italic>), spotlighting <bold>complete-combustion signatures</bold> (very low CO) combined with <bold>operational context</bold> (higher speed/RPM) and the <bold>dual fuel-use pattern</bold> (low per-distance, high per-hour).</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>Feature omission implies normality</bold>: MAP, TPS, Lambda, AFR do <bold>not</bold> appear&#x2014;consistent with <bold>near-typical values</bold> that offer little local discrimination for this instance.</p></list-item>
</list></td>
</tr>
<tr>
<td><bold>3.2 SHAP (waterfall decomposition with magnitudes)</bold></td>
</tr>
<tr>
<td><list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>Ranked contributions with additivity</bold>: SHAP quantifies <bold>relative strengths</bold>&#x2014;<bold>CO (&#x002B;0.64)</bold> is dominant, followed by <bold>RPM (&#x002B;0.26)</bold>, <bold>L/H (&#x002B;0.20)</bold>, with <bold>CO</bold><sub><bold>2</bold></sub> <bold>(&#x002B;0.11)</bold>, <bold>L/100 km (&#x002B;0.09)</bold>, <bold>Speed (&#x002B;0.07)</bold>, <bold>HC (&#x002B;0.06)</bold> as supportive. Negatives include <bold>Force (&#x2212;0.06)</bold>, <bold>O</bold><sub><bold>2</bold></sub> <bold>(&#x2212;0.06)</bold>, <bold>TPS (&#x2212;0.04)</bold>, <bold>Power (&#x2212;0.02)</bold>, <bold>Lambda (&#x2212;0.02)</bold>. The <bold>stepwise path</bold> shows a clear <bold>upward trajectory</bold> from the class baseline to a substantially higher final score, consistent with a confident lean assignment.</p></list-item>
<list-item><label>&#x25CB;</label><p><bold>Minor but present variables</bold>: Unlike LIME, SHAP <bold>still registers</bold> small effects for <bold>MAP (&#x002B;0.03)</bold> and <bold>Lambda/Power (&#x2248;&#x2212;0.02)</bold>, indicating <bold>weak but non-zero</bold> local influence rather than pure omission.</p></list-item>
</list></td>
</tr>
<tr>
<td><bold>3.3 Reconciliation of an apparent discrepancy on CO direction</bold></td>
</tr>
<tr>
<td><list list-type="simple">
<list-item>
<p>&#x2022; The SHAP narrative text describes <bold>&#x201C;abnormally high CO&#x201D;</bold> as pro-lean, which conflicts with both <bold>domain knowledge</bold> and <bold>the LIME rule (CO &#x2264; 0.02)</bold>. Given the <bold>large positive SHAP value for CO</bold> and the <bold>LIME threshold indicating very low CO</bold>, the <bold>consistent interpretation</bold> is:
<list list-type="simple">
<list-item><label>&#x25CB;</label><p><bold>This instance&#x2019;s CO is low</bold>, and <bold>the model has learned that lower-than-typical CO increases the Lean probability</bold>.</p></list-item>
<list-item><label>&#x25CB;</label><p>The phrase &#x201C;abnormally high CO&#x201D; in the SHAP narrative is best treated as a <bold>textual slip</bold> from the VLM description; the <bold>attribution sign</bold> (positive toward Lean) aligns with <bold>low CO favoring Lean</bold>.</p></list-item>
</list></p></list-item>
<list-item><label>&#x2212;</label><p>More broadly, <bold>SHAP measures contribution to the Lean score</bold>, not the raw direction of the feature value. LIME&#x2019;s threshold shows the <bold>actual side of the split</bold> (CO very low). Together, they indicate <bold>low CO &#x2192; positive push toward Lean</bold>.</p></list-item>
</list></td>
</tr>
<tr>
<td><bold>3.4 Scale/metric note</bold></td>
</tr>
<tr>
<td><list list-type="simple">
<list-item><label>&#x25CB;</label><p>LIME reports <bold>probability (0.84)</bold>, whereas SHAP&#x2019;s waterfall typically operates in <bold>model output space (often log-odds or margin)</bold>. The <bold>large cumulative positive SHAP shift</bold> complements, rather than contradicts, the <bold>high LIME probability</bold>.</p></list-item>
</list></td>
</tr>
<tr>
<td><bold>(4) Integrated decision logic (synthesized, instance-specific)</bold></td>
</tr>
<tr>
<td><list list-type="order">
<list-item>
<p><bold>Primary signal (combustion quality): Very low CO</bold>&#x2014;supported by <bold>moderate CO</bold><sub><bold>2</bold></sub> and <bold>slightly elevated HC</bold>&#x2014;indicates <bold>air-dominant, relatively complete combustion</bold> with some local lean-zone misfires.</p></list-item>
<list-item>
<p><bold>Operational context under load: High RPM</bold> and <bold>elevated speed</bold> point to <bold>lean behavior during acceleration/cruise</bold>, not idle.</p></list-item>
<list-item>
<p><bold>Fuel-use pattern consistent with lean: Low L/100 km</bold> alongside <bold>high L/H</bold> reflects <bold>efficient per-distance burn</bold> (air surplus) amid <bold>high throughput per time</bold> (due to RPM/airflow).</p></list-item>
<list-item>
<p><bold>Counter-evidence insufficient: Moderate torque (Force)</bold> and <bold>O</bold><sub><bold>2</bold></sub> <bold>in a mid range</bold> slightly oppose a lean label but are <bold>small in magnitude</bold> relative to the dominant emission and efficiency signals.</p></list-item>
<list-item>
<p><bold>Other sensors near nominal: MAP, TPS, Lambda, AFR</bold> show <bold>minimal local impact</bold>, indicating values close to typical ranges and adding little discriminative power here.</p></list-item>
</list></td>
</tr>
<tr>
<td><bold>Conclusion (local, for this instance):</bold></td>
</tr>
<tr>
<td>The convergent evidence from LIME and SHAP supports a <bold>Lean Mixture</bold> diagnosis driven chiefly by <bold>very low CO</bold> (complete-combustion signature), <bold>moderate CO</bold><sub><bold>2</bold></sub> and <bold>slightly elevated HC</bold>, combined with <bold>high RPM/speed</bold> and a <bold>fuel-use profile</bold> of <bold>low per-distance but high per-hour consumption</bold>. Minor counter-signals (Force, O<sub>2</sub>) do not overcome the dominant pattern. The integrated interpretation is <bold>technically consistent</bold>, <bold>diagnostically coherent</bold>, and <bold>faithful to the model&#x2019;s local decision path</bold> for this specific sample.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The LLM also reconciled method-specific differences between LIME and SHAP. LIME emphasized interpretable threshold-based rules, such as ranges of CO and fuel consumption, while SHAP quantified additive feature contributions in a stepwise manner. In cases of apparent discrepancies, such as the interpretation of CO directionality, the LLM clarified that both methods ultimately supported the conclusion that lower than typical CO values increased the probability of lean operation. Moreover, the report distinguished between features omitted in LIME due to near-normal values and those assigned minimal weights by SHAP, thereby ensuring consistent representation of low impact evidence across methods.</p>
<p>Through this synthesis, the LLM generated an integrated decision logic that contextualized the prediction as the result of convergent signals from emissions, operational load, and fuel efficiency, moderated by weak counter-signals such as torque and oxygen levels. The resulting explanation formed a coherent narrative that remained diagnostically faithful to the model&#x2019;s reasoning process while enhancing accessibility for human interpretation.</p>
<p>This experiment illustrates the potential of LLM-based report generators to serve as interpreters that unify heterogeneous explanation modalities. By combining the threshold-oriented clarity of LIME with the magnitude-based decomposition of SHAP, the method enhances transparency, mitigates the risk of conflicting interpretations, and provides domain experts with a more robust and interpretable account of ML&#x2013;driven fault diagnosis.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Evaluation of Explanatory Quality and User Satisfaction</title>
<p>To assess the explanatory quality and user satisfaction of the proposed method, we employed a dual evaluation method that combined VLM-as-a-Judge-based assessment with a user study incorporating quantitative ratings and qualitative feedback. The evaluation employed a five-point Likert scale survey administered to 10 domain experts with professional experience in vehicle engine systems. The study compared three explanation settings: VLM-based LIME analysis, VLM-based SHAP analysis, and the proposed method. The evaluation aimed to assess five key criteria, including comprehensibility, faithfulness, domain relevance, explanatory value, and reliability. In addition, open-ended responses were collected to identify strengths and areas for improvement.</p>
<p>As shown in <xref ref-type="table" rid="table-9">Tables 9</xref> and <xref ref-type="table" rid="table-10">10</xref>, the VLM-as-a-Judge evaluation results indicate that the proposed method achieved the highest scores across all evaluation criteria under both GPT-4o and LLaMA 3.2-Vision. In particular, improvements in comprehensibility and faithfulness highlighted the effectiveness of synthesizing LIME- and SHAP-based explanations into a unified narrative that enhances clarity and alignment with the model&#x2019;s diagnostic reasoning. Although LLaMA 3.2-Vision produced slightly more conservative absolute scores than GPT-4o, the relative performance trends remained consistent. In both evaluations, the proposed method outperformed standalone LIME and SHAP analyses, especially in domain relevance and explanatory value, suggesting that the observed gains are robust to evaluator-specific bias. Reliability scores remained lower than other criteria, reflecting the inherent challenge of establishing trust in AI-driven diagnostics. Nevertheless, the proposed method demonstrated consistent improvements in reliability over individual explanation approaches. Overall, the agreement between two independent VLM evaluators provided strong evidence for the robustness and generalizability of the proposed method.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>VLM-as-a-Judge evaluation results using GPT-4o based on a 5-point Likert scal.e</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Criterion</th>
<th>VLM-Based LIME Analysis (Mean &#x00B1; SD)</th>
<th>VLM-Based SHAP Analysis (Mean &#x00B1; SD)</th>
<th>Proposed Method (Mean &#x00B1; SD)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Comprehensibility</td>
<td>4.53 &#x00B1; 0.57</td>
<td>4.63 &#x00B1; 0.49</td>
<td>4.70 &#x00B1; 0.46</td>
</tr>
<tr>
<td>Faithfulness</td>
<td>4.66 &#x00B1; 0.47</td>
<td>4.56 &#x00B1; 0.50</td>
<td>4.73 &#x00B1; 0.44</td>
</tr>
<tr>
<td>Domain relevance</td>
<td>4.03 &#x00B1; 0.61</td>
<td>4.20 &#x00B1; 0.66</td>
<td>4.33 &#x00B1; 0.54</td>
</tr>
<tr>
<td>Explanatory value</td>
<td>4.60 &#x00B1; 0.49</td>
<td>4.66 &#x00B1; 0.47</td>
<td>4.63 &#x00B1; 0.55</td>
</tr>
<tr>
<td>Reliability</td>
<td>3.96 &#x00B1; 0.66</td>
<td>4.13 &#x00B1; 0.73</td>
<td>4.20 &#x00B1; 0.48</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>VLM-as-a-Judge evaluation results using LLaMA 3.2-Vision based on a 5-point Likert scale.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Criterion</th>
<th>VLM-Based Lime Analysis (Mean &#x00B1; SD)</th>
<th>VLM-Based Shap Analysis (Mean &#x00B1; SD)</th>
<th>Proposed Method (Mean &#x00B1; SD)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Comprehensibility</td>
<td>4.34 &#x00B1; 0.57</td>
<td>4.41 &#x00B1; 0.55</td>
<td>4.46 &#x00B1; 0.52</td>
</tr>
<tr>
<td>Faithfulness</td>
<td>4.28 &#x00B1; 0.62</td>
<td>4.31 &#x00B1; 0.57</td>
<td>4.47 &#x00B1; 0.59</td>
</tr>
<tr>
<td>Domain relevance</td>
<td>3.82 &#x00B1; 0.68</td>
<td>3.96 &#x00B1; 0.70</td>
<td>4.12 &#x00B1; 0.61</td>
</tr>
<tr>
<td>Explanatory value</td>
<td>4.39 &#x00B1; 0.55</td>
<td>4.35 &#x00B1; 0.58</td>
<td>4.48 &#x00B1; 0.50</td>
</tr>
<tr>
<td>Reliability</td>
<td>3.71 &#x00B1; 0.72</td>
<td>3.86 &#x00B1; 0.75</td>
<td>3.98 &#x00B1; 0.60</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-11">Table 11</xref> presents the results of the user study. Although the absolute scores were lower than those obtained in the VLM-as-a-Judge evaluation, the relative trends remained consistent. The proposed method again outperformed both standalone LIME and SHAP analyses across all evaluation criteria. Participants particularly appreciated improvements in faithfulness and explanatory value, underscoring the effectiveness of the unified report in conveying diagnostic reasoning. Nevertheless, the comparatively lower scores for reliability across all methods highlight the persistent challenge of fostering user confidence in automated explanations within safety-critical diagnostic environments.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Evaluation of proposed method through the user study.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Criterion</th>
<th>VLM-Based LIME Analysis (Mean &#x00B1; SD)</th>
<th>VLM-Based SHAP Analysis (Mean &#x00B1; SD)</th>
<th>Proposed method (Mean &#x00B1; SD)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Comprehensibility</td>
<td>3.76 &#x00B1; 0.81</td>
<td>4.06 &#x00B1; 0.78</td>
<td>4.13 &#x00B1; 0.57</td>
</tr>
<tr>
<td>Faithfulness</td>
<td>3.83 &#x00B1; 0.87</td>
<td>4.10 &#x00B1; 0.66</td>
<td>4.20 &#x00B1; 0.66</td>
</tr>
<tr>
<td>Domain relevance</td>
<td>3.46 &#x00B1; 0.86</td>
<td>3.56 &#x00B1; 0.85</td>
<td>4.03 &#x00B1; 0.71</td>
</tr>
<tr>
<td>Explanatory value</td>
<td>3.80 &#x00B1; 0.84</td>
<td>4.16 &#x00B1; 0.79</td>
<td>4.23 &#x00B1; 0.56</td>
</tr>
<tr>
<td>Reliability</td>
<td>3.56 &#x00B1; 0.89</td>
<td>3.63 &#x00B1; 0.76</td>
<td>3.66 &#x00B1; 0.47</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-12">Table 12</xref> summarizes the open-ended feedback. Positive remarks highlighted the accessibility of the unified report, its clear and structured format, and the sense of trust fostered by presenting consistent evidence across methods. Participants noted that the integration of visual indicators and textual reasoning enhanced decision-making and communication in practical diagnostic contexts. On the other hand, suggestions for improvement emphasized the need for more domain-specific diagnostic guidance (e.g., tailored recommendations for sensor replacement or fuel system checks), improved handling of discrepancies between LIME and SHAP outputs, and options to adjust report length and complexity for different user groups. Participants also recommended emphasizing threshold deviations more explicitly to enhance interpretability for less experienced users.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Opinions of users on the proposed method.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Type</th>
<th>Opinion</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="5">Positive feedback</td>
<td>It is much easier to understand since I can see LIME and SHAP results in a single consolidated report.</td>
</tr>
<tr>

<td>The structured format with clear sections improves readability and understanding.</td>
</tr>
<tr>

<td>The unified report highlights consistencies across methods, which increases trust in the explanation.</td>
</tr>
<tr>

<td>The integration of visualization-derived evidence into textual reasoning is helpful for decision-making.</td>
</tr>
<tr>

<td>The report style makes it directly applicable for documentation and communication with colleagues.</td>
</tr>
<tr>
<td rowspan="5">Improvement suggestion</td>
<td>The report could be improved by offering domain-specific diagnostic guidance, such as tailored recommendations for sensor replacement or fuel system checks.</td>
</tr>
<tr>

<td>Rather than focusing solely on generic fault classes (Lean/Rich/Voltage), adapting the explanations to specific vehicle or engine types would enhance their applicability.</td>
</tr>
<tr>

<td>A deeper analysis of why LIME and SHAP sometimes disagree would improve reliability.</td>
</tr>
<tr>

<td>More emphasis on relative deviations from thresholds would make the explanations more informative.</td>
</tr>
<tr>

<td>Offering a concise executive summary or adjustable report length would improve usability.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Taken together, these findings confirm that the proposed method delivers superior explanatory performance by enhancing clarity, consistency, and user relevance compared to single-method explanations. At the same time, the qualitative feedback highlights directions for refinement, particularly in domain customization, flexible presentation, and reliability enhancement. Addressing these areas is expected to further improve the usability and practical adoption of the proposed method in real-world diagnostic scenarios.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Discussion</title>
<sec id="s6_1">
<label>6.1</label>
<title>Consistent and Discrepancy Analysis across XAI Methods</title>
<p>Although LIME and SHAP are jointly integrated into a unified diagnostic report, differences in their attribution mechanisms may lead to variations in feature importance and ranking [<xref ref-type="bibr" rid="ref-65">65</xref>]. To move beyond purely descriptive reconciliation and to enhance cross-XAI reliability, we conducted a quantitative consistency and discrepancy analysis between LIME- and SHAP-based local explanations using an LLM.</p>
<p>As summarized in <xref ref-type="table" rid="table-18">Table A5</xref>, a structured analysis prompt was designed to objectively compare feature-level attributions produced by LIME and SHAP for the same diagnostic instance. Overlapping features were first identified, and their importance values were independently normalized to account for scale differences between LIME coefficients and SHAP values. Based on these normalized representations, four quantitative metrics were computed: feature overlap ratio [<xref ref-type="bibr" rid="ref-66">66</xref>], Spearman rank correlation [<xref ref-type="bibr" rid="ref-67">67</xref>], directional consistency score [<xref ref-type="bibr" rid="ref-68">68</xref>], and normalized importance divergence [<xref ref-type="bibr" rid="ref-69">69</xref>].</p>

<p><xref ref-type="table" rid="table-13">Table 13</xref> presents the analysis results for the same lean mixture instance previously analyzed in <xref ref-type="table" rid="table-8">Table 8</xref>. Strong structural agreement was observed, with nine overlapping features out of eleven unique features, yielding a feature overlap ratio of 0.818. Both methods consistently assigned the highest relative importance to CO, confirming agreement on the dominant role of emission-related indicators. Other shared features, including fuel consumption metrics, vehicle speed, RPM, and exhaust gas variables, further reinforced this common explanatory basis.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Result of consistent and discrepancy analysis across XAI methods.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th colspan="4">Result</th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="4"><bold>1. Overlapping feature set</bold></td>
</tr>
<tr>
<td colspan="4"><bold>LIME features (10):</bold></td>
</tr>
<tr>
<td colspan="4">CO, L/100 km, Speed, L/H, HC, CO2, Power, RPM, Force, O2</td>
</tr>
<tr>
<td colspan="4"><bold>SHAP features (10):</bold></td>
</tr>
<tr>
<td colspan="4">CO, RPM, L/H, CO2, L/100 km, Speed, HC, Force, O2, TPS</td>
</tr>
<tr>
<td colspan="4"><bold>Overlapping features (9):</bold></td>
</tr>
<tr>
<td colspan="4"><bold>CO, L/100 km, Speed, L/H, HC, CO2, RPM, Force, O2</bold></td>
</tr>
<tr>
<td colspan="4"><bold>Union of features (11):</bold></td>
</tr>
<tr>
<td colspan="4">{CO, L/100 km, Speed, L/H, HC, CO2, Power, RPM, Force, O2, TPS}</td>
</tr>
<tr>
<td colspan="4"><bold>2. Normalization of feature importances</bold></td>
</tr>
<tr>
<td colspan="4">Feature importances are <bold>min&#x2013;max normalized independently</bold> within each method over the overlapping feature set.</td>
</tr>
<tr>
<td colspan="4"><bold>Normalized importances (overlapping features only)</bold></td>
</tr>
<tr>
<td><bold>Feature</bold></td>
<td><bold>LIME (norm)</bold></td>
<td><bold>SHAP (norm)</bold></td>
<td><bold>Direction (LIME/SHAP)</bold></td>
</tr>
<tr>
<td>CO</td>
<td>1.000</td>
<td>1.000</td>
<td>&#x002B;/&#x002B;</td>
</tr>
<tr>
<td>L/100 km</td>
<td>0.508</td>
<td>0.086</td>
<td>&#x002B;/&#x002B;</td>
</tr>
<tr>
<td>Speed</td>
<td>0.439</td>
<td>0.018</td>
<td>&#x002B;/&#x002B;</td>
</tr>
<tr>
<td>L/H</td>
<td>0.269</td>
<td>0.241</td>
<td>&#x002B;/&#x002B;</td>
</tr>
<tr>
<td>HC</td>
<td>0.130</td>
<td>0.000</td>
<td>&#x002B;/&#x002B;</td>
</tr>
<tr>
<td>CO2</td>
<td>0.042</td>
<td>0.121</td>
<td>&#x002B;/&#x002B;</td>
</tr>
<tr>
<td>RPM</td>
<td>0.000</td>
<td>0.345</td>
<td>&#x002B;/&#x002B;</td>
</tr>
<tr>
<td>Force</td>
<td>0.394</td>
<td>0.000</td>
<td>&#x2212;/&#x2212;</td>
</tr>
<tr>
<td>O2</td>
<td>0.163</td>
<td>0.000</td>
<td>&#x2212;/&#x2212;</td>
</tr>
<tr>
<td colspan="4"><bold>3. Consistency metrics</bold></td>
</tr>
<tr>
<td colspan="4"><bold>a) Feature Overlap Ratio (FOR)</bold></td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mtext>FOR</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">&#x2223;</mml:mo><mml:mrow><mml:mtext>Overlap</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2223;</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">&#x2223;</mml:mo><mml:mrow><mml:mtext>Union</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2223;</mml:mo></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mrow><mml:mn>9</mml:mn></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">0.818</mml:mtext></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td colspan="4"><bold>b) Rank consistency (Spearman&#x2019;s &#x03C1;)</bold></td>
</tr>
<tr>
<td colspan="4">Computed over the 9 overlapping features using their respective ranks:</td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>&#x03C1;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">0.48</mml:mtext></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td colspan="4">&#x2192; <bold>Moderate rank agreement</bold></td>
</tr>
<tr>
<td colspan="4"><bold>c) Directional Consistency Score (DCS)</bold></td>
</tr>
<tr>
<td colspan="4">All overlapping features have <bold>identical contribution directions</bold>.</td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;<inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mtext>DCS</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mn>9</mml:mn></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mrow><mml:mn>9</mml:mn></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">1.00</mml:mtext></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td colspan="4">&#x2192; <bold>Perfect directional alignment</bold></td>
</tr>
<tr>
<td colspan="4"><bold>d) Normalized Importance Divergence (NID)</bold></td>
</tr>
<tr>
<td colspan="4">Mean absolute difference between normalized importances:</td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;<inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mrow><mml:mtext>NID</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">0.236</mml:mtext></mml:mrow></mml:math></inline-formula>&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;&#x2002;</td>
</tr>
<tr>
<td colspan="4"><bold>4. Metric summary</bold></td>
</tr>
<tr>
<td colspan="4">{</td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x201C;Feature_Overlap_Ratio&#x201D;: 0.818,</td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x201C;Spearman_Rank_Correlation&#x201D;: 0.48,</td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x201C;Directional_Consistency_Score&#x201D;: 1.00,</td>
</tr>
<tr>
<td colspan="4">&#x2002;&#x2002;&#x201C;Normalized_Importance_Divergence&#x201D;: 0.236</td>
</tr>
<tr>
<td colspan="4">}</td>
</tr>
<tr>
<td colspan="4"><bold>5. Interpretation and consistency assessment</bold></td>
</tr>
<tr>
<td colspan="4">
<list list-type="bullet">
<list-item>
<p><bold>High structural consistency</bold> is observed between LIME and SHAP, as evidenced by a <bold>high Feature Overlap Ratio (0.818)</bold> and <bold>perfect Directional Consistency (1.00)</bold>. Both methods unanimously agree on <italic>which features push the prediction toward or away from the Lean Mixture class</italic>.</p></list-item>
<list-item>
<p><bold>Moderate rank consistency (&#x03C1; &#x003D; 0.48)</bold> indicates that while the two methods identify largely the same influential variables, they <bold>prioritize them differently</bold>. For example:
<list list-type="simple">
<list-item><label>&#x25CB;</label>
<p><bold>RPM</bold> is highly ranked in SHAP (rank 2) but appears with minimal importance in LIME.</p></list-item>
<list-item><label>&#x25CB;</label>
<p><bold>L/100 km and Speed</bold> receive stronger emphasis in LIME than in SHAP.</p></list-item>
</list></p></list-item>
<list-item>
<p>The <bold>Normalized Importance Divergence (0.236)</bold> suggests a <bold>non-negligible difference in attribution magnitude</bold>, reflecting methodological distinctions:
<list list-type="simple">
<list-item><label>&#x25CB;</label>
<p>LIME emphasizes <italic>local linear sensitivity</italic> around the instance.</p></list-item>
<list-item><label>&#x25CB;</label>
<p>SHAP reflects <italic>game-theoretic marginal contributions</italic> relative to a baseline.</p></list-item>
</list></p></list-item>
<list-item>
<p>Importantly, <bold>no directional conflicts</bold> are present. All discrepancies are confined to <bold>relative importance scaling and ranking</bold>, not interpretive conclusions.</p></list-item>
</list>
</td>
</tr>
<tr>
<td colspan="4"><bold>Overall conclusion</bold></td>
</tr>
<tr>
<td colspan="4">The two explanation methods demonstrate <bold>strong qualitative agreement and moderate quantitative alignment</bold>. They should be interpreted as <bold>complementary local explanation tools</bold>, with SHAP providing a more globally anchored attribution structure and LIME offering sharper local sensitivity, rather than as competing or contradictory explanations.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Directional consistency analysis revealed perfect semantic alignment, with a directional consistency score of 1.00, indicating that all overlapping features contributed in the same direction to the predicted class. In contrast, rank- and magnitude-based metrics showed moderate divergence, with a Spearman rank correlation of 0.48 and a normalized importance divergence of 0.236. These differences reflect methodological distinctions, LIME&#x2019;s local sensitivity vs. SHAP&#x2019;s additive marginal contributions, rather than contradictory diagnostic reasoning.</p>
<p>Overall, the results demonstrate strong qualitative consistency and moderate quantitative alignment between LIME and SHAP. The two methods should therefore be interpreted as complementary, and explicitly quantifying their agreement and divergence strengthens the reliability of the proposed method.</p>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Robustness against Prompt Sensitivity and Hallucination</title>
<p>Recent studies have shown that both VLMs and LLMs are susceptible to prompt sensitivity and may occasionally produce misleading or hallucinated outputs [<xref ref-type="bibr" rid="ref-70">70</xref>,<xref ref-type="bibr" rid="ref-71">71</xref>]. To address this concern, we conducted an explicit hallucination analysis on the generated explanations and introduced an additional verification mechanism to impose domain-specific constraints on the explanation process.</p>
<p>First, we examined all VLM and LLM-generated explanations used in the experiments to identify potential hallucinations, defined as statements that are inconsistent with the underlying XAI visualizations, model outputs, or domain knowledge of vehicle engine fault diagnosis. This examination was conducted through manual and qualitative inspection of the evaluated instances, in which the generalized textual explanations were systematically compared against the corresponding LIME or SHAP visualizations and diagnostic evidence.</p>
<p>Within the scope of the instances analyzed in this study, no hallucinated explanations were observed. In particular, all referenced features, contribution directions, and diagnostic interpretations were grounded in the corresponding XAI results and aligned with well-known combustion and engine operation principles. These observations indicate that the structured prompt design adopted in this study effectively constrained the generative behavior of the models under the evaluated conditions.</p>
<p>Nevertheless, to further mitigate the risk of erroneous explanations, we introduce an additional verification stage based on a domain-aware review agent [<xref ref-type="bibr" rid="ref-72">72</xref>]. The prompt of this agent is presented in <xref ref-type="table" rid="table-19">Table A6</xref>. Specifically, explanations generated by the VLMs for LIME and SHAP interpretation and the unified reports produced by the LLMs are subsequently re-evaluated by an independent verification agent. This agent assesses factual consistency with the original XAI outputs, checks adherence to domain-specific constraints, and flags unsupported claims or speculative reasoning.</p>

<p>By incorporating this verification stage, the proposed method enforces an explicit separation between explanation generation and explanation validation. This design reduces reliance on a single generative pass and enhances robustness against prompt sensitivity and latent hallucination risks. Although the present study did not observe hallucinated outputs, the verification mechanism provides an additional safeguard and establishes a scalable foundation for deploying the proposed method in safety-critical and energy-related diagnostic systems.</p>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Implications for Automotive Maintenance Workflows</title>
<p>While the proposed DRIVE focuses on fault classification and explainable report generation, its outputs are directly applicable to real-world automotive maintenance and troubleshooting workflows [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>]. To illustrate this practical relevance, we discuss how the unified diagnostic report supports technician decision-making using a representative lean mixture case.</p>
<p>As shown in <xref ref-type="table" rid="table-8">Table 8</xref>, the unified report consolidates LIME and SHAP-based explanations into a structured and instance-level diagnostic narrative that highlights dominant signals, operating context, and counter evidence. From a maintenance perspective, this report bridges model inference and component-level inspection. For example, the consistent identification of low CO levels combined with moderate CO<sub>2</sub> and elevated HC emissions directs technicians toward air-dominant combustion issues, guiding inspection priorities such as intake air leaks, airflow sensor calibration, or fuel delivery constraints.</p>

<p>The report further contextualizes the fault under dynamic operating conditions by linking elevated RPM and vehicle speed with a dual fuel-consumption pattern. This helps technicians distinguish transient driving effects from persistent system-level faults, thereby narrowing troubleshooting pathways and reducing unnecessary component replacement. Explicit identification of counter signals, such as torque and oxygen levels, also supports ruling out competing fault hypotheses.</p>
<p>Overall, the stepwise and structured nature of the unified report aligns well with automotive diagnostic workflows, reducing the cognitive burden associated with interpreting raw XAI visualizations. Although an end-to-end deployment is beyond the scope of this study, the presented case demonstrates that DRIVE can function as an effective decision-support layer in practical automotive maintenance pipelines.</p>
</sec>
<sec id="s6_4">
<label>6.4</label>
<title>Computational Considerations and Deployment Feasibility</title>
<p>Real-world vehicle diagnostics often impose constraints on inference latency and computational resources, particularly in scenarios requiring near-real-time decision support [<xref ref-type="bibr" rid="ref-73">73</xref>,<xref ref-type="bibr" rid="ref-74">74</xref>]. It should be noted that the primary objective of DRIVE is not real-time engine control, but interpretable diagnostic reasoning and decision support for maintenance, troubleshooting, and post-event analysis.</p>
<p>From a computational standpoint, DRIVE follows a modular architecture with distinct latency characteristics. The core fault classification is performed using XGBoost, whose inference time is on the order of milliseconds under embedded hardware assumptions and is well suited for deployment on edge devices or in-vehicle diagnostic units. This enables timely detection of potential fault states without reliance on external computation.</p>
<p>In contrast, the explainability and report generation stages, including LIME and SHAP analysis, VLM-based interpretation, and LLM-based synthesis, incur higher computational overhead. These components are intended for near-real-time or offline diagnostic contexts, such as fleet maintenance systems and engineering analysis platforms, where response times of several seconds are acceptable and interpretability is prioritized over strict latency constraints.</p>
<p>The modular design of DRIVE allows flexible deployment strategies. For example, a hierarchical configuration can be adopted in which fault classification runs on the vehicle or edge device, while XAI computation and VLM&#x2013;LLM reasoning are executed on a gateway, on-premise server, or cloud infrastructure [<xref ref-type="bibr" rid="ref-75">75</xref>,<xref ref-type="bibr" rid="ref-76">76</xref>]. Future work will focus on latency-aware optimization, lightweight explanation generation, and partial on-device deployment to further improve engineering feasibility in energy and safety-critical automotive systems [<xref ref-type="bibr" rid="ref-77">77</xref>,<xref ref-type="bibr" rid="ref-78">78</xref>].</p>
</sec>
<sec id="s6_5">
<label>6.5</label>
<title>Extension to Additional XAI Methods</title>
<p>The experimental evaluation in this study focused on LIME and SHAP as representative instance-level XAI techniques, as the primary objective of DRIVE is to enhance local interpretability by translating instance-specific visual explanations into coherent textual reasoning and unified diagnostic reports. Nevertheless, other widely used XAI methods, such as partial dependence plots (PDP) [<xref ref-type="bibr" rid="ref-19">19</xref>] and individual conditional expectation (ICE) [<xref ref-type="bibr" rid="ref-79">79</xref>] plots, offer complementary explanatory perspectives.</p>
<p>PDP and ICE are effective in capturing global and semi-local relationships between input variables and model predictions. PDP illustrates the average marginal effect of a feature across the dataset, while ICE reveals instance-level heterogeneity in model responses. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> presents representative PDP and ICE visualizations for key engine-related variables, demonstrating how changes in sensor values influence predicted fault probabilities.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Visualization of PDP and ICE analysis: (<bold>a</bold>) PDP analysis; (<bold>b</bold>) ICE analysis.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_76888-fig-6.tif"/>
</fig>
<p>Although PDP and ICE were not integrated into DRIVE, these results indicate that the proposed method is not limited to LIME and SHAP. In principle, PDP and ICE visualizations can be interpreted using the same VLM-based translation mechanism and subsequently synthesized by the LLM together with LIME- and SHAP-based explanations. Incorporating additional XAI techniques, including PDP, ICE, and counterfactual explanations, therefore represents a natural extension of DRIVE and will be explored in future work to further enhance interpretability and practical decision support in vehicle engine fault diagnosis and energy-related systems [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
</sec>
<sec id="s6_6">
<label>6.6</label>
<title>Generalizability across Vision-Language Models</title>
<p>In the current implementation, GPT-4o was employed as the VLM to interpret LIME and SHAP visualizations and generate explanatory reports, owing to its strong multimodal reasoning capability. Nevertheless, reliance on a single VLM raises questions regarding the generalizability of the proposed method when alternative models are used.</p>
<p>Importantly, DRIVE is designed to be VLM-agnostic. The VLM is strictly constrained to translating structured XAI visualizations into textual explanations under predefined prompts and domain-specific guidelines. Consequently, the core diagnostic evidence derived from the ML model and XAI methods remains unchanged regardless of the selected VLM. Replacing GPT-4o with other VLMs, such as LLaMA 3.2-Vision [<xref ref-type="bibr" rid="ref-63">63</xref>], is therefore expected to mainly affect linguistic style and conservativeness, rather than the underlying explanatory content or diagnostic conclusions [<xref ref-type="bibr" rid="ref-80">80</xref>].</p>
<p>Auxiliary evaluations with alternative VLMs indicate that, while absolute Likert-scale scores may vary slightly, the relative performance trends and qualitative interpretations remain consistent under identical prompt structures. This suggests that the explanatory effectiveness of DRIVE is primarily driven by structured XAI integration and prompt design rather than by VLM-specific characteristics. Future work will include systematic benchmarking of multiple VLMs to further assess robustness, faithfulness, and hallucination resistance in safety-critical diagnostic settings [<xref ref-type="bibr" rid="ref-81">81</xref>].</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Conclusion</title>
<p>In this study, we proposed a novel method, DRIVE, that enhances the interpretability and practical usability of vehicle engine fault diagnosis by integrating ML models with VLMs and LLMs. After selecting the best-performing ML classifier for engine fault data, we employed VLMs to transform the outputs of two widely used XAI methods, LIME and SHAP, into structured textual explanations. These complementary explanations were subsequently synthesized by an LLM into a unified diagnostic report, providing users with a coherent and human-readable narrative that clarifies the decision-making process of the ML model.</p>
<p>Experimental results demonstrated that the proposed method not only maintains high diagnostic accuracy but also substantially improves the transparency and accessibility of model reasoning. For example, LIME-based explanations highlighted local feature perturbations influencing specific predictions, while SHAP-based explanations quantified the cumulative contribution of sensor variables. By integrating these perspectives, the LLM-generated reports provided a comprehensive interpretation of the diagnostic outcome and facilitated a clearer understanding for domain experts.</p>
<p>The findings of this study highlight the potential of combining VLM-based interpretability with LLM-based synthesis to advance explainability in AI-driven fault diagnosis. Specifically, the proposed method demonstrates how LIME and SHAP outputs, which traditionally require expert interpretation of visual plots, can be automatically converted into coherent textual narratives and further integrated into unified diagnostic reports. This multi-stage process not only reduces the cognitive burden on end-users but also provides a more comprehensive and reliable account of the model&#x2019;s reasoning. By enabling clear communication of diagnostic logic and aligning model outputs with domain expertise, this method contributes to the digital transformation of vehicle maintenance through more transparent, trustworthy, and user-oriented diagnostic systems that support informed decision-making in both engineering and operational contexts.</p>
<p>While LIME and SHAP analysis were utilized as the primary XAI method in this study, other post-hoc method such as PDP [<xref ref-type="bibr" rid="ref-19">19</xref>] ICE plots [<xref ref-type="bibr" rid="ref-79">79</xref>] provide additional valuable perspectives on model interpretability. Future research will aim to integrate multiple XAI methods with VLMs and LLMs to further enhance the interpretability of ML models. In addition, considering the prompt sensitivity of the outputs of VLMs and LLMs, future studies will investigate the development of adaptive prompt templates customized to individual input instances, corresponding visual outputs, and model confidence levels, thereby improving the precision, relevance, and stability of generated explanations. Such an approach would enable the explanation tone and certainty to be moderated in low-confidence predictions, reducing the risk of overly deterministic interpretations. By explicitly reflecting model confidence in the explanation generation process, the resulting reports can better communicate diagnostic uncertainty and support cautious, expert-driven decision making rather than definitive fault assertions.</p>
<p>From a broader perspective, future work will explore the extension of the proposed method toward energy systems, where explainable vehicle-level diagnostic can contribute to improving fuel efficiency, reducing unnecessary energy consumption, and supporting emission-aware energy management [<xref ref-type="bibr" rid="ref-82">82</xref>]. Such integration would enable the proposed method to serve as a building block for explainable deep learning applications in transportation-related energy systems.</p>
</sec>
<sec sec-type="supplementary-material" id="s8">
<title>Supplementary Materials</title>
<supplementary-material id="SD1">
<media xlink:href="CMES_76888-s001.docx"/>
</supplementary-material></sec>
</body>
<back>
<ack>
<p>This research was supported by Korea Institute of Planning and Evaluation for Technology in Food, Agriculture and Forestry (IPET) through the High Value-added Food Technology Development Program, funded by the Ministry of Agriculture, Food and Rural Affairs (MAFRA) (RS-2024-00403286).</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2025-00516023).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Jaeseung Lee and Jehyeok Rew; methodology, Jehyeok Rew; software, Jaeseung Lee; validation, Jehyeok Rew; formal analysis, Jehyeok Rew; investigation, Jaeseung Lee; resources, Jehyeok Rew; data curation, Jaeseung Lee; writing&#x2014;original draft preparation, Jaeseung Lee; writing&#x2014;review and editing, Jehyeok Rew; visualization, Jaeseung Lee; supervision, Jehyeok Rew; project administration, Jehyeok Rew; funding acquisition, Jehyeok Rew. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support of the findings of this study are openly available in EngineFaultDB at <ext-link ext-link-type="uri" xlink:href="https://github.com/Leo-Thomas/EngineFaultDB">https://github.com/Leo-Thomas/EngineFaultDB</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<sec>
<title>Supplementary Materials</title>
<p>The supplementary material is available online at <ext-link ext-link-type="uri" xlink:href="https://www.techscience.com/doi/10.32604/journal.2026.076888/s1">https://www.techscience.com/doi/10.32604/journal.2026.076888/s1</ext-link>. The following supporting information is provided at Supplementary Materials. Figure S1: LIME plot of proposed method&#x2014;classified as No Fault; Figure S2: SHAP plot of proposed method&#x2014;classified as No Fault; Figure S3: LIME plot of proposed method&#x2014;classified as Rich Mixture; Figure S4: SHAP plot of proposed method&#x2014;classified as Rich Mixture; Figure S5: LIME plot of proposed method&#x2014;classified as Low Voltage; Figure S6: SHAP plot of proposed method&#x2014;classified as Low Voltage; Table S1: Interpretation of LIME analysis using the VLM&#x2014;corresponding to Fig. S1; Table S2: Interpretation of SHAP analysis using the VLM&#x2014;corresponding to Fig. S2; Table S3: Interpretation of LIME and SHAP analysis using the LLM&#x2014;integrating Tables S1 and S2; Table S4: Interpretation of LIME analysis using the VLM&#x2014;corresponding to Fig. S3; Table S5: Interpretation of SHAP analysis using the VLM&#x2014;corresponding to Fig. S4; Table S6: Interpretation of LIME and SHAP analysis using the LLM&#x2014;integrating Tables S4 and S5; Table S7: Interpretation of LIME analysis using the VLM&#x2014;corresponding to Fig. S5; Table S8: Interpretation of SHAP analysis using the VLM&#x2014;corresponding to Fig. S6; Table S9: Interpretation of LIME and SHAP analysis using the LLM&#x2014;integrating Tables S7 and S8.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1"> 
<title>Abbreviations</title>
<def-list>
<def-item>
<term>AHP</term>
<def>
<p>Analytic hierarchy process</p>
</def>
</def-item>
<def-item>
<term>AFR</term>
<def>
<p>Air-fuel ratio</p>
</def>
</def-item>
<def-item>
<term>CO</term>
<def>
<p>Carbon monoxide</p>
</def>
</def-item>
<def-item>
<term>HC</term>
<def>
<p>Hydrocarbons</p>
</def>
</def-item>
<def-item>
<term>ICE</term>
<def>
<p>Individual conditional expectation</p>
</def>
</def-item>
<def-item>
<term>LIME</term>
<def>
<p>Local interpretable model-agnostic explanations</p>
</def>
</def-item>
<def-item>
<term>LLM</term>
<def>
<p>Large language model</p>
</def>
</def-item>
<def-item>
<term>MAP</term>
<def>
<p>Manifold absolute pressure</p>
</def>
</def-item>
<def-item>
<term>ML</term>
<def>
<p>Machine learning</p>
</def>
</def-item>
<def-item>
<term>MLP</term>
<def>
<p>Multi-layer perceptron</p>
</def>
</def-item>
<def-item>
<term>NGA</term>
<def>
<p>Non-dispersive gas analyzer</p>
</def>
</def-item>
<def-item>
<term>NLP</term>
<def>
<p>Natural language processing</p>
</def>
</def-item>
<def-item>
<term>OCR</term>
<def>
<p>Optical character recognition</p>
</def>
</def-item>
<def-item>
<term>PDP</term>
<def>
<p>Partial dependence plot</p>
</def>
</def-item>
<def-item>
<term>RPM</term>
<def>
<p>Revolutions per minute</p>
</def>
</def-item>
<def-item>
<term>SHAP</term>
<def>
<p>Shapley additive explanations</p>
</def>
</def-item>
<def-item>
<term>TPS</term>
<def>
<p>Throttle position sensor</p>
</def>
</def-item>
<def-item>
<term>VLM</term>
<def>
<p>Vision-language model</p>
</def>
</def-item>
<def-item>
<term>XAI</term>
<def>
<p>Explainable artificial intelligence</p>
</def>
</def-item>
<def-item>
<term>XGBoost</term>
<def>
<p>Extreme gradient boosting</p>
</def>
</def-item>
</def-list>
</glossary>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A</title>
<p>The appendix provides the complete set of prompt templates employed throughout this study, with the aim of enhancing methodological transparency and supporting reproducibility. These templates define the structured interactions used by the VLMs and LLMs at different stages of DRIVE, spanning explanation generation, evaluation, consistency analysis, and verification.</p>
<p><xref ref-type="table" rid="table-14">Table A1</xref> presents the prompt template used for analyzing LIME results, which guides the VLM to produce instance-level textual explanations by approximating the local behavior of the diagnostic model. <xref ref-type="table" rid="table-15">Table A2</xref> provides the corresponding prompt template for integrating SHAP results, focusing on stepwise feature attributions and contribution magnitudes derived from Shapley value decomposition. <xref ref-type="table" rid="table-16">Table A3</xref> describes the prompt template for unified report generation, in which an LLM synthesizes VLM-generated LIME and SHAP explanations into a coherent and structured diagnostic narrative.</p>
<p>In addition to explanation generation, the appendix includes prompts for evaluation and robustness analysis. <xref ref-type="table" rid="table-17">Table A4</xref> presents the prompt template used for Likert-scale evaluation of explanatory reports, enabling both VLM-as-a-Judge and human-aligned assessment based on predefined criteria. <xref ref-type="table" rid="table-18">Table A5</xref> introduces the prompt template for structured consistency and discrepancy analysis across XAI methods, supporting quantitative comparison of LIME- and SHAP-based explanations. Finally, <xref ref-type="table" rid="table-19">Table A6</xref> provides the prompt template for the verification agent, which performs a secondary review of generated explanations to check factual consistency, domain adherence, and potential hallucinations.</p>
<table-wrap id="table-14">
<label>Table A1</label>
<caption>
<title>Template for prompt in analyzing the result of LIME analysis.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Type</th>
<th>Prompt Detail</th>
</tr>
</thead>
<tbody>
<tr>
<td>Role description</td>
<td>You are a domain expert in vehicle engine diagnostics, machine learning, and explainable artificial intelligence. Your objective is to analyze the local interpretable model-agnostic explanations (LIME) visualization for a specific data instance classified by a vehicle engine fault diagnosis model, and to produce a precise and insightful textual explanation suitable for inclusion in a report.</td>
</tr>
<tr>
<td rowspan="21">Dataset description</td>
<td>&#x003C;Dataset description&#x003E;</td>
</tr>
<tr>

<td>1. Input variables</td>
</tr>
<tr>

<td>- The dataset contains input variables that capture both engine behavior and emission characteristics during vehicle acceleration.</td>
</tr>
<tr>

<td>- Engine behavior variables include manifold absolute pressure (MAP), throttle position sensor (TPS), engine torque (Force), power output, engine revolutions per minute (RPM), fuel consumption in liters per hour (L/H) and per 100 km (L/100 km), and vehicle speed.</td>
</tr>
<tr>

<td>- Emission-related variables comprise carbon monoxide (CO), hydrocarbons (HC), carbon dioxide (CO<sub>2</sub>), oxygen (O<sub>2</sub>), lambda, and air-fuel ratio (AFR).</td>
</tr>
<tr>

<td>- All input features are continuous numerical values, and their normality is determined based on how much they deviate from statistical norms.</td>
</tr>
<tr>

<td>- Engine-related measurements such as MAP, TPS, Force, Power, and RPM are considered abnormal when their values are excessively high relative to standard distribution patterns.</td>
</tr>
<tr>

<td>- Fuel consumption metrics such as L/H and L/100 km are also treated as abnormal if they exceed typical ranges observed in the dataset.</td>
</tr>
<tr>

<td>- Vehicle speed is regarded as abnormal when it is either too low or too high compared to the statistical average.</td>
</tr>
<tr>

<td>- Gas emission variables are labeled abnormal if their values significantly surpass thresholds derived from the summary statistics.</td>
</tr>
<tr>

<td>- All input variables were normalized to a range between 0 and 1 using Min-Max normalization.</td>
</tr>
<tr>

<td>- The mean, standard deviation, and median values for each input variable are as follows:</td>
</tr>
<tr>

<td>{Mean, standard deviation, median values for each input variable}</td>
</tr>
<tr>

<td>2. Output variables</td>
</tr>
<tr>

<td>- The output variable is a categorical label that represents the engine&#x2019;s operational state. There are four classes, each indicating a specific engine condition relevant to diagnostic analysis:</td>
</tr>
<tr>

<td>(1) No fault: Indicates a normal engine operating state with no observable anomalies in sensor readings or fuel-air mixture ratios. All engine and emission parameters fall within statistically expected ranges.</td>
</tr>
<tr>

<td>(2) Rich mixture: Characterized by excessive fuel relative to the amount of air in the combustion chamber. This condition may result from&#x2014;Incorrect sensor performance&#x2014;High fuel pressure&#x2014;Defective fuel injector&#x2014;Malfunctioning pressure regulator&#x2014;Clogged air filter&#x2014;Clogged fuel return line</td>
</tr>
<tr>

<td>(3) Lean mixture: Occurs when there is too much air relative to fuel in the combustion process. This fault type may be due to&#x2014;Incorrect sensor performance&#x2014;Low fuel pressure&#x2014;Defective injector&#x2014;Faulty pressure regulator</td>
</tr>
<tr>

<td>4) Low voltage: Refers to ignition-related issues resulting in misfires or weak combustion. Common causes include:&#x2014;Worn spark plugs&#x2014;Faulty ignition cables&#x2014;Defective coil&#x2014;Faulty sensor wiring</td>
</tr>
<tr>

<td>- This dataset is intended for multi-class classification tasks and is designed to support model training, evaluation, and interpretability in vehicle engine fault diagnosis.</td>
</tr>
<tr>

<td>&#x003C;/Dataset description&#x003E;</td>
</tr>
<tr>
<td rowspan="7">Task specification</td>
<td>&#x003C;Task&#x003E;</td>
</tr>
<tr>

<td>This task involves interpreting the output of LIME, an explainable artificial intelligence technique that approximates complex model behavior using locally faithful, interpretable models. LIME helps identify how individual features influence a model&#x2019;s prediction for a specific instance by generating a set of weighted feature conditions. The image shows the LIME explanation for a specific instance classified by {Prediction model} model trained to identify vehicle engine faults.</td>
</tr>
<tr>

<td>The LIME visualization includes:</td>
</tr>
<tr>

<td>- A bar plot of &#x002A;&#x002A;prediction probabilities&#x002A;&#x002A;, showing the model&#x2019;s confidence across the four target classes. In this case, the instance was classified as class &#x002A;&#x002A;{Predicted class}&#x002A;&#x002A; with the &#x002A;&#x002A;highest probability of {Value of highest probability}&#x002A;&#x002A;</td>
</tr>
<tr>

<td>- A &#x002A;&#x002A;decision path&#x002A;&#x002A; composed of feature-based rules that contributed to the prediction. Each horizontal bar represents a feature condition satisfied for this input. The &#x002A;&#x002A;value on the right side&#x002A;&#x002A; of each bar indicates the &#x002A;&#x002A;local contribution (weight)&#x002A;&#x002A; to the final class prediction. Positive values support the predicted class, while negative values push against it. Longer bars indicate stronger influence on the prediction decision. The actual input feature values used in this instance.</td>
</tr>
<tr>

<td>The following LIME feature conditions and their contribution values were observed: {LIME feature conditions and contribution values}</td>
</tr>
<tr>

<td>&#x003C;/Task&#x003E;</td>
</tr>
<tr>
<td rowspan="10">Interpretation guideline</td>
<td>&#x003C;Interpretation guideline&#x003E;</td>
</tr>
<tr>

<td>- Please provide a structured explanation that includes:</td>
</tr>
<tr>

<td>1. &#x002A;&#x002A;Key Features&#x002A;&#x002A;: A summary of the most influential input features in the local explanation, including those with both positive and negative contributions to the predicted fault class, along with their respective contribution values.</td>
</tr>
<tr>

<td>2. &#x002A;&#x002A;Feature Prioritization&#x002A;&#x002A;: An explanation of how the model locally prioritizes specific sensor variables to arrive at the fault classification for this particular input instance.</td>
</tr>
<tr>

<td>3. &#x002A;&#x002A;Low-Impact Features&#x002A;&#x002A;: An interpretation of features that exhibit minimal or near-zero local contribution, with possible reasons for their limited impact on the model&#x2019;s decision in this context.</td>
</tr>
<tr>

<td>4. &#x002A;&#x002A;Decision Logic&#x002A;&#x002A;: A logical analysis of the decision-making process inferred from the LIME explanation, focusing on how combinations of sensor measurements led the model to assign the instance to the predicted fault class.</td>
</tr>
<tr>

<td>- The explanation should focus exclusively on the local behavior of the model for the given input instance and must not generalize the importance of features across the dataset.</td>
</tr>
<tr>

<td>- The analysis should be clear, technically rigorous, and appropriate for inclusion in a peer-reviewed research paper.</td>
</tr>
<tr>

<td>- Use formal language suitable for an academic audience in the fields of machine learning, vehicle engine fault diagnostics, and explainable artificial intelligence.</td>
</tr>
<tr>

<td>&#x003C;/Interpretation guideline&#x003E;</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-15">
<label>Table A2</label>
<caption>
<title>Template for prompt in analyzing the result of SHAP analysis.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Type</th>
<th>Prompt Detail</th>
</tr>
</thead>
<tbody>
<tr>
<td>Role description</td>
<td>You are a domain expert in vehicle engine diagnostics, machine learning, and explainable artificial intelligence. Your objective is to analyze the Shapley additive explanations (SHAP) visualization for a specific data instance classified by a vehicle engine fault diagnosis model, and to produce a precise and insightful textual explanation suitable for inclusion in a report.</td>
</tr>
<tr>
<td rowspan="21">Dataset description</td>
<td>&#x003C;Dataset description&#x003E;</td>
</tr>
<tr>

<td>1. Input variables</td>
</tr>
<tr>

<td>- The dataset contains input variables that capture both engine behavior and emission characteristics during vehicle acceleration.</td>
</tr>
<tr>

<td>- Engine behavior variables include manifold absolute pressure (MAP), throttle position sensor (TPS), engine torque (Force), power output, engine revolutions per minute (RPM), fuel consumption in liters per hour (L/H) and per 100 km (L/100 km), and vehicle speed.</td>
</tr>
<tr>

<td>- Emission-related variables comprise carbon monoxide (CO), hydrocarbons (HC), carbon dioxide (CO<sub>2</sub>), oxygen (O<sub>2</sub>), lambda, and air-fuel ratio (AFR).</td>
</tr>
<tr>

<td>- All input features are continuous numerical values, and their normality is determined based on how much they deviate from statistical norms.</td>
</tr>
<tr>

<td>- Engine-related measurements such as MAP, TPS, Force, Power, and RPM are considered abnormal when their values are excessively high relative to standard distribution patterns.</td>
</tr>
<tr>

<td>- Fuel consumption metrics such as L/H and L/100 km are also treated as abnormal if they exceed typical ranges observed in the dataset.</td>
</tr>
<tr>

<td>- Vehicle speed is regarded as abnormal when it is either too low or too high compared to the statistical average.</td>
</tr>
<tr>

<td>- Gas emission variables are labeled abnormal if their values significantly surpass thresholds derived from the summary statistics.</td>
</tr>
<tr>

<td>- All input variables were normalized to a range between 0 and 1 using Min-Max normalization.</td>
</tr>
<tr>

<td>- The mean, standard deviation, and median values for each input variable are as follows:</td>
</tr>
<tr>

<td>{Mean, standard deviation, median values for each input variable}</td>
</tr>
<tr>

<td>2. Output variables</td>
</tr>
<tr>

<td>- The output variable is a categorical label that represents the engine&#x2019;s operational state. There are four classes, each indicating a specific engine condition relevant to diagnostic analysis:</td>
</tr>
<tr>

<td>(1) No fault: Indicates a normal engine operating state with no observable anomalies in sensor readings or fuel-air mixture ratios. All engine and emission parameters fall within statistically expected ranges.</td>
</tr>
<tr>

<td>(2) Rich mixture: Characterized by excessive fuel relative to the amount of air in the combustion chamber. This condition may result from&#x2014;Incorrect sensor performance&#x2014;High fuel pressure&#x2014;Defective fuel injector&#x2014;Malfunctioning pressure regulator&#x2014;Clogged air filter&#x2014;Clogged fuel return line</td>
</tr>
<tr>

<td>(3) Lean mixture: Occurs when there is too much air relative to fuel in the combustion process. This fault type may be due to&#x2014;Incorrect sensor performance&#x2014;Low fuel pressure&#x2014;Defective injector&#x2014;Faulty pressure regulator</td>
</tr>
<tr>

<td>(4) Low voltage: Refers to ignition-related issues resulting in misfires or weak combustion. Common causes include&#x2014;Worn spark plugs&#x2014;Faulty ignition cables&#x2014;Defective coil&#x2014;Faulty sensor wiring</td>
</tr>
<tr>

<td>- This dataset is intended for multi-class classification tasks and is designed to support model training, evaluation, and interpretability in vehicle engine fault diagnosis.</td>
</tr>
<tr>

<td>&#x003C;/Dataset description&#x003E;</td>
</tr>
<tr>
<td rowspan="9">Task specification</td>
<td>&#x003C;Task&#x003E;</td>
</tr>
<tr>

<td>This task involves interpreting the output of SHAP, an explainable artificial intelligence technique based on cooperative game theory that assigns Shapley values to quantify each feature&#x2019;s contribution to the model&#x2019;s prediction. SHAP helps identify both the direction and magnitude of influence that input features exert on a specific model outcome. The image shows the SHAP Waterfall Plot for an instance classified by the {Prediction model} trained to detect vehicle engine faults.</td>
</tr>
<tr>

<td>The SHAP visualization includes:</td>
</tr>
<tr>

<td>- A stepwise decomposition of the prediction, beginning from the model&#x2019;s expected (baseline) value and sequentially adding the positive and negative contributions of individual features until the final prediction is reached. In this case, the instance was classified as class &#x002A;&#x002A;{Predicted class}&#x002A;&#x002A; with the &#x002A;&#x002A;highest probability of {Value of highest probability}&#x002A;&#x002A;</td>
</tr>
<tr>

<td>- Horizontal axis values representing the model output, where positive contributions increase the predicted value and negative contributions decrease it.</td>
</tr>
<tr>

<td>- Color-coded bars that indicate the direction and magnitude of each feature&#x2019;s effect (e.g., pink for positive impact, blue for negative).</td>
</tr>
<tr>

<td>- Feature labels aligned with each bar, together with their corresponding Shapley values, which specify the quantitative contribution of each feature to the overall prediction.</td>
</tr>
<tr>

<td>The following SHAP feature contributions were observed: {SHAP feature names and contribution values}</td>
</tr>
<tr>

<td>&#x003C;/Task&#x003E;</td>
</tr>
<tr>
<td rowspan="10">Interpretation guideline</td>
<td>&#x003C;Interpretation guideline&#x003E;</td>
</tr>
<tr>

<td>- Please provide a structured explanation that includes:</td>
</tr>
<tr>

<td>1. &#x002A;&#x002A;Key Features&#x002A;&#x002A;: An identification of the most influential input features in the SHAP Waterfall Plot, highlighting both positive and negative Shapley values along with their respective contribution magnitudes to the predicted fault class.</td>
</tr>
<tr>

<td>2. &#x002A;&#x002A;Stepwise Contribution Process&#x002A;&#x002A;: An explanation of how the model&#x2019;s prediction is constructed from the baseline (expected value) as individual feature contributions are sequentially added or subtracted, with emphasis on the order and cumulative impact of these contributions.</td>
</tr>
<tr>

<td>3. &#x002A;&#x002A;Low-Impact Features&#x002A;&#x002A;: A discussion of features with near-zero or minimal Shapley values, including possible reasons for their limited influence on the model&#x2019;s decision in this specific instance.</td>
</tr>
<tr>

<td>4. &#x002A;&#x002A;Decision Logic&#x002A;&#x002A;: An analysis of the local decision-making process inferred from the SHAP Waterfall Plot, focusing on how the combined contributions of sensor features led the model to assign the instance to the predicted engine fault class.</td>
</tr>
<tr>

<td>- The explanation should focus exclusively on the local interpretation for the given input instance, without generalizing feature importance across the dataset.</td>
</tr>
<tr>

<td>- The analysis should be clear, technically rigorous, and appropriate for inclusion in a peer-reviewed research paper.</td>
</tr>
<tr>

<td>- Use formal language suitable for an academic audience in the fields of machine learning, vehicle engine fault diagnostics, and explainable artificial intelligence.</td>
</tr>
<tr>

<td>&#x003C;/Interpretation guideline&#x003E;</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-16">
<label>Table A3</label>
<caption>
<title>Template for prompt in generating unified report using the LLM.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Type</th>
<th>Prompt Detail</th>
</tr>
</thead>
<tbody>
<tr>
<td>Role description</td>
<td>You are a domain expert in vehicle engine diagnostics, machine learning, and explainable artificial intelligence. Your objective is to synthesize textual explanations generated from the vision-language model (VLM) based local interpretable model-agnostic explanations (LIME) and Shapley additive explanations (SHAP) analyses into a unified and coherent diagnostic report. The report should highlight consistencies, reconcile differences, and present the reasoning process in a structured and accessible form suitable for technical reporting.</td>
</tr>
<tr>
<td rowspan="21">Dataset description</td>
<td>&#x003C;Dataset description&#x003E;</td>
</tr>
<tr>

<td>1. Input variables</td>
</tr>
<tr>

<td>- The dataset contains input variables that capture both engine behavior and emission characteristics during vehicle acceleration.</td>
</tr>
<tr>

<td>- Engine behavior variables include manifold absolute pressure (MAP), throttle position sensor (TPS), engine torque (Force), power output, engine revolutions per minute (RPM), fuel consumption in liters per hour (L/H) and per 100 km (L/100 km), and vehicle speed.</td>
</tr>
<tr>

<td>- Emission-related variables comprise carbon monoxide (CO), hydrocarbons (HC), carbon dioxide (CO<sub>2</sub>), oxygen (O<sub>2</sub>), lambda, and air-fuel ratio (AFR).</td>
</tr>
<tr>

<td>- All input features are continuous numerical values, and their normality is determined based on how much they deviate from statistical norms.</td>
</tr>
<tr>

<td>- Engine-related measurements such as MAP, TPS, Force, Power, and RPM are considered abnormal when their values are excessively high relative to standard distribution patterns.</td>
</tr>
<tr>

<td>- Fuel consumption metrics such as L/H and L/100 km are also treated as abnormal if they exceed typical ranges observed in the dataset.</td>
</tr>
<tr>

<td>- Vehicle speed is regarded as abnormal when it is either too low or too high compared to the statistical average.</td>
</tr>
<tr>

<td>- Gas emission variables are labeled abnormal if their values significantly surpass thresholds derived from the summary statistics.</td>
</tr>
<tr>

<td>- All input variables were normalized to a range between 0 and 1 using Min-Max normalization.</td>
</tr>
<tr>

<td>- The mean, standard deviation, and median values for each input variable are as follows:</td>
</tr>
<tr>

<td>{Mean, standard deviation, median values for each input variable}</td>
</tr>
<tr>

<td>2. Output variables</td>
</tr>
<tr>

<td>- The output variable is a categorical label that represents the engine&#x2019;s operational state. There are four classes, each indicating a specific engine condition relevant to diagnostic analysis:</td>
</tr>
<tr>

<td>(1) No fault: Indicates a normal engine operating state with no observable anomalies in sensor readings or fuel-air mixture ratios. All engine and emission parameters fall within statistically expected ranges.</td>
</tr>
<tr>

<td>(2) Rich mixture: Characterized by excessive fuel relative to the amount of air in the combustion chamber. This condition may result from&#x2014;Incorrect sensor performance&#x2014;High fuel pressure&#x2014;Defective fuel injector&#x2014;Malfunctioning pressure regulator&#x2014;Clogged air filter&#x2014;Clogged fuel return line</td>
</tr>
<tr>

<td>(3) Lean mixture: Occurs when there is too much air relative to fuel in the combustion process. This fault type may be due to&#x2014;Incorrect sensor performance&#x2014;Low fuel pressure&#x2014;Defective injector&#x2014;Faulty pressure regulator</td>
</tr>
<tr>

<td>(4) Low voltage: Refers to ignition-related issues resulting in misfires or weak combustion. Common causes include&#x2014;Worn spark plugs&#x2014;Faulty ignition cables&#x2014;Defective coil&#x2014;Faulty sensor wiring</td>
</tr>
<tr>

<td>- This dataset is intended for multi-class classification tasks and is designed to support model training, evaluation, and interpretability in vehicle engine fault diagnosis.</td>
</tr>
<tr>

<td>&#x003C;/Dataset description&#x003E;</td>
</tr>
<tr>
<td rowspan="12">Task specification</td>
<td>&#x003C;Task&#x003E;</td>
</tr>
<tr>

<td>This task involves integrating VLM-based LIME and SHAP analyses for the same prediction instance into a single and unified diagnostic report.</td>
</tr>
<tr>

<td>The input to this task consists of two textual explanations:</td>
</tr>
<tr>

<td>1. VLM-based explanation of the LIME visualization results</td>
</tr>
<tr>

<td>2. VLM-based explanation of the SHAP waterfall plot results</td>
</tr>
<tr>

<td>The unified report should:</td>
</tr>
<tr>

<td>- Consolidate overlapping findings from LIME and SHAP to highlight strong, consistent evidence.</td>
</tr>
<tr>

<td>- Reconcile differences between the two methods, clarifying complementary insights and potential sources of discrepancy.</td>
</tr>
<tr>

<td>- Maintain logical consistency, technical rigor, and diagnostic relevance, ensuring that the final explanation accurately reflects the decision-making process of the underlying model.</td>
</tr>
<tr>

<td>The following VLM-based explanation of the LIME visualization results were observed: {VLM-based explanation of the LIME visualization results}</td>
</tr>
<tr>

<td>The following VLM-based explanation of the SHAP waterfall plot results were observed: {VLM-based explanation of the SHAP waterfall plot results}</td>
</tr>
<tr>

<td>&#x003C;/Task&#x003E;</td>
</tr>
<tr>
<td rowspan="9">Interpretation guideline</td>
<td>&#x003C;Interpretation guideline&#x003E;</td>
</tr>
<tr>

<td>- Please provide a structured explanation that includes:</td>
</tr>
<tr>

<td>1. &#x002A;&#x002A;Summary of Predicted Class &#x002A;&#x002A;: A clear statement of the model&#x2019;s predicted engine fault condition for the given instance.</td>
</tr>
<tr>

<td>2. &#x002A;&#x002A;Consistent Feature Contributions&#x002A;&#x002A;: Features identified by both LIME and SHAP as influential, with discussion of their relative importance.</td>
</tr>
<tr>

<td>3. &#x002A;&#x002A;Method-specific Insights&#x002A;&#x002A;: Unique contributions highlighted by either LIME or SHAP, including potential reasons for these differences.</td>
</tr>
<tr>

<td>4. &#x002A;&#x002A;Integrated Decision Logic&#x002A;&#x002A;: A synthesized explanation of how the combined evidence from both methods supports the final diagnostic interpretation.</td>
</tr>
<tr>

<td>- The explanation should focus exclusively on the local interpretation for the given input instance, without generalizing across the dataset.</td>
</tr>
<tr>

<td>- The report must be written in clear, formal academic language, ensuring precision, coherence, and suitability for inclusion in a peer-reviewed research paper.</td>
</tr>
<tr>

<td>&#x003C;/Interpretation guideline&#x003E;</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-17">
<label>Table A4</label>
<caption>
<title>Template for prompt in Likert-scale evaluation of explanatory reports.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Type</th>
<th>Prompt detail</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">Role description</td>
<td>You are an expert evaluator with professional knowledge in vehicle engine systems, machine learning, and explainable artificial intelligence (XAI). Your task is to objectively assess the quality of a generated textual explanation that interprets the diagnostic output of a vehicle engine fault classification model.</td>
</tr>
<tr>

<td>You must evaluate the explanation strictly based on the provided content, without introducing new assumptions or external information.</td>
</tr>
<tr>
<td rowspan="5">Input description</td>
<td>&#x003C;Input description&#x003E;</td>
</tr>
<tr>

<td>You are given:</td>
</tr>
<tr>

<td><list list-type="order">
<list-item>
<p>A visual explanation output (e.g., local interpretable model-agnostic explanations (LIME) or Shapley additive explanations (SHAP) plot) generated from a vehicle engine fault diagnosis model.</p></list-item>
<list-item>
<p>A textual explanation generated by a Vision-Language Model (VLM) or Large Language Model (LLM) that describes and interprets the visual explanation.</p></list-item>
</list></td>
</tr>
<tr>

<td>The explanation concerns an instance-level vehicle engine fault diagnosis task.</td>
</tr>
<tr>

<td>&#x003C;/Input description&#x003E;</td>
</tr>
<tr>
<td rowspan="6">Task specification</td>
<td>&#x003C;Task specification&#x003E;</td>
</tr>
<tr>

<td>Evaluate the provided textual explanation using a 5-point Likert scale according to the five predefined criteria listed below.</td>
</tr>
<tr>

<td>Each criterion must be scored independently on a scale from:</td>
</tr>
<tr>

<td><list list-type="bullet">
<list-item>
<p>1 &#x003D; Very dissatisfied</p></list-item>
<list-item>
<p>2 &#x003D; Dissatisfied</p></list-item>
<list-item>
<p>3 &#x003D; Neutral</p></list-item>
<list-item>
<p>4 &#x003D; Satisfied</p></list-item>
<list-item>
<p>5 &#x003D; Very satisfied</p></list-item>
</list></td>
</tr>
<tr>

<td>Your evaluation should reflect how well the explanation satisfies each criterion from the perspective of a domain expert.</td>
</tr>
<tr>

<td>&#x003C;/Task specification&#x003E;</td>
</tr>
<tr>
<td rowspan="4">Evaluation criteria and guideline</td>
<td>&#x003C;Evaluation criteria and guideline&#x003E;</td>
</tr>
<tr>

<td>Please assess the explanation according to the following criteria:</td>
</tr>
<tr>

<td><list list-type="order">
<list-item>
<p>Comprehensibility</p>
<p>Evaluate whether the explanation is clearly written, logically structured, and easy to understand for users with technical background in vehicle diagnostics.</p></list-item>
<list-item>
<p>Faithfulness</p>
<p>Assess the extent to which the explanation accurately represents the information conveyed in the visual explanation (e.g., feature importance, contribution direction, and relative magnitude), without distortion or hallucination.</p></list-item>
<list-item>
<p>Domain Relevance</p>
<p>Determine whether the explanation aligns with established principles and expert knowledge in vehicle engine fault diagnosis and combustion behavior.</p></list-item>
<list-item>
<p>Explanatory Value</p>
<p>Evaluate how useful the textual explanation is in improving user understanding compared to the visual explanation alone.</p></list-item>
<list-item>
<p>Reliability</p>
<p>Assess the perceived credibility and trustworthiness of the explanation from a domain expert&#x2019;s perspective, considering consistency, technical rigor, and absence of misleading statements.</p></list-item>
</list></td>
</tr>
<tr>

<td>&#x003C;/Evaluation criteria and guideline&#x003E;</td>
</tr>
<tr>
<td rowspan="10">Output format</td>
<td>&#x003C;Output format&#x003E;</td>
</tr>
<tr>

<td>Provide the evaluation results in the following JSON format:</td>
</tr>
<tr>

<td>{</td>
</tr>
<tr>

<td>&#x2002;&#x2002;&#x201C;Comprehensibility&#x201D;: &#x003C;score&#x003E;,</td>
</tr>
<tr>

<td>&#x2002;&#x2002;&#x201C;Faithfulness&#x201D;: &#x003C;score&#x003E;,</td>
</tr>
<tr>

<td>&#x2002;&#x2002;&#x201C;Domain_Relevance&#x201D;: &#x003C;score&#x003E;,</td>
</tr>
<tr>

<td>&#x2002;&#x2002;&#x201C;Explanatory_Value&#x201D;: &#x003C;score&#x003E;,</td>
</tr>
<tr>

<td>&#x2002;&#x2002;&#x201C;Reliability&#x201D;: &#x003C;score&#x003E;</td>
</tr>
<tr>

<td>}</td>
</tr>
<tr>

<td>&#x003C;/Output format&#x003E;</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-18">
<label>Table A5</label>
<caption>
<title>Template for prompt in analyzing structured consistency.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="left"/> </colgroup>
<thead>
<tr>
<th>Type</th>
<th>Prompt detail</th>
</tr>
</thead>
<tbody>
<tr>
<td>Role description</td>
<td>You are an explainable artificial intelligence (XAI) consistency analysis agent specializing in vehicle engine fault diagnosis, machine learning interpretability, and post-hoc explanation methods. Your role is to quantitatively evaluate the consistency and discrepancies between local interpretable model-agnostic explanations (LIME)-based and Shapley additive explanations (SHAP)-based local explanations generated for the same prediction instance of a vehicle engine fault diagnosis model. You are required to analyze feature-level attributions produced by LIME and SHAP, and to provide objective, metric-based assessments of their agreement and divergence.</td>
</tr>
<tr>
<td rowspan="18">Input description</td>
<td>&#x003C;Input description&#x003E;</td>
</tr>
<tr>

<td>The input consists of two structured explanation results derived from the same vehicle engine fault diagnosis instance:</td>
</tr>
<tr>

<td>(1) LIME-based explanation:</td>
</tr>
<tr>

<td>- Predicted fault class and prediction probability</td>
</tr>
<tr>

<td>- A list of input features with:</td>
</tr>
<tr>

<td>&#x2022;&#x2002;feature name</td>
</tr>
<tr>

<td>&#x2022;&#x2002;importance magnitude</td>
</tr>
<tr>

<td>&#x2022;&#x2002;contribution direction (positive or negative)</td>
</tr>
<tr>

<td>&#x2022;&#x2002;importance rank</td>
</tr>
<tr>

<td>(2) SHAP-based explanation:</td>
</tr>
<tr>

<td>- Baseline prediction value and final output value</td>
</tr>
<tr>

<td>- A list of input features with:</td>
</tr>
<tr>

<td>&#x2022;&#x2002;feature name</td>
</tr>
<tr>

<td>&#x2022;&#x2002;importance magnitude</td>
</tr>
<tr>

<td>&#x2022;&#x2002;contribution direction (positive or negative)</td>
</tr>
<tr>

<td>&#x2022;&#x2002;importance rank</td>
</tr>
<tr>

<td>Both explanations are generated from the same ML-based diagnostic model and refer to the same input instance. Feature importance values may differ in scale across methods and should be normalized before comparison.</td>
</tr>
<tr>

<td>&#x003C;/Input description&#x003E;</td>
</tr>
<tr>
<td rowspan="18">Task specification</td>
<td>&#x003C;Task specification&#x003E;</td>
</tr>
<tr>

<td>Your task is to perform a quantitative consistency analysis between the LIME-based and SHAP-based explanations by executing the following steps:</td>
</tr>
<tr>

<td>1. Identify the overlapping feature set shared by both LIME and SHAP.</td>
</tr>
<tr>

<td>2. Normalize feature importance values independently within each method.</td>
</tr>
<tr>

<td>3. Compute the following consistency metrics:</td>
</tr>
<tr>

<td>(a) Feature Overlap Ratio (FOR)</td>
</tr>
<tr>

<td>&#x2013; the ratio of shared features to the union of features.</td>
</tr>
<tr>

<td>(b) Rank Consistency</td>
</tr>
<tr>

<td>&#x2013; Spearman rank correlation coefficient computed on overlapping features.</td>
</tr>
<tr>

<td>(c) Directional Consistency Score (DCS)</td>
</tr>
<tr>

<td>&#x2013; the proportion of overlapping features with matching contribution</td>
</tr>
<tr>

<td>directions.</td>
</tr>
<tr>

<td>(d) Normalized Importance Divergence (NID)</td>
</tr>
<tr>

<td>&#x2013; the mean absolute difference between normalized importance values</td>
</tr>
<tr>

<td>across overlapping features.</td>
</tr>
<tr>

<td>4. Summarize all metric values in a structured format.</td>
</tr>
<tr>

<td>5. Provide a concise interpretation of the overall consistency and highlight any notable discrepancies between the two explanation methods.</td>
</tr>
<tr>

<td>&#x003C;/Task specification&#x003E;</td>
</tr>
<tr>
<td rowspan="10">Interpretation guideline</td>
<td>&#x003C;Interpretation guideline&#x003E;</td>
</tr>
<tr>

<td>When interpreting the results, follow these guidelines:</td>
</tr>
<tr>

<td>- Focus strictly on local explanations for the given instance.</td>
</tr>
<tr>

<td>- Do not introduce new features or external assumptions.</td>
</tr>
<tr>

<td>- Treat LIME and SHAP as complementary methods with different attribution mechanisms.</td>
</tr>
<tr>

<td>- Interpret discrepancies in feature importance magnitude as potential methodological</td>
</tr>
<tr>

<td>differences rather than errors unless directions conflict.</td>
</tr>
<tr>

<td>- Use quantitative metrics as the primary evidence for assessing consistency.</td>
</tr>
<tr>

<td>- Provide interpretations in a neutral, academically rigorous tone.</td>
</tr>
<tr>

<td>&#x003C;/Interpretation guideline&#x003E;</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-19">
<label>Table A6</label>
<caption>
<title>Template for prompt in verification agent.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Type</th>
<th>Prompt detail</th>
</tr>
</thead>
<tbody>
<tr>
<td>Role description</td>
<td>You are a domain expert in vehicle engine diagnostics, machine learning, and explainable artificial intelligence (XAI). Your role is to verify the factual correctness, domain consistency, and logical validity of explanation texts generated by a vision-language model (VLM) or a large language model (LLM).</td>
</tr>
<tr>
<td rowspan="21">Dataset description</td>
<td>&#x003C;Dataset description&#x003E;</td>
</tr>
<tr>

<td>1. Input variables</td>
</tr>
<tr>

<td>- The dataset contains input variables that capture both engine behavior and emission characteristics during vehicle acceleration.</td>
</tr>
<tr>

<td>- Engine behavior variables include manifold absolute pressure (MAP), throttle position sensor (TPS), engine torque (Force), power output, engine revolutions per minute (RPM), fuel consumption in liters per hour (L/H) and per 100 km (L/100 km), and vehicle speed.</td>
</tr>
<tr>

<td>- Emission-related variables comprise carbon monoxide (CO), hydrocarbons (HC), carbon dioxide (CO<sub>2</sub>), oxygen (O<sub>2</sub>), lambda, and air-fuel ratio (AFR).</td>
</tr>
<tr>

<td>- All input features are continuous numerical values, and their normality is determined based on how much they deviate from statistical norms.</td>
</tr>
<tr>

<td>- Engine-related measurements such as MAP, TPS, Force, Power, and RPM are considered abnormal when their values are excessively high relative to standard distribution patterns.</td>
</tr>
<tr>

<td>- Fuel consumption metrics such as L/H and L/100 km are also treated as abnormal if they exceed typical ranges observed in the dataset.</td>
</tr>
<tr>

<td>- Vehicle speed is regarded as abnormal when it is either too low or too high compared to the statistical average.</td>
</tr>
<tr>

<td>- Gas emission variables are labeled abnormal if their values significantly surpass thresholds derived from the summary statistics.</td>
</tr>
<tr>

<td>- All input variables were normalized to a range between 0 and 1 using Min-Max normalization.</td>
</tr>
<tr>

<td>- The mean, standard deviation, and median values for each input variable are as follows:</td>
</tr>
<tr>

<td>{Mean, standard deviation, median values for each input variable}</td>
</tr>
<tr>

<td>2. Output variables</td>
</tr>
<tr>

<td>- The output variable is a categorical label that represents the engine&#x2019;s operational state. There are four classes, each indicating a specific engine condition relevant to diagnostic analysis:</td>
</tr>
<tr>

<td>(1) No fault: Indicates a normal engine operating state with no observable anomalies in sensor readings or fuel-air mixture ratios. All engine and emission parameters fall within statistically expected ranges.</td>
</tr>
<tr>

<td>(2) Rich mixture: Characterized by excessive fuel relative to the amount of air in the combustion chamber. This condition may result from&#x2014;Incorrect sensor performance&#x2014;High fuel pressure&#x2014;Defective fuel injector&#x2014;Malfunctioning pressure regulator&#x2014;Clogged air filter&#x2014;Clogged fuel return line</td>
</tr>
<tr>

<td>(3) Lean mixture: Occurs when there is too much air relative to fuel in the combustion process. This fault type may be due to&#x2014;Incorrect sensor performance&#x2014;Low fuel pressure&#x2014;Defective injector&#x2014;Faulty pressure regulator</td>
</tr>
<tr>

<td>(4) Low voltage: Refers to ignition-related issues resulting in misfires or weak combustion. Common causes include&#x2014;Worn spark plugs&#x2014;Faulty ignition cables&#x2014;Defective coil&#x2014;Faulty sensor wiring</td>
</tr>
<tr>

<td>- This dataset is intended for multi-class classification tasks and is designed to support model training, evaluation, and interpretability in vehicle engine fault diagnosis.</td>
</tr>
<tr>

<td>&#x003C;/Dataset description&#x003E;</td>
</tr>
<tr>
<td rowspan="6">Task specification</td>
<td>&#x003C;Task specification&#x003E;</td>
</tr>
<tr>

<td>You are given:</td>
</tr>
<tr>

<td><list list-type="order">
<list-item>
<p>The original XAI output (local interpretable model-agnostic explanations (LIME) rules or Shapley additive explanations (SHAP) waterfall information),</p></list-item>
<list-item>
<p>The model prediction (predicted class and probability),</p></list-item>
<list-item>
<p>A generated textual explanation produced by a VLM or LLM.</p></list-item>
</list></td>
</tr>
<tr>

<td>Your task is to verify whether the explanation:</td>
</tr>
<tr>

<td><list list-type="bullet">
<list-item>
<p>Faithfully reflects the provided XAI outputs,</p></list-item>
<list-item>
<p>Accurately represents feature contribution directions and relative importance,</p></list-item>
<list-item>
<p>Avoids introducing unsupported claims, speculative reasoning, or domain-inconsistent statements.</p></list-item>
</list></td>
</tr>
<tr>

<td>&#x003C;/Task specification&#x003E;</td>
</tr>
<tr>
<td rowspan="5">Interpretation guideline</td>
<td>&#x003C;Interpretation guideline&#x003E;</td>
</tr>
<tr>

<td>Perform the verification by following these steps:</td>
</tr>
<tr>

<td><list list-type="order">
<list-item>
<p>Factual Consistency Check
<list list-type="simple">
<list-item><label>&#x25CB;</label>
<p>Confirm that all mentioned features, contribution directions, and importance claims are grounded in the given XAI output.</p></list-item>
</list></p></list-item>
<list-item>
<p>Domain Consistency Check
<list list-type="simple">
<list-item><label>&#x25CB;</label>
<p>Assess whether the explanation aligns with established principles of vehicle engine operation and combustion diagnostics.</p></list-item>
</list></p></list-item>
<list-item>
<p>Hallucination Detection
<list list-type="simple">
<list-item><label>&#x25CB;</label>
<p>Identify any statements that cannot be traced back to the XAI output or domain knowledge.</p></list-item>
</list></p></list-item>
<list-item>
<p>Verdict Generation
<list list-type="simple">
<list-item><label>&#x25CB;</label>
<p>Classify the explanation as:
<list list-type="simple">
<list-item><label>&#x025AA;</label>
<p>Verified, or</p></list-item>
<list-item><label>&#x025AA;</label>
<p>Minor Issues Detected, or</p></list-item>
<list-item><label>&#x025AA;</label>
<p>Hallucination Detected.</p></list-item>
</list></p></list-item>
<list-item><label>&#x25CB;</label>
<p>If issues are detected, explicitly describe which parts are unsupported and why.</p></list-item>
</list></p></list-item>
</list></td>
</tr>
<tr>

<td>Output your assessment in a structured textual format without introducing new diagnostic information.</td>
</tr>
<tr>

<td>&#x003C;/Interpretation guideline&#x003E;</td>
</tr>
</tbody>
</table>
</table-wrap>


<p>Collectively, these prompt templates constitute a core component of DRIVE. By explicitly formalizing each stage of explanation generation, evaluation, and verification, the appendix ensures consistent interpretation behavior across models and experiments, while facilitating transparency, extensibility, and reproducibility of the proposed method.</p>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Grzesiak</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sulich</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Car engines comparative analysis: sustainable approach</article-title>. <source>Energies</source>. <year>2022</year>;<volume>15</volume>(<issue>14</issue>):<fpage>5170</fpage>. doi:<pub-id pub-id-type="doi">10.3390/en15145170</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Akbal&#x0131;k</surname> <given-names>F</given-names></string-name>, <string-name><surname>Y&#x0131;ld&#x0131;z</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ertu&#x011F;rul</surname> <given-names>&#x00D6;F</given-names></string-name>, <string-name><surname>Zan</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Engine fault detection by sound analysis and machine learning</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>15</issue>):<fpage>6532</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app14156532</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cui</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Research on common faults and maintenance countermeasures of automobile engines</article-title>. <source>J Educ Teach Soc Stud</source>. <year>2023</year>;<volume>5</volume>(<issue>2</issue>):<fpage>68</fpage>. doi:<pub-id pub-id-type="doi">10.22158/jetss.v5n2p68</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Fault diagnosis of vehicle engine based on analytic hierarchy process and neural network</article-title>. In: <conf-name>Proceedings of the 2018 International Conference on Mechanical, Electronic, Control and Automation Engineering (MECAE 2018)</conf-name>. <publisher-loc>Paris, France</publisher-loc>: <publisher-name>Atlantis Press</publisher-name>; <year>2018</year>. p. <fpage>143</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.2991/mecae-18.2018.32</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>MK</given-names></string-name>, <string-name><surname>Chin</surname> <given-names>RKY</given-names></string-name>, <string-name><surname>Chua</surname> <given-names>BL</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Teo</surname> <given-names>KTK</given-names></string-name></person-group>. <article-title>Engine fault diagnosis using probabilistic neural network</article-title>. In: <conf-name>Proceedings of the 2021 IEEE International Conference on Artificial Intelligence in Engineering and Technology (IICAIET); 2021 Sep 13&#x2013;15</conf-name>; <publisher-loc>Kota Kinabalu, Malaysia</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iicaiet51634.2021.9573654</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Fault diagnosis of new energy vehicles based on PSO-IBP neural network</article-title>. <source>Int J Low Carbon Technol</source>. <year>2025</year>;<volume>20</volume>:<fpage>1104</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1093/ijlct/ctaf052</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hoxha</surname> <given-names>J</given-names></string-name>, <string-name><surname>&#x00C7;odur</surname> <given-names>MY</given-names></string-name>, <string-name><surname>Mustafaraj</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kanj</surname> <given-names>H</given-names></string-name>, <string-name><surname>El Masri</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Prediction of transportation energy demand in T&#x00FC;rkiye using stacking ensemble models: methodology and comparative analysis</article-title>. <source>Appl Energy</source>. <year>2023</year>;<volume>350</volume>(<issue>3</issue>):<fpage>121765</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.apenergy.2023.121765</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Engine fault diagnosis based on the weighted DS evidence theory</article-title>. In: <conf-name>2014 IEEE 7th International Workshop on Computational Intelligence and Applications (IWCIA); 2014 Nov 7&#x2013;8</conf-name>; <publisher-loc>Hiroshima, Japan</publisher-loc>. p. <fpage>219</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/IWCIA.2014.6988110</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nixon</surname> <given-names>S</given-names></string-name>, <string-name><surname>Weichel</surname> <given-names>R</given-names></string-name>, <string-name><surname>Reichard</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kozlowski</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Machine learning approach to diesel engine health prognostics using engine controller data</article-title>. <source>Annu Conf PHM Soc</source>. <year>2018</year>;<volume>10</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.36001/phmconf.2018.v10i1.587</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Baik</surname> <given-names>N</given-names></string-name>, <string-name><surname>Rew</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Multi-scale graph neural network for multivariate time-series anomaly detection</article-title>. <source>J Korea Soc Ind Inf Syst</source>. <year>2025</year>;<volume>30</volume>(<issue>1</issue>):<fpage>51</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.9723/jksiis.2025.30.1.051</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>FSAMLM: a few-shot adaptation multimodal large model for cross-domain fault diagnosis</article-title>. <source>Appl Soft Comput</source>. <year>2025</year>;<volume>185</volume>:<fpage>113985</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2025.113985</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>You</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Channel-adaptive generative reconstruction and fusion for multi-sensor graph features in few-shot fault diagnosis</article-title>. <source>Inf Fusion</source>. <year>2026</year>;<volume>127</volume>:<fpage>103742</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2025.103742</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hassija</surname> <given-names>V</given-names></string-name>, <string-name><surname>Chamola</surname> <given-names>V</given-names></string-name>, <string-name><surname>Mahapatra</surname> <given-names>A</given-names></string-name>, <string-name><surname>Singal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goel</surname> <given-names>D</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Interpreting black-box models: a review on explainable artificial intelligence</article-title>. <source>Cogn Comput</source>. <year>2024</year>;<volume>16</volume>(<issue>1</issue>):<fpage>45</fpage>&#x2013;<lpage>74</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s12559-023-10179-8</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Parimbelli</surname> <given-names>E</given-names></string-name>, <string-name><surname>Buonocore</surname> <given-names>TM</given-names></string-name>, <string-name><surname>Nicora</surname> <given-names>G</given-names></string-name>, <string-name><surname>Michalowski</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wilk</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bellazzi</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Why did AI get this one wrong?&#x2014;Tree-based explanations of machine learning model predictions</article-title>. <source>Artif Intell Med</source>. <year>2023</year>;<volume>135</volume>:<fpage>102471</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.artmed.2022.102471</pub-id>; <pub-id pub-id-type="pmid">36628785</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mersha</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lam</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wood</surname> <given-names>J</given-names></string-name>, <string-name><surname>AlShami</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Kalita</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Explainable artificial intelligence: a survey of needs, techniques, applications, and future direction</article-title>. <source>Neuro Comput</source>. <year>2024</year>;<volume>599</volume>:<fpage>128111</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2024.128111</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Saranya</surname> <given-names>A</given-names></string-name>, <string-name><surname>Subhashini</surname> <given-names>R</given-names></string-name></person-group>. <article-title>A systematic review of Explainable Artificial Intelligence models and applications: recent developments and future trends</article-title>. <source>Decis Anal J</source>. <year>2023</year>;<volume>7</volume>:<fpage>100230</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dajour.2023.100230</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ribeiro</surname> <given-names>M</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guestrin</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Why should I trust you?: explaining the predictions of any classifier</article-title>. In: <conf-name>Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations; 2016 Jun 12&#x2013;16</conf-name>; <publisher-loc>San Diego, CA, USA</publisher-loc>. p. <fpage>97</fpage>&#x2013;<lpage>101</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/n16-3020</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lundberg</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>SI</given-names></string-name></person-group>. <article-title>A unified approach to interpreting model predictions</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2017</year>;<volume>2</volume>:<fpage>4766</fpage>&#x2013;<lpage>75</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Muschalik</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fumagalli</surname> <given-names>F</given-names></string-name>, <string-name><surname>Jagtani</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hammer</surname> <given-names>B</given-names></string-name>, <string-name><surname>H&#x00FC;llermeier</surname> <given-names>E</given-names></string-name></person-group>. <chapter-title>iPDP: on partial dependence plots in dynamic modeling scenarios</chapter-title>. In: <source>Explainable artificial intelligence</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer Nature</publisher-name>; <year>2023</year>. p. <fpage>177</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-44064-9_11</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ca&#x00E7;&#x00E3;o</surname> <given-names>J</given-names></string-name>, <string-name><surname>Santos</surname> <given-names>J</given-names></string-name>, <string-name><surname>Antunes</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Explainable AI for industrial fault diagnosis: a systematic review</article-title>. <source>J Ind Inf Integr</source>. <year>2025</year>;<volume>47</volume>:<fpage>100905</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jii.2025.100905</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Leemann</surname> <given-names>T</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>TT</given-names></string-name>, <string-name><surname>Fiedler</surname> <given-names>L</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>P</given-names></string-name>, <string-name><surname>Unhelkar</surname> <given-names>V</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Towards human-centered explainable AI: a survey of user studies for model explanations</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2024</year>;<volume>46</volume>(<issue>4</issue>):<fpage>2104</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2023.3331846</pub-id>; <pub-id pub-id-type="pmid">37956008</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>dos Santos</surname> <given-names>PC</given-names></string-name>, <string-name><surname>Rocha</surname> <given-names>MB</given-names></string-name>, <string-name><surname>Krohling</surname> <given-names>RA</given-names></string-name></person-group>. <article-title>Combining SHAP and causal analysis for interpretable fault detection in industrial processes</article-title>. <comment>arXiv:2510.23817</comment>. <year>2025</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zereen</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Das</surname> <given-names>A</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Machine fault diagnosis using audio sensors data and explainable AI techniques-LIME and SHAP</article-title>. <source>Comput Mater Contin</source>. <year>2024</year>;<volume>80</volume>(<issue>3</issue>):<fpage>3463</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2024.054886</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lundberg</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mowla</surname> <given-names>NI</given-names></string-name>, <string-name><surname>Abedin</surname> <given-names>SF</given-names></string-name>, <string-name><surname>Thar</surname> <given-names>K</given-names></string-name>, <string-name><surname>Mahmood</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gidlund</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Experimental analysis of trustworthy in-vehicle intrusion detection system using eXplainable artificial intelligence (XAI)</article-title>. <source>IEEE Access</source>. <year>2022</year>;<volume>10</volume>:<fpage>102831</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2022.3208573</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Vision-language models for vision tasks: a survey</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2024</year>;<volume>46</volume>(<issue>8</issue>):<fpage>5625</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2024.3369699</pub-id>; <pub-id pub-id-type="pmid">38408000</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ghosh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Acharya</surname> <given-names>A</given-names></string-name>, <string-name><surname>Saha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>V</given-names></string-name>, <string-name><surname>Chadha</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Exploring the frontier of vision-language models: a survey of current methodologies and future directions</article-title>. <comment>arXiv:2404.07214. 2024</comment>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rew</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Memory-augmented large language model for enhanced chatbot services in university learning management systems</article-title>. <source>Appl Sci</source>. <year>2025</year>;<volume>15</volume>(<issue>17</issue>):<fpage>9775</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app15179775</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rew</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Automated change history analysis of software bill of materials(SBOM) using large language model-based structural analysis and forgery detection agents</article-title>. <source>J Korea Soc Ind Inf Syst</source>. <year>2025</year>;<volume>30</volume>(<issue>4</issue>):<fpage>39</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.9723/jksiis.2025.30.4.039</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Klissarov</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hjelm</surname> <given-names>D</given-names></string-name>, <string-name><surname>Toshev</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mazoure</surname> <given-names>B</given-names></string-name></person-group>. <article-title>On the modeling capabilities of large language models for sequential decision making</article-title>. <comment>arXiv:2410.05656. 2024</comment>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guestrin</surname> <given-names>C</given-names></string-name></person-group>. <article-title>XGBoost: a scalable tree boosting system</article-title>. In: <conf-name>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016 Aug 13&#x2013;17</conf-name>; <publisher-loc>San Francisco, CA, USA</publisher-loc>. p. <fpage>785</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Lyu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Data-driven fault early warning model of automobile engines based on soft classification</article-title>. <source>Electronics</source>. <year>2023</year>;<volume>12</volume>(<issue>3</issue>):<fpage>511</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics12030511</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hasan</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Sohaib</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>An explainable AI-based fault diagnosis model for bearings</article-title>. <source>Sensors</source>. <year>2021</year>;<volume>21</volume>(<issue>12</issue>):<fpage>4070</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s21124070</pub-id>; <pub-id pub-id-type="pmid">34199163</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Pilario</surname> <given-names>KES</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>N</given-names></string-name>, <string-name><surname>Moon</surname> <given-names>I</given-names></string-name>, <string-name><surname>Na</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Explainable artificial intelligence for fault diagnosis of industrial processes</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2025</year>;<volume>21</volume>(<issue>1</issue>):<fpage>4</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2023.3240601</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Roberts</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wong</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Yasunaga</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Image2Struct: benchmarking structure extraction for vision-language models</article-title>. <comment>arXiv:2410.22456</comment>. <year>2024</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xia</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Vision language models for spreadsheet understanding: challenges and opportunities</article-title>. <comment>arXiv:2405.16234</comment>. <year>2024</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kuang</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Ocean-OCR: towards general OCR application via a vision-language model</article-title>. <volume>arXiv</volume>:<fpage>2501</fpage>.<lpage>15558</lpage>. <year>2025</year>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shu</surname> <given-names>K</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Can large language model agents simulate human trust behaviors?</article-title> <comment>arXiv:2402.04559. 2024</comment>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name></person-group>. <article-title>ChatLog: recording and analyzing ChatGPT across time</article-title>. <comment>arXiv:2304.14106</comment>. <year>2023</year>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sui</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>M</given-names></string-name>, <string-name><surname>Han</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Table meets LLM: can large language models understand structured table data? a benchmark and empirical study</article-title>. In: <conf-name>Proceedings of the 17th ACM International Conference on Web Search and Data Mining; 2024 Mar 4&#x2013;8</conf-name>; <publisher-loc>Merida, Mexico</publisher-loc>. p. <fpage>645</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3616855.3635752</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Trizoglou</surname> <given-names>P</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Fault detection by an ensemble framework of Extreme Gradient Boosting (XGBoost) in the operation of offshore wind turbines</article-title>. <source>Renew Energy</source>. <year>2021</year>;<volume>179</volume>:<fpage>945</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.renene.2021.07.085</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Fault diagnosis of vibration sensors based on triage loss function-improved XGBoost</article-title>. <source>Electronics</source>. <year>2023</year>;<volume>12</volume>(<issue>21</issue>):<fpage>4442</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics12214442</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Intelligent fault diagnosis of diesel engines <italic>via</italic> extreme gradient boosting and high-accuracy time-frequency information of vibration signals</article-title>. <source>Sensors</source>. <year>2019</year>;<volume>19</volume>(<issue>15</issue>):<fpage>3280</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s19153280</pub-id>; <pub-id pub-id-type="pmid">31349707</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Raghuvira</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Panda</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sameera</surname> <given-names>GS</given-names></string-name></person-group>. <article-title>Predictive maintenance for two-wheeler vehicles using XGBoost</article-title>. In: <conf-name>Proceedings of the 2024 10th International Conference on Advanced Computing and Communication Systems (ICACCS); 2024 Mar 14&#x2013;15</conf-name>; <publisher-loc>Coimbatore, India</publisher-loc>. p. <fpage>746</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICACCS60874.2024.10717187</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Beirami</surname> <given-names>A</given-names></string-name>, <string-name><surname>He</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A systematic survey of prompt engineering on vision-language foundation models</article-title>. <comment>arXiv:2307.12980</comment>. <year>2023</year>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vergara</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ramos</surname> <given-names>L</given-names></string-name>, <string-name><surname>Rivera-Campoverde</surname> <given-names>ND</given-names></string-name>, <string-name><surname>Rivas-Echeverr&#x00ED;a</surname> <given-names>F</given-names></string-name></person-group>. <article-title>EngineFaultDB: a novel dataset for automotive engine fault classification and baseline results</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>126155</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2023.3331316</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gurram</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>Studies on spark ignition engine&#x2014;a review</article-title>. <source>Int J Therm Technol</source>. <year>2016</year>;<volume>6</volume>(<issue>3</issue>):<fpage>266</fpage>&#x2013;<lpage>72</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Patro</surname> <given-names>SGK</given-names></string-name>, <string-name><surname>Sahu</surname> <given-names>KK</given-names></string-name></person-group>. <article-title>Normalization: a preprocessing stage</article-title>. <comment>arXiv:1503.06462. 2015</comment>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Ingersoll</surname> <given-names>GM</given-names></string-name></person-group>. <article-title>An introduction to logistic regression analysis and reporting</article-title>. <source>J Educ Res</source>. <year>2002</year>;<volume>96</volume>(<issue>1</issue>):<fpage>3</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.1080/00220670209598786</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Cunningham</surname> <given-names>P</given-names></string-name>, <string-name><surname>Delany</surname> <given-names>SJ</given-names></string-name></person-group>. <source>k-nearest neighbour classifiers</source>: <edition>2nd edition (with Python examples)</edition>; <year>2020</year>. Vol. <volume>1</volume>, p. <fpage>1</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3459665</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Marius-Constantin</surname> <given-names>P</given-names></string-name>, <string-name><surname>Balas</surname> <given-names>VE</given-names></string-name>, <string-name><surname>Perescu-Popescu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Mastorakis</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Multilayer perceptron and neural networks</article-title>. <source>WSEAS Trans Circuits Syst</source>. <year>2009</year>;<volume>8</volume>(<issue>7</issue>):<fpage>579</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.1002/9781394268993.ch3</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Louppe</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Understanding random forests: from theory to practice</article-title>. <comment>arXiv:1407.7502. 2014</comment>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Quinlan</surname> <given-names>JR</given-names></string-name></person-group>. <article-title>Induction of decision trees</article-title>. <source>Mach Learn</source>. <year>1986</year>;<volume>1</volume>(<issue>1</issue>):<fpage>81</fpage>&#x2013;<lpage>106</lpage>. doi:<pub-id pub-id-type="doi">10.1007/BF00116251</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Geurts</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ernst</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wehenkel</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Extremely randomized trees</article-title>. <source>Mach Learn</source>. <year>2006</year>;<volume>63</volume>(<issue>1</issue>):<fpage>3</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10994-006-6226-1</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Beja-Battais</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Overview of AdaBoost: reconciling its views to better understand its dynamics</article-title>. <comment>arXiv:2310.18323. 2023</comment>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tharwat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gaber</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ibrahim</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hassanien</surname> <given-names>AE</given-names></string-name></person-group>. <article-title>Linear discriminant analysis: a detailed tutorial</article-title>. <source>AI Commun</source>. <year>2017</year>;<volume>30</volume>(<issue>2</issue>):<fpage>169</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.3233/aic-170729</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ghojogh</surname> <given-names>B</given-names></string-name>, <string-name><surname>Crowley</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Linear and quadratic discriminant analysis: tutorial</article-title>. <comment>arXiv:1906.02590. 2019</comment>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lau</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Gradient boosting machine: a survey</article-title>. <comment>arXiv:1908.06951. 2019</comment>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fault diagnosis method for high-voltage direct current transmission system based on multimodal sensor feature-LightGBM algorithm: a case study in China</article-title>. <source>Energies</source>. <year>2025</year>;<volume>18</volume>(<issue>23</issue>):<fpage>6253</fpage>. doi:<pub-id pub-id-type="doi">10.3390/en18236253</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prokhorenkova</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gusev</surname> <given-names>G</given-names></string-name>, <string-name><surname>Vorobev</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dorogush</surname> <given-names>AV</given-names></string-name>, <string-name><surname>Gulin</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Catboost: unbiased boosting with categorical features</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2018</year>;<volume>2018</volume>:<fpage>6638</fpage>&#x2013;<lpage>48</lpage>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>OpenAI</collab>, <string-name><surname>Achiam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Adler</surname> <given-names>S</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>L</given-names></string-name>, <string-name><surname>Akkaya</surname> <given-names>I</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GPT-4 technical report</article-title>. <comment>arXiv:2303.08774. 2023</comment>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vujovic</surname> <given-names>&#x017D;&#x00D0;</given-names></string-name></person-group>. <article-title>Classification model evaluation metrics</article-title>. <source>Int J Adv Comput Sci Appl</source>. <year>2021</year>;<volume>12</volume>(<issue>6</issue>):<fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.14569/ijacsa.2021.0120670</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name>, <string-name><surname>Park</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>G</given-names></string-name>, <string-name><surname>Seo</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Prometheus-vision: vision-language model as a judge for fine-grained evaluation</article-title>. In: <conf-name>Proceedings of the Findings of the Association for Computational Linguistics ACL 2024</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2024</year>. p. <fpage>11286</fpage>&#x2013;<lpage>315</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2024.findings-acl.672</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Grattafiori</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dubey</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jauhri</surname> <given-names>A</given-names></string-name>, <string-name><surname>Pandey</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kadian</surname> <given-names>A</given-names></string-name>, <string-name><surname>Al-Dahle</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The Llama 3 herd of models</article-title>. <comment>arXiv:2407.21783</comment>. <year>2024</year>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Joshi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kale</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chandel</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pal</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Likert scale: explored and explained</article-title>. <source>Br J Appl Sci Technol</source>. <year>2015</year>;<volume>7</volume>(<issue>4</issue>):<fpage>396</fpage>&#x2013;<lpage>403</lpage>. doi:<pub-id pub-id-type="doi">10.9734/bjast/2015/14975</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Devireddy</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A comparative study of explainable AI methods: model-agnostic vs. model-specific approaches</article-title>. <comment>arXiv:2504.04276. 2025</comment>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Feature selections based on two-type overlap degrees and three-view granulation measures for k-nearest-neighbor rough sets</article-title>. <source>Pattern Recognit</source>. <year>2024</year>;<volume>156</volume>:<fpage>110837</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2024.110837</pub-id>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Shepherd</surname> <given-names>BE</given-names></string-name></person-group>. <article-title>Between- and within-cluster spearman rank correlations</article-title>. <source>Stat Med</source>. <year>2025</year>;<volume>44</volume>(<issue>3&#x2013;4</issue>):<fpage>e10326</fpage>. doi:<pub-id pub-id-type="doi">10.1002/sim.10326</pub-id>; <pub-id pub-id-type="pmid">39853810</pub-id></mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Barron</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Interpreting and improving optimal control problems with directional corrections</article-title>. <source>IEEE Robot Autom Lett</source>. <year>2025</year>;<volume>10</volume>(<issue>5</issue>):<fpage>4986</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1109/LRA.2025.3557226</pub-id>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Coeurjolly</surname> <given-names>JF</given-names></string-name>, <string-name><surname>Drouilhet</surname> <given-names>R</given-names></string-name>, <string-name><surname>Robineau</surname> <given-names>JF</given-names></string-name></person-group>. <article-title>Normalized information-based divergences</article-title>. <source>Probl Inf Transm</source>. <year>2007</year>;<volume>43</volume>(<issue>3</issue>):<fpage>167</fpage>&#x2013;<lpage>89</lpage>. doi:<pub-id pub-id-type="doi">10.1134/s0032946007030015</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions</article-title>. <source>ACM Trans Inf Syst</source>. <year>2025</year>;<volume>43</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3703155</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on hallucination in large vision-language models</article-title>. <comment>arXiv:2402.00253</comment>. <year>2024</year>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>J</given-names></string-name>, <string-name><surname>Buntine</surname> <given-names>W</given-names></string-name>, <string-name><surname>Shareghi</surname> <given-names>E</given-names></string-name></person-group>. <chapter-title>VerifiAgent: a unified verification agent in language model reasoning</chapter-title>. In: <source>Findings of the association for computational linguistics: EMNLP 2025</source>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2025</year>. p. <fpage>16410</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2025.findings-emnlp.891</pub-id>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zielonka</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sikora</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wo&#x017A;niak</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Fuzzy rules intelligent car real-time diagnostic system</article-title>. <source>Eng Appl Artif Intell</source>. <year>2024</year>;<volume>135</volume>:<fpage>108648</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2024.108648</pub-id>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rew</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Vision-language model-based local interpretable model-agnostic explanations analysis for explainable in-vehicle controller area network intrusion detection</article-title>. <source>Sensors</source>. <year>2025</year>;<volume>25</volume>(<issue>10</issue>):<fpage>3020</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s25103020</pub-id>; <pub-id pub-id-type="pmid">40431814</pub-id></mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chodnekar</surname> <given-names>V</given-names></string-name>, <string-name><surname>Roy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A cost-benefit analysis of on-premise large language model deployment: breaking even with commercial LLM services</article-title>. <comment>arXiv:2509.18101</comment>. <year>2025</year>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mokhtarian</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kampmann</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lueer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kowalewski</surname> <given-names>S</given-names></string-name>, <string-name><surname>Alrifaee</surname> <given-names>B</given-names></string-name></person-group>. <article-title>A cloud architecture for networked and autonomous vehicles</article-title>. <source>IFAC PapersOnLine</source>. <year>2021</year>;<volume>54</volume>(<issue>2</issue>):<fpage>233</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ifacol.2021.06.028</pub-id>.</mixed-citation></ref>
<ref id="ref-77"><label>[77]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Haque</surname> <given-names>ME</given-names></string-name>, <string-name><surname>Zabin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>J</given-names></string-name></person-group>. <article-title>EnsembleXAI-motor: a lightweight framework for fault classification in electric vehicle drive motors using feature selection, ensemble learning, and explainable AI</article-title>. <source>Machines</source>. <year>2025</year>;<volume>13</volume>(<issue>4</issue>):<fpage>314</fpage>. doi:<pub-id pub-id-type="doi">10.3390/machines13040314</pub-id>.</mixed-citation></ref>
<ref id="ref-78"><label>[78]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Empowering edge intelligence: a comprehensive survey on on-device AI models</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>57</volume>(<issue>9</issue>):<fpage>1</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3724420</pub-id>.</mixed-citation></ref>
<ref id="ref-79"><label>[79]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Goldstein</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kapelner</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bleich</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pitkin</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Peeking inside the black box: visualizing statistical learning with plots of individual conditional expectation</article-title>. <source>J Comput Graph Stat</source>. <year>2015</year>;<volume>24</volume>(<issue>1</issue>):<fpage>44</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1080/10618600.2014.907095</pub-id>.</mixed-citation></ref>
<ref id="ref-80"><label>[80]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Du</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Nghiem</surname> <given-names>H</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A survey of state of the art large vision language models: alignment, benchmark, evaluations and challenges</article-title>. In: <conf-name>Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 2025 Jun 11&#x2013;12</conf-name>; <publisher-loc>Nashville, TN, USA</publisher-loc>. p. <fpage>1578</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPRW67362.2025.00147</pub-id>.</mixed-citation></ref>
<ref id="ref-81"><label>[81]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Vision language models: a survey of 26K papers</article-title>. <comment>arXiv:2510.09586</comment>. <year>2025</year>.</mixed-citation></ref>
<ref id="ref-82"><label>[82]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hua</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A survey of vehicle system and energy models</article-title>. <source>Actuators</source>. <year>2025</year>;<volume>14</volume>(<issue>1</issue>):<fpage>10</fpage>. doi:<pub-id pub-id-type="doi">10.3390/act14010010</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>