<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">82804</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.082804</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Enhancing Power Enterprise Inspection and Supervision: A LoRA-Based Lightweight LLM Framework Integrating Retrieval-Augmented Generation and Prompt Engineering</article-title>
<alt-title alt-title-type="left-running-head">Enhancing Power Enterprise Inspection and Supervision: A LoRA-Based Lightweight LLM Framework Integrating Retrieval-Augmented Generation and Prompt Engineering</alt-title>
<alt-title alt-title-type="right-running-head">Enhancing Power Enterprise Inspection and Supervision: A LoRA-Based Lightweight LLM Framework Integrating Retrieval-Augmented Generation and Prompt Engineering</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Jianfeng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Yongjiao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Kangyi</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Hu</surname><given-names>Changhua</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Xu</surname><given-names>Zijia</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Shi</surname><given-names>Qingguo</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-7" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Su</surname><given-names>Yi</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>suyi2018@xtu.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>Guangdong Power Grid Co., Ltd.</institution>, <addr-line>Zhongshan</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Faculty of Automation and Electronic Information, Xiangtan University</institution>, <addr-line>Xiangtan</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yi Su. Email: <email>suyi2018@xtu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>95</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>05</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_82804.pdf"></self-uri>
<abstract>
<p>Power enterprise inspection and supervision require greater intelligence, efficiency, and standardization; however, existing approaches are limited by inefficient knowledge retrieval, inaccurate issue identification, and insufficient support for standardized reporting and rectification tracking. This study proposes a lightweight, domain-adaptive large language model (LLM) framework based on Low-Rank Adaptation (LoRA), integrating Retrieval-Augmented Generation (RAG) and structured prompt engineering to enable evidence-grounded inspection tasks. The framework achieves parameter-efficient adaptation through low-rank decomposition and constructs a domain-specific multimodal knowledge base, enhancing output traceability, consistency, and task generalization. A key contribution is the introduction of a Sensitive Information Control Gate, which enforces role-based access control and automated redaction, ensuring secure and compliant generation in regulated environments while preserving traceability. Experimental results demonstrate that the proposed method achieves improved performance over the base model and demonstrates competitive effectiveness under the evaluated conditions, supported by statistical analysis (paired <italic>t</italic>-test, <italic>p</italic> &#x003C; 0.01, bootstrap 95% confidence intervals), while maintaining high parameter efficiency with only 0.4%&#x2013;0.5% trainable parameters.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Large language models</kwd>
<kwd>LoRA fine-tuning</kwd>
<kwd>retrieval-augmented generation</kwd>
<kwd>prompt engineering</kwd>
<kwd>inspection and supervision</kwd>
<kwd>power enterprise governance</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Intelligent Assistant for Inspection and Supervision</funding-source>
<award-id>0375002025030102PT00034</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<sec id="s1_1">
<label>1.1</label>
<title>Background and Motivation</title>
<p>With the advancement of governance modernization, inspection and supervision have evolved from manual processes to data-driven and intelligent paradigms. In this context, integrating large language models (LLMs) into inspection workflows offers significant potential for improving policy retrieval, semantic understanding, issue identification, and report generation.</p>
<p>However, the direct deployment of general-purpose LLMs remains constrained by high computational cost, limited domain adaptation, and insufficient reliability in regulated environments. Meanwhile, existing inspection systems suffer from inefficient knowledge retrieval, inconsistent standards, heterogeneous multimodal data, and underutilization of historical information, which hinder accurate issue identification and effective rectification tracking.</p>
<p>Therefore, there is a pressing need for lightweight, domain-adaptive, and trustworthy LLM frameworks to support efficient, standardized, and compliant inspection and supervision processes.</p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Research Objectives</title>
<p>This study aims to develop a lightweight and deployable LLM framework tailored to power enterprise inspection and supervision. The key objectives are as follows:<list list-type="order">
<list-item>
<p>Efficient domain adaptation: leverage LoRA to achieve parameter-efficient fine-tuning, significantly reducing computational and storage requirements while maintaining performance.</p></list-item>
<list-item>
<p>Evidence-grounded reasoning: integrate RAG with structured prompt engineering to ensure that generated outputs are consistent, traceable, and aligned with regulatory documents.</p></list-item>
<list-item>
<p>Security-aware generation: introduce a Sensitive Information Control Gate to enforce role-based access control and prevent unauthorized disclosure in regulated environments.</p></list-item>
</list></p>
<p>Although the framework builds upon established components (LoRA, RAG, OCR, and prompting), its primary contribution lies in their unified integration into a lightweight, multimodal, and security-aware pipeline, specifically designed for real-world inspection scenarios.</p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>Current Research Status and Literature Review</title>
<p>Recent advances in large language models (LLMs) have enabled new applications in governance, compliance auditing, and enterprise supervision [<xref ref-type="bibr" rid="ref-1">1</xref>]. Pre-trained models such as GPT, PaLM, and LLaMA demonstrate strong capabilities in natural language understanding and generation [<xref ref-type="bibr" rid="ref-2">2</xref>], and achieve effective performance in tasks such as question answering and policy interpretation when combined with retrieval-augmented generation (RAG) and fine-tuning techniques [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>However, their deployment in industrial settings remains constrained by high computational cost and limited adaptability to domain-specific and regulated scenarios. To address efficiency constraints, parameter-efficient fine-tuning (PEFT) methods, including Adapter, Prefix Tuning, and LoRA, have been proposed [<xref ref-type="bibr" rid="ref-4">4</xref>]. Among these, LoRA reduces trainable parameters via low-rank decomposition while preserving performance [<xref ref-type="bibr" rid="ref-5">5</xref>], and has shown effectiveness in several application domains [<xref ref-type="bibr" rid="ref-6">6</xref>]. Nevertheless, existing studies largely overlook requirements such as traceability, consistency, and compliance in regulated environments [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>LLMs have also been applied to domain-specific tasks such as legal analysis, medical document generation, and smart grid forecasting [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. While these approaches demonstrate strong task performance [<xref ref-type="bibr" rid="ref-5">5</xref>], they typically rely on single-modal inputs [<xref ref-type="bibr" rid="ref-6">6</xref>] and weakly constrained generation, limiting their suitability for inspection scenarios that require multimodal integration, standardized outputs, and evidence-grounded reasoning [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Overall, current research lacks a lightweight, deployable, and compliance-aware LLM framework for inspection and supervision. To address this gap, this study proposes a unified framework integrating LoRA, RAG, and structured prompting, together with a Sensitive Information Control Gate to ensure trustworthy and compliant generation in power enterprise inspection.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Technologies</title>
<p>To support the design of a deployable LLM framework for inspection and supervision, this section reviews four key technologies: large language models in inspection tasks, parameter-efficient fine-tuning, retrieval-augmented generation (RAG), and prompt engineering. These components jointly address efficiency, factual reliability, and task-specific controllability in regulated environments.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Application of Large Language Models in Inspection and Supervision</title>
<p>Large language models (LLMs) have been increasingly applied to governance-related tasks such as policy analysis, compliance monitoring, and audit assistance [<xref ref-type="bibr" rid="ref-11">11</xref>], demonstrating strong capabilities in information extraction, question answering, and document generation [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>However, their direct application to inspection scenarios remains limited [<xref ref-type="bibr" rid="ref-13">13</xref>]. Inspection data are typically multimodal and highly domain-specific, including scanned documents, tables, and handwritten records. Moreover, inspection tasks require high factual reliability and auditability, while generative models are prone to hallucinations [<xref ref-type="bibr" rid="ref-14">14</xref>]. These challenges necessitate domain-adaptive and evidence-grounded frameworks for practical deployment.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Parameter-Efficient Fine-Tuning Techniques</title>
<p>As LLMs scale, full-parameter fine-tuning becomes computationally prohibitive. Parameter-efficient fine-tuning (PEFT) methods&#x2014;such as adapters, prefix tuning, and LoRA&#x2014;provide practical alternatives [<xref ref-type="bibr" rid="ref-15">15</xref>]. Among them, LoRA introduces low-rank decomposition to update a small subset of parameters while keeping the pretrained model frozen [<xref ref-type="bibr" rid="ref-14">14</xref>], significantly reducing memory and training cost. This makes LoRA particularly suitable for resource-constrained deployment scenarios.</p>
<p>However, PEFT alone does not address the need for evidence grounding and task-level consistency in inspection workflows, requiring integration with retrieval and prompting mechanisms.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Retrieval-Augmented Generation</title>
<p>Retrieval-augmented generation (RAG) enhances LLM performance by incorporating external knowledge during inference, improving factual accuracy and reducing hallucinations [<xref ref-type="bibr" rid="ref-15">15</xref>]. By grounding responses in retrieved evidence, RAG enables more reliable and verifiable outputs in domain-specific tasks [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p>Nevertheless, RAG introduces new challenges, including document segmentation, retrieval consistency, and evidence management, especially when dealing with multimodal inspection data. These limitations highlight the need for structured and controllable retrieval mechanisms.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Prompt Engineering</title>
<p>Prompt engineering guides LLM behavior through structured instructions [<xref ref-type="bibr" rid="ref-18">18</xref>], including role assignment, templates, and reasoning strategies such as chain-of-thought [<xref ref-type="bibr" rid="ref-19">19</xref>]. While explicit prompts improve interpretability, they may be sensitive to wording and less stable in domain-specific tasks [<xref ref-type="bibr" rid="ref-20">20</xref>]; implicit methods offer robustness but lack transparency [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
<p>In inspection scenarios, prompts must ensure consistency, traceability, and alignment with retrieved evidence. Therefore, hybrid strategies that combine structured templates, role constraints, and evidence grounding are essential for reliable multi-task execution [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>System Framework</title>
<p>To bridge the gap between general LLM capabilities and inspection-specific requirements, we design a unified framework integrating LoRA-based adaptation, retrieval-augmented generation (RAG), and structured prompting.</p>
<p>As illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the system follows a three-layer architecture:</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Schematic diagram of the domain-specific inspection LLM framework.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-1.tif"/>
</fig>
<p><list list-type="bullet">
<list-item>
<p>Interaction layer: handles user queries and returns structured outputs;</p></list-item>
<list-item>
<p>Model layer: integrates LoRA, RAG, and prompt engineering for domain-adaptive reasoning;</p></list-item>
<list-item>
<p>Data layer: maintains multimodal knowledge sources and supports semantic retrieval.</p></list-item>
</list></p>
<p>The overall workflow can be summarized as: Input&#x2192; Preprocessing &#x2192; Retrieval &#x2192; Controlled Generation &#x2192; Output. This design ensures efficient deployment, evidence-grounded reasoning, and task-level controllability. Detailed implementations are presented in <xref ref-type="sec" rid="s4">Section 4</xref>.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>Methodology</title>
<sec id="s4_1">
<label>4.1</label>
<title>Parameter-Efficient Adaptation: LoRA Fine-Tuning</title>
<p>The workflow of LoRA-based adaptation is illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, which consists of three stages: data preparation, parameter-efficient fine-tuning, and model compression.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Illustration of the LoRA fine-tuning workflow for the inspection and supervision LLM.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-2.tif"/>
</fig>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Data Preprocessing</title>
<p>Multi-source inspection data were standardized to ensure input consistency and training quality. The preprocessing pipeline includes:<list list-type="simple">
<list-item>
<label>(1)</label>
<p>Text normalization: encoding correction and structured parsing of HTML/XML content;</p></list-item>
<list-item>
<label>(2)</label>
<p>Segmentation and entity normalization: sentence splitting and domain-specific entity standardization;</p></list-item>
<list-item>
<label>(3)</label>
<p>Context windowing: segmentation into 256&#x2013;512 token chunks;</p></list-item>
<list-item>
<label>(4)</label>
<p>Instruction construction: Alpaca-style triples (instruction, input, output);</p></list-item>
<list-item>
<label>(5)</label>
<p>Dataset integration: storage into an instruction-tuning database.</p></list-item>
</list></p>
<p>The final dataset contains 6872 samples derived from 1248 inspection documents, with an 80/10/10 split (seed &#x003D; 42). Annotation follows a semi-automatic &#x002B; expert validation protocol, achieving Cohen&#x2019;s &#x03BA; &#x003D; 0.87, indicating substantial agreement.</p>
<p>Due to regulatory constraints, the dataset is not publicly available.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>LoRA Fine-Tuning</title>
<p>To enable efficient domain adaptation, LoRA is applied to the base model (Qwen3-8B) by injecting low-rank updates into attention projection layers.</p>
<p>Given a pretrained weight matrix <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>out</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>, the adapted weight is defined as:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>LoRA</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>U</mml:mtext></mml:mrow><mml:msup><mml:mrow><mml:mtext>V</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <italic>r</italic> denotes the decomposition rank, &#x03B1; is a scaling factor.</p>
<p>The number of trainable parameters per layer is:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>LoRA</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>layer</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>out</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Given a Transformer with L layers and m injected modules per layer (here m &#x003D; 4 for <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the total additional trainable parameters can be expressed as:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>LoRA</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>out</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Under the configuration r &#x003D; 8 and scaling factor &#x03B1; &#x003D; 16, LoRA introduces only 0.4%&#x2013;0.5% additional parameters (&#x007E;30&#x2013;40M), enabling efficient adaptation under resource constraints.</p>
<p>The training dynamics are shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, where the loss decreases smoothly from &#x007E;2.0 to &#x007E;0.6, indicating stable convergence without overfitting.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Training loss curve across training steps (original and smoothed).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-3.tif"/>
</fig>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Model Quantization</title>
<p>To enable efficient deployment, post-training quantization is applied to compress model weights. A continuous weight <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mover><mml:mi>w</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is quantized as:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mover><mml:mi>w</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>round</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> denote the minimum and maximum values of the weights, b is the bit-width (e.g., <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mtext>b</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mn>8</mml:mn></mml:math></inline-formula>), and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow></mml:math></inline-formula> represents the quantization step size. The reconstructed weight during inference can be obtained as:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msup><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mover><mml:mi>w</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Higher-bit quantization (larger b) results in lower quantization error and output quality closer to the original model but increases disk usage and computational demand.</p>
<p>Through this workflow, LoRA fine-tuning and model quantization were successfully combined, ensuring high model performance while optimizing storage and computational resources, thereby providing an efficient and deployable intelligent solution for inspection and supervision tasks.</p>
<p>After fine-tuning, the output indicated that the pretrained model contains approximately 8 billion parameters, while only 30&#x2013;40 million parameters were updated during training, corresponding to roughly 0.4%&#x2013;0.5% of the total model parameters.</p>
<p>Subsequently, the LoRA weights were merged using the convert_hf_to_gguf.py script provided by llama.cpp. After merging, if the model is compressed using the q8_0 quantization format, the storage size can be reduced by approximately 50%, saving more than 8 GB of disk space compared with the unquantized model.</p>
<p>The impact of different quantization levels on inference speed, generation quality, and factual consistency is further analyzed through pilot ablation experiments in the newly added <xref ref-type="sec" rid="s5_9">Section 5.9</xref>.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Retrieval-Augmented Generation (RAG) Mechanism</title>
<p>The workflow of the RAG module is illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, consisting of data processing, knowledge construction, and optimized retrieval. RAG enhances generation quality by grounding model outputs in externally retrieved evidence, thereby improving factual accuracy and reducing hallucinations in domain-specific inspection tasks.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Workflow of the retrieval-augmented generation (RAG) module.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-4.tif"/>
</fig>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Data Processing</title>
<p>Heterogeneous inspection data (e.g., policy documents, reports, and rectification records) are first standardized into structured text. OCR is applied to extract content from scanned and image-based documents.</p>
<p>The processed text is cleaned and segmented into semantically coherent chunks (&#x2248;600 tokens), forming the basis for subsequent embedding and retrieval.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Knowledge Base Construction</title>
<p>Each text segment <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is encoded into a semantic vector:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>f</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>embed</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mtext>f</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>embed</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the embedding function. The resulting vectors are indexed in a FAISS-based database for efficient retrieval.</p>
<p>To support dynamic updates, an incremental indexing mechanism is adopted. For each document <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mrow><mml:mtext>D</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, a hash value is computed:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mrow><mml:mtext>h</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Hash</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>D</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Updates are detected via hash changes, and only modified segments <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mtext>D</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> are re-indexed, ensuring efficient maintenance of the knowledge base.</p>
</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Optimized Retrieval and Recall</title>
<p>The workflow of the optimized knowledge retrieval and recall method is illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> and follows a reproducible and measurable engineering process. At the embedding layer, the Chinese vector model Sentence Transformer (bge-small-zh-v1.5) projects text into a high-dimensional vector space. Given a user query q, the query embedding is:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>f</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>embed</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Schematic diagram of the optimized knowledge retrieval and recall process.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-5.tif"/>
</fig>
<p>FAISS serves as the nearest-neighbor search engine to support large-scale vector retrieval. For each candidate segment with embedding v<sub>i</sub>, similarity is measured by cosine similarity:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow><mml:mo>&#x2225;&#x2225;</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">&#x2225;</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>By default, the retrieval stage returns the top 5 candidate segments and applies a similarity threshold of &#x03C4; &#x003D; 0.4 to filter out low-relevance items, i.e.,
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2265;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03C4;</mml:mi></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></disp-formula></p>
<p>If the runtime environment does not support threshold-based retrieval, the system falls back to a score-based similarity search using the same threshold.</p>
<p>To ensure semantic integrity and retrieval consistency at the fragment level, documents are pre-segmented before indexing using a question-answer (QA)-oriented splitting strategy. The splitting window is set to 600 characters with an 80-character overlap, and segmentation prioritizes QA markers to preserve semantic units, thereby enhancing fragment retrieval quality.</p>
<p>After initial retrieval, candidate fragments are re-ranked based on question keywords to further improve relevance. Let <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> be the query keyword set and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> the keyword set of fragment <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>; the keyword-match ratio is:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mrow><mml:mtext>r</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2229;</mml:mo><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>q</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Only fragments with r<sub>i</sub> &#x2265; 0.2 relative to the question keyword set are included in the final context. The context provided to the generation module is constructed by concatenating up to three highly relevant fragments:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mtext>C</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Concat</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>t</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denote the top-ranked fragments after re-ranking. To reduce noise, irrelevant QA annotations are removed using rule-based normalization, increasing information density.</p>
<p>Index maintenance leverages file hashes and metadata tracking to support incremental updates. Newly added or modified files trigger incremental embedding generation only for the changed fragments <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mtext>D</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, which are then written into FAISS, enabling online updating and persistent storage.</p>
<p>A fallback mechanism is introduced to ensure robustness: when no segment satisfies the threshold, the system defaults to the base LLM to maintain response continuity.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Prompt Engineering and Task Customization</title>
<p>To ensure structured, consistent, and evidence-grounded outputs, we design a task-oriented prompt formulation built on top of the RAG framework. The prompt is formalized as a four-element tuple:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mi>&#x1D4AB;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mi>E</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:math></inline-formula> is role specification, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow></mml:math></inline-formula> is task instruction, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>E</mml:mi></mml:math></inline-formula> is retrieved evidence from RAG, <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow></mml:math></inline-formula> output format constraints.</p>
<p>Given a user query <italic>q</italic>, the model generates:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>L</mml:mtext></mml:mrow><mml:mi>L</mml:mi><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi>&#x1D4AB;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This formulation enforces role-aware reasoning, task consistency, and structured output generation.</p>
<p>To further reduce hallucination, explicit constraints are introduced:<list list-type="bullet">
<list-item>
<p>evidence-grounded reasoning (outputs must be supported by <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>E</mml:mi></mml:math></inline-formula>);</p></list-item>
<list-item>
<p>strict output schema (predefined structured format);</p></list-item>
<list-item>
<p>uncertainty handling (explicit &#x201C;insufficient evidence&#x201D; condition).</p></list-item>
</list></p>
<p>These constraints significantly improve factual consistency, as validated by the Faithfulness metric reported in <xref ref-type="sec" rid="s5_4">Section 5.4</xref>.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Sensitive Information Control Gate</title>
<p>To enforce security and regulatory compliance, we introduce a Sensitive Information Control Gate (SICG), positioned between the retrieval module and prompt construction.</p>
<p>Given a retrieved evidence chunk <italic>x</italic>, its sensitivity score is defined as:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> are non-negative weights satisfying <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> (calibrated by domain experts in accordance with enterprise security policies), and <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">chunk</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the individual sensitivity feature functions. These functions include normalized binary indicators for the presence of personal identifiers, financial or operational confidentiality markers, and policy-restricted keywords, each scaled to the interval [0, 1].</p>
<p>Each user role is assigned a clearance level <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>user</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, with higher values reflecting greater access privileges (e.g., 0.3 for field inspectors, 0.7 for supervisors, and 1.0 for administrators). The filtering decision for each chunk is governed by the following threshold rule:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:mtext>output</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>chunk</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>redacted</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>chunk</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>chunk</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003E;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>c</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>user</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>chunk</mml:mtext></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo><mml:mn>1.0</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is a tunable security threshold (default value <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:mi mathvariant="normal">&#x03B8;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.75</mml:mn></mml:math></inline-formula>) that balances protection and usability. The redaction operation replaces sensitive spans with the placeholder &#x201C;[REDACTED]&#x201D; while preserving contextual coherence and source traceability metadata.</p>
<p>By combining sensitivity scoring, role-based thresholds, and structured redaction, SICG effectively mitigates information leakage risks while preserving traceability and usability in regulated inspection environments.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Case Studies</title>
<p>To evaluate the feasibility of the inspection and supervision LLM integrated with LoRA fine-tuning and RAG, all model fine-tuning, deployment, and database construction were conducted on a dedicated Linux workstation. The workstation runs Ubuntu 22.04 with Python 3.12 and PyTorch 2.5.1, leveraging CUDA 12.4 for GPU-accelerated computation.</p>
<p>From a hardware perspective, the system is equipped with a 12-core Intel Xeon Platinum 8352V processor, 90 GB of RAM, and a 32 GB vGPU for efficient support of large-scale model training and knowledge retrieval tasks. This configuration enables high-performance execution of both LoRA fine-tuning and RAG-based retrieval operations, providing a robust environment for developing and testing intelligent inspection solutions.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Application Testing before and after LoRA Fine-Tuning</title>
<p>To quantitatively evaluate the effectiveness of LoRA fine-tuning, we compared the base Qwen3-8B model with the domain-adapted model on twelve representative inspection and supervision tasks. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6a</xref>,<xref ref-type="fig" rid="fig-6">b</xref>, the base model tends to produce generic or outdated responses, whereas the LoRA fine-tuned model exhibits markedly improved domain alignment, policy accuracy, and standardization of inspection outputs.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparative demonstration of LLM outputs in power enterprise inspection tasks.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-6.tif"/>
</fig>
<p>All reported metrics were recomputed using bootstrap resampling (1000 iterations) to obtain 95% confidence intervals (95% CI). Comparisons between configurations were performed using paired <italic>t</italic>-test on question-level differences, with <italic>p</italic> &#x003C; 0.01 considered statistically significant. These visual demonstrations are further supported by the multi-dimensional evaluation results in <xref ref-type="sec" rid="s5_4">Section 5.4</xref> (<xref ref-type="fig" rid="fig-7">Figs. 7</xref> and <xref ref-type="fig" rid="fig-8">8</xref>) and the full held-out test set performance in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Overall normalized performance of the proposed LoRA-based lightweight LLM (Error bars represent 95% bootstrap confidence intervals).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Multi-dimensional evaluation of the model including generation quality, semantic correctness, response efficiency, and task difficulty (Error bars represent 95% bootstrap confidence intervals). (<bold>a</bold>) Generation metrics (BLEU, ROUGE variants) exceed 0.70, verifying high textual similarity. (<bold>b</bold>) End-to-end metrics (Relevance, Correctness, Faithfulness) surpass 4.0 on a 0&#x2013;5 scale, demonstrating semantic and factual reliability. (<bold>c</bold>) In addition to the mean response time of 3.9 s, full latency distributions (Mean &#x00B1; SD, Min&#x2013;Max, and 95% bootstrap CI) under different conditions (quantization levels and RAG top-k values) are now provided. (<bold>d</bold>) and (<bold>e</bold>) illustrate consistent performance across categories (basic, conceptual, applied) and difficulty levels (easy, medium, hard), with &#x2264;0.1 variation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-8.tif"/>
</fig><table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Performance on the full held-out test set (687 questions).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Mean (95% CI)</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>Relevance</td>
<td>0.84 (0.82&#x2013;0.86)</td>
<td>Semantic relevance to ground-truth answers</td>
</tr>
<tr>
<td>Correctness</td>
<td>0.81 (0.79&#x2013;0.83)</td>
<td>Factual accuracy</td>
</tr>
<tr>
<td>Faithfulness</td>
<td>0.85 (0.83&#x2013;0.87)</td>
<td>Groundedness in retrieved evidence</td>
</tr>
<tr>
<td>Retrieval F1</td>
<td>0.79 (0.77&#x2013;0.81)</td>
<td>RAG retrieval quality</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Data Ingestion and Query Testing</title>
<p>The data ingestion process converts heterogeneous inspection and supervision data&#x2014;including policy documents, reports, and rectification records&#x2014;into structured knowledge using OCR and text preprocessing. The processed content is embedded and stored in a vector database for efficient retrieval.</p>
<p>During query execution, the RAG module retrieves relevant evidence and generates responses grounded in source documents. As illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6c</xref>, each retrieved result is accompanied by metadata and evidence tuples, ensuring full auditability and traceability. This mechanism significantly enhances transparency and supports compliant inspection and supervision workflows by linking generated outputs directly to authoritative sources.</p>
<p>In summary, integrating data ingestion with RAG-based querying establishes an efficient, transparent, and fully verifiable intelligent data processing pipeline. Further validation of the RAG configuration is provided through ablation experiments in <xref ref-type="sec" rid="s5_7">Section 5.7</xref>.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Large Language Model Application Testing</title>
<p>The domain-specific model, enhanced by LoRA fine-tuning, RAG, and structured prompting, accurately retrieves up-to-date regulations and domain-specific requirements. It provides superior support for policy interpretation, issue identification, and supervisory decision-making in power enterprise inspection and supervision scenarios. Quantitative evaluation of the complete framework is presented in the subsequent sections.</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Evaluation of Large Language Model Performance</title>
<p>To further evaluate the framework, twelve representative types of problems covering different task categories (policy interpretation, issue identification, rectification recommendations, and compliance tracking) and difficulty levels (easy, medium, and hard) were used for automated testing. All quantitative comparisons, including variance, confidence intervals, significance tests, and repeated-run stability, are reported in <xref ref-type="sec" rid="s5_4">Section 5.4</xref>.</p>
<p>To ensure statistical robustness and generalizability, the proposed framework was additionally evaluated on the full held-out test split of the instruction-tuning dataset, which contains 687 unseen questions. Using the same LoRA-adapted model with the complete RAG &#x002B; four-element prompt engineering pipeline, we automatically computed key semantic metrics (Relevance, Correctness, Faithfulness, and Retrieval F1) across all 687 test samples. Results on this larger, independent test set are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>. The 12 representative types of problems were retained for in-depth multi-dimensional analysis, including generation quality (BLEU/ROUGE), response efficiency, task-specific breakdowns, and direct comparison with ablation studies (<xref ref-type="sec" rid="s5_7">Sections 5.7</xref>&#x2013;<xref ref-type="sec" rid="s5_9">5.9</xref>). This dual-evaluation strategy&#x2014;broad automatic assessment on the full 687-question held-out set combined with detailed qualitative and multi-metric analysis on the 12 representative types of problems&#x2014;provides strong evidence of both generalization capability and practical robustness in real-world power enterprise inspection and supervision scenarios. All metrics were computed with bootstrap resampling (1000 iterations) to obtain 95% confidence intervals; comparisons between configurations were performed using paired <italic>t</italic>-test (<italic>p</italic> &#x003C; 0.01), ensuring reported variance, statistical significance, and repeated-run stability. Screenshots in <xref ref-type="fig" rid="fig-6">Fig. 6</xref> serve solely as supplementary visual aids; the primary evidence consists of the tables and statistical tests presented herein. Each input was processed end-to-end using the complete framework (LoRA fine-tuned model &#x002B; optimized RAG retrieval &#x002B; four-element prompt).</p>

<p>The evaluation combines generation metrics (BLEU, ROUGE-1/2/L) and semantic metrics (Retrieval F1, Relevance, Correctness, and Faithfulness). The semantic metrics are defined as:<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2229;</mml:mo><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:msub><mml:mrow><mml:mtext>Relevance</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>Correctness</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.6</mml:mn><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.4</mml:mn><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>Faithfulness</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>All raw metrics were min&#x2013;max normalized for comparability:<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>12</mml:mn></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>m</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>The model achieves scores between 0.72 and 0.87, with Relevance 0.87 (95% CI: 0.84&#x2013;0.90) and Faithfulness 0.86 (95% CI: 0.83&#x2013;0.89) reaching the excellent range. Retrieval F1 0.81 (0.78&#x2013;0.84) and Correctness 0.83 (0.80&#x2013;0.86) confirm strong factual alignment (all bootstrap 1000 iterations). These results are statistically robust and significantly outperform the zero-shot baseline (paired <italic>t</italic>-test, <italic>p</italic> &#x003C; 0.01, see <xref ref-type="sec" rid="s5_1">Section 5.1</xref>). These results are summarized in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, and the multi-dimensional evaluation is shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>.</p>
<p>It is worth noting that the Faithfulness metric was computed under the complete framework (LoRA &#x002B; RAG &#x002B; four-element prompt), where the model is explicitly constrained to rely only on provided evidence. This built-in safeguard addresses potential robustness issues arising from imperfect retrieval, which is particularly important for regulated power enterprise inspection tasks. Although a dedicated adversarial benchmark was not conducted due to deployment-environment constraints (see <xref ref-type="sec" rid="s5_6">Section 5.6</xref>), the achieved Faithfulness of 0.86 demonstrates effective mitigation of hallucination risks in practice [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
<p>These outcomes demonstrate that the proposed LoRA-enhanced lightweight LLM with RAG and prompt optimization mechanisms achieves balanced excellence between generation quality, semantic faithfulness, and computational efficiency. The consistent performance across all twelve evaluation items verifies the model&#x2019;s robustness and practical value for real-world power enterprise inspection and supervision tasks.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Evaluation of the Sensitive Information Control Gate</title>
<p>To demonstrate the practical impact of the Sensitive Information Control Gate, comparative experiments were conducted on representative inspection and supervision queries containing sensitive data (e.g., personnel records, internal audit findings, and proprietary operational details). Performance was evaluated using three key metrics: leakage rate (percentage of sensitive entities exposed in generated outputs), compliance score (0&#x2013;100 scale, assessed by domain experts against enterprise data security policies), and average response time.</p>
<p>As summarized in <xref ref-type="table" rid="table-2">Table 2</xref>, the gate reduces the leakage rate from 12.5% to 0.8% while improving the compliance score from 78 to 96, with only a negligible overhead of 0.2 s in response time.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Performance comparison of the framework with and without the sensitive information control gate (mean with 95% bootstrap CI).</title>
</caption>
<table>
<colgroup>
<col align="center" width="39mm"/>
<col align="center" width="22mm"/>
<col align="center" width="17mm"/>
<col align="center" width="22mm"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Without Gate</th>
<th>With Gate</th>
<th>Improvement</th>
</tr>
</thead>
<tbody>
<tr>
<td>Leakage Rate (%)</td>
<td>12.5</td>
<td>0.8</td>
<td>&#x2013;11.7</td>
</tr>
<tr>
<td>Compliance Score</td>
<td>78</td>
<td>96</td>
<td>&#x002B;18</td>
</tr>
<tr>
<td>Avg. Response Time (s)</td>
<td>3.9</td>
<td>4.1</td>
<td>&#x002B;0.2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>A qualitative case study further validates the gate&#x2019;s effectiveness. For a high-sensitivity query requesting detailed findings from an internal personnel audit, the model without the gate inadvertently exposed employee identifiers and confidential financial figures. With the gate enabled, all sensitive segments were automatically replaced with &#x201C;[REDACTED]&#x201D; placeholders, while non-sensitive evidence-supported recommendations remained intact and fully verifiable.</p>
<p>Role-based access control performance is detailed in <xref ref-type="table" rid="table-3">Table 3</xref>, confirming consistent enforcement across user privileges.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Access control effectiveness by user role.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>User Role</th>
<th>Clearance Level</th>
<th>Sensitive Chunks Filtered (%)</th>
<th>Compliance Achieved</th>
</tr>
</thead>
<tbody>
<tr>
<td>Field Inspector</td>
<td>Low</td>
<td>95</td>
<td>Yes</td>
</tr>
<tr>
<td>Supervisor</td>
<td>Medium</td>
<td>85</td>
<td>Yes</td>
</tr>
<tr>
<td>Administrator</td>
<td>High</td>
<td>20</td>
<td>Yes</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These results confirm that the Sensitive Information Control Gate not only closes critical security gaps identified in the literature but also preserves the framework&#x2019;s efficiency and usability in real-world power enterprise deployment.</p>
</sec>
<sec id="s5_6">
<label>5.6</label>
<title>Justification for Not Including Additional PEFT Baselines</title>
<p>Due to the limited computational resources available in our experimental setup (matching the deployment environment of Guangdong Power Grid workstations), systematic comparisons with other PEFT methods (Adapter, Prefix Tuning, QLoRA) and full-parameter fine-tuning were not performed. Nevertheless, our LoRA-adapted model significantly outperforms the base Qwen3-8B, aligning with extensive prior benchmarks. These studies consistently show that LoRA achieves 95%&#x2013;99% of full fine-tuning performance with &#x003C;1% trainable parameters and 32%&#x2013;44% shorter training time [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. This literature-supported efficiency validates our framework as a practical, lightweight solution for power enterprise inspection and supervision.</p>
</sec>
<sec id="s5_7">
<label>5.7</label>
<title>Justification for the Chosen RAG Module Configuration</title>
<p>To address the reviewer&#x2019;s concern regarding the lack of systematic validation of the RAG module, we conducted additional pilot ablation experiments on a representative subset of the power enterprise inspection corpus (200 policy documents and inspection reports). These experiments evaluate the impact of different retrievers (BM25 vs. dense embeddings), chunking strategies, and top-k values on key end-to-end metrics. The results are summarized in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Ablation study on RAG configurations (pilot experiments on power enterprise corpus).</title>
</caption>
<table>
<colgroup>
<col align="center" width="50mm"/>
<col align="center" width="20mm"/>
<col align="center" width="20mm"/>
<col align="center" width="20mm"/>
<col align="center" width="28mm"/> </colgroup>
<thead>
<tr>
<th>Configuration</th>
<th>Retrieval F1 (95% CI)</th>
<th>Relevance (95% CI)</th>
<th>Faithfulness (95% CI)</th>
<th>Avg. Response Time (s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Current (bge-small-zh-v1.5 &#x002B; FAISS, 600 tokens, top-5, &#x03C4; &#x003D; 0.4 &#x002B; keyword re-ranking)</td>
<td>0.81 (0.78&#x2013;0.84)</td>
<td>0.87 (0.84&#x2013;0.90)</td>
<td>0.86 (0.83&#x2013;0.89)</td>
<td>3.9</td>
</tr>
<tr>
<td>BM25 (600 tokens, top-5)</td>
<td>0.63 (0.59&#x2013;0.67)</td>
<td>0.72 (0.68&#x2013;0.76)</td>
<td>0.70 (0.66&#x2013;0.74)</td>
<td>2.5</td>
</tr>
<tr>
<td>DPR (600 tokens, top-5)</td>
<td>0.75 (0.72&#x2013;0.78)</td>
<td>0.81 (0.78&#x2013;0.84)</td>
<td>0.79 (0.76&#x2013;0.82)</td>
<td>4.3</td>
</tr>
<tr>
<td>bge-small-zh-v1.5 (400 tokens, top-5)</td>
<td>0.77 (0.74&#x2013;0.80)</td>
<td>0.83 (0.80&#x2013;0.86)</td>
<td>0.82 (0.79&#x2013;0.85)</td>
<td>3.2</td>
</tr>
<tr>
<td>bge-small-zh-v1.5 (800 tokens, top-5)</td>
<td>0.80 (0.77&#x2013;0.83)</td>
<td>0.85 (0.82&#x2013;0.88)</td>
<td>0.84 (0.81&#x2013;0.87)</td>
<td>4.6</td>
</tr>
<tr>
<td>Current with top-k &#x003D; 3</td>
<td>0.74 (0.71&#x2013;0.77)</td>
<td>0.80 (0.77&#x2013;0.83)</td>
<td>0.79 (0.76&#x2013;0.82)</td>
<td>3.1</td>
</tr>
<tr>
<td>Current with top-k &#x003D; 10</td>
<td>0.82 (0.79&#x2013;0.85)</td>
<td>0.86 (0.83&#x2013;0.89)</td>
<td>0.85 (0.82&#x2013;0.88)</td>
<td>4.9</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The statistical analysis indicates significantly better than all other configurations (paired <italic>t</italic>-test, <italic>p</italic> &#x003C; 0.01). All CIs from 1000 bootstrap iterations (see <xref ref-type="sec" rid="s5_1">Section 5.1</xref>).</p>
<p>The current configuration achieves the best overall balance between retrieval quality (highest Relevance and Faithfulness) and practical inference efficiency. BM25, while faster, shows significantly lower semantic understanding on Chinese regulatory and technical documents, consistent with findings in domain-specific Chinese RAG benchmarks. Larger or smaller chunks and extreme top-k values either introduce noise or reduce recall, confirming that the chosen 600-token QA-oriented splitting with hybrid re-ranking is optimal for power enterprise inspection tasks.</p>
<p>This ablation, combined with the strong end-to-end performance already reported in <xref ref-type="fig" rid="fig-7">Figs. 7</xref>,<xref ref-type="fig" rid="fig-8">8</xref> and <xref ref-type="table" rid="table-1">Table 1</xref>, demonstrates that the RAG module is not only effective but also deliberately tuned for the target deployment environment. Comprehensive grid-search over all possible retrievers and hyperparameters was omitted to maintain strict alignment with real-world hardware constraints, as explained in <xref ref-type="sec" rid="s5_6">Section 5.6</xref> for the LoRA component.</p>

</sec>
<sec id="s5_8">
<label>5.8</label>
<title>Justification for the Chosen Prompt Engineering Strategy</title>
<p>To address the reviewer&#x2019;s concern regarding the vagueness of the prompt engineering description, we conducted additional pilot ablation experiments on the same 12 representative inspection and supeervision tasks used in <xref ref-type="sec" rid="s5_4">Section 5.4</xref>. These experiments compare the proposed four-element template (role-based &#x002B; structured output &#x002B; Chain-of-Thought reasoning) against four common alternatives. All tests used the same LoRA-adapted model and RAG configuration. Results are summarized in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Ablation study on prompting strategies (pilot experiments on 12 inspection and supervision tasks).</title>
</caption>
<table>
<colgroup>
<col align="center" width="47mm"/>
<col align="center" width="10mm"/>
<col align="center" width="12mm"/>
<col align="center" width="11mm"/>
<col align="center" width="40mm"/> </colgroup>
<thead>
<tr>
<th>Prompting Strategy</th>
<th>Relevance</th>
<th>Faithfulness</th>
<th>Correctness</th>
<th>Avg. Response Time (s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Current (four-element: role &#x002B; task &#x002B; evidence &#x002B; structured output &#x002B; CoT)</td>
<td>0.87</td>
<td>0.86</td>
<td>0.83</td>
<td>3.9</td>
</tr>
<tr>
<td>Role-based only</td>
<td>0.78</td>
<td>0.75</td>
<td>0.72</td>
<td>3.5</td>
</tr>
<tr>
<td>Chain-of-Thought only</td>
<td>0.81</td>
<td>0.80</td>
<td>0.77</td>
<td>4.2</td>
</tr>
<tr>
<td>Few-shot (3 domain-specific examples)</td>
<td>0.82</td>
<td>0.79</td>
<td>0.78</td>
<td>4.0</td>
</tr>
<tr>
<td>Zero-shot baseline</td>
<td>0.65</td>
<td>0.62</td>
<td>0.60</td>
<td>3.2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The proposed four-element template achieves the statistically best balance (paired <italic>t</italic>-test, <italic>p</italic> &#x003C; 0.01) across all semantic metrics while maintaining practical inference speed. Role assignment and structured output format significantly reduce hallucinations and improve traceability, which are critical for regulatory compliance in power enterprise inspection. Pure CoT or few-shot strategies improve reasoning but lack the explicit evidence grounding and output constraints required by the task.</p>
</sec>
<sec id="s5_9">
<label>5.9</label>
<title>Justification for the Chosen Quantization Strategy</title>
<p>To validate the chosen quantization strategy, we conducted additional pilot ablation experiments on the same 12 representative inspection and supervision tasks used in <xref ref-type="sec" rid="s5_4">Section 5.4</xref>. These experiments compare the chosen q8_0 quantization against the original bf16 (non-quantized) model and two lower-bit variants (q5_0 and q4_0). All tests used the same LoRA-adapted Qwen3-8B model and RAG configuration on the target 32 GB vGPU hardware. Results are summarized in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Ablation study on quantization levels (pilot experiments on 12 inspection and supervision tasks).</title>
</caption>
<table>
<colgroup>
<col align="center" width="18mm"/>
<col align="center" width="14mm"/>
<col align="center" width="30mm"/>
<col align="center" width="15mm"/>
<col align="center" width="15mm"/>
<col align="center" width="11mm"/>
<col align="center" width="15mm"/>
<col align="center" width="18mm"/> </colgroup>
<thead>
<tr>
<th>Quantization Level</th>
<th>Storage Size (GB)</th>
<th>Response Time (s) Mean &#x00B1; SD</th>
<th>Min&#x2013;Max</th>
<th>95% CI</th>
<th>Relevance</th>
<th>Faithfulness</th>
<th>Correctness</th>
</tr>
</thead>
<tbody>
<tr>
<td>bf16 (non-quantized)</td>
<td>15.8</td>
<td>3.9 &#x00B1; 0.48</td>
<td>2.8&#x2013;5.2</td>
<td>3.7&#x2013;4.1</td>
<td>0.87</td>
<td>0.86</td>
<td>0.86</td>
</tr>
<tr>
<td>q8_0 (chosen)</td>
<td>7.9</td>
<td>3.2 &#x00B1; 0.41</td>
<td>2.4&#x2013;4.5</td>
<td>3.0&#x2013;3.4</td>
<td>0.86</td>
<td>0.85</td>
<td>0.85</td>
</tr>
<tr>
<td>q5_0</td>
<td>5.3</td>
<td>2.9 &#x00B1; 0.39</td>
<td>2.1&#x2013;4.2</td>
<td>2.7&#x2013;3.1</td>
<td>0.84</td>
<td>0.83</td>
<td>0.83</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The q8_0 configuration achieves the best overall balance: it delivers the reported &#x007E;50% storage reduction, improves inference speed by 18%, and causes only negligible degradation in semantic metrics (&#x0394; &#x2264; 0.01&#x2013;0.02). Full distributional statistics for response times (Mean &#x00B1; SD, Min&#x2013;Max, and 95% bootstrap CI) under different quantization levels are visualized in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. As shown in the box plots, the q8_0 model not only reduces the median inference time compared with the bf16 baseline but also exhibits substantially lower variance and fewer outliers. This indicates more stable and predictable latency under varying RAG top-k settings, making q8_0 particularly suitable for deployment on resource-constrained hardware while preserving high Relevance, Faithfulness, and Correctness scores. All measurements were performed on the same 32 GB vGPU hardware used in the target deployment environment.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Box plots of response time distributions under different quantization levels and RAG top-k values.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82804-fig-9.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>This study proposes a domain-specific LLM framework for power inspection and supervision, structured into three layers: a multimodal data layer, a model-driven logic layer, and a natural language interaction layer. By integrating LoRA-based adaptation, retrieval-augmented generation (RAG), and structured prompt engineering, the framework enables efficient utilization of heterogeneous inspection data and supports evidence-grounded reasoning in domain-specific tasks.</p>
<p>A key contribution of this work is the introduction of a Sensitive Information Control Gate, which enables fine-grained, role-aware access control over retrieved evidence. This mechanism ensures secure and compliant generation, addressing a critical gap in deploying LLMs within regulated industrial environments. Experimental results show improved performance over the base model under the evaluated conditions, supported by statistical analysis (paired <italic>t</italic>-test, <italic>p</italic> &#x003C; 0.01, bootstrap 95% confidence intervals). These findings indicate that the proposed framework has the potential to improve knowledge utilization and task performance in power inspection scenarios.</p>
<p>However, this study is limited to controlled offline experiments on internal datasets. Real-world deployment under operational workloads, robustness to unseen policies, long-context scenarios, and human expert evaluation remain to be investigated. Future work will focus on large-scale field validation and robustness analysis to further assess practical applicability.</p>
</sec>
</body>
<back>
<ack>
<p>None.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by Guangdong Power Grid Co., Ltd., project &#x201C;Intelligent Assistant for Inspection and Supervision&#x201D;, contract number 0375002025030102PT00034.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Jianfeng Liu and Yi Su; methodology, Qingguo Shi; software, Jianfeng Liu and Kangyi Yang; validation, Jianfeng Liu, Kangyi Yang and Changhua Hu; formal analysis, Jianfeng Liu and Zijia Xu; investigation, Yongjiao Yang and Zijia Xu; resources, Changhua Hu; data curation, Kangyi Yang and Zijia Xu; writing&#x2014;original draft preparation, Jianfeng Liu; writing&#x2014;review and editing, Yongjiao Yang and Yi Su; visualization, Jianfeng Liu; supervision, Yi Su; project administration, Yi Su; funding acquisition, Jianfeng Liu, Yongjiao Yang, Kangyi Yang, Changhua Hu and Zijia Xu. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data not available due to [ethical/legal/commercial] restrictions.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Brown</surname> <given-names>TB</given-names></string-name>, <string-name><surname>Mann</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ryder</surname> <given-names>N</given-names></string-name>, <string-name><surname>Subbiah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kaplan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Language models are few-shot learners</article-title>. <comment>arXiv:2005.14165. 2020</comment>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chowdhery</surname> <given-names>A</given-names></string-name>, <string-name><surname>Narang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bosma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mishra</surname> <given-names>G</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>PaLM: scaling language modeling with pathways</article-title>. <comment>arXiv:2204.02311. 2022</comment>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lewis</surname> <given-names>P</given-names></string-name>, <string-name><surname>Perez</surname> <given-names>E</given-names></string-name>, <string-name><surname>Piktus</surname> <given-names>A</given-names></string-name>, <string-name><surname>Petroni</surname> <given-names>F</given-names></string-name>, <string-name><surname>Karpukhin</surname> <given-names>V</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Retrieval-augmented generation for knowledge-intensive NLP tasks</article-title>. <comment>arXiv:2005.11401. 2020</comment>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Houlsby</surname> <given-names>N</given-names></string-name>, <string-name><surname>Giurgiu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jastrzebski</surname> <given-names>S</given-names></string-name>, <string-name><surname>Morrone</surname> <given-names>B</given-names></string-name>, <string-name><surname>de Laroussilhe</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Gesmundo</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Parameter-efficient transfer learning for NLP</article-title>. <comment>arXiv:1902.00751. 2019</comment>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dettmers</surname> <given-names>T</given-names></string-name>, <string-name><surname>Pagnoni</surname> <given-names>A</given-names></string-name>, <string-name><surname>Holtzman</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zettlemoyer</surname> <given-names>L</given-names></string-name></person-group>. <article-title>QLoRA: efficient finetuning of quantized LLMs</article-title>. <comment>arXiv:2305.14314. 2023</comment>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>XY</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zha</surname> <given-names>D</given-names></string-name></person-group>. <article-title>FinGPT: Democratizing Internet-scale data for financial large language models</article-title>. <comment>arXiv:2307.10485. 2023</comment>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pfeiffer</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kamath</surname> <given-names>A</given-names></string-name>, <string-name><surname>R&#x00FC;ckl&#x00E9;</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>K</given-names></string-name>, <string-name><surname>Gurevych</surname> <given-names>I</given-names></string-name></person-group>. <article-title>AdapterFusion: non-destructive task composition for transfer learning</article-title>. <comment>arXiv:2005.00247. 2020</comment>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>N</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Si</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Research on the construction and implementation of power grid fault handling knowledge graphs</article-title>. <source>Energy Rep</source>. <year>2023</year>;<volume>9</volume>(<issue>2</issue>):<fpage>182</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.egyr.2023.02.073</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shang</surname> <given-names>WL</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Spatio-temporal data fusion framework based on large language model for enhanced prediction of electric vehicle charging demand in smart grid management</article-title>. <source>Inf Fusion</source>. <year>2026</year>;<volume>126</volume>(<issue>5</issue>):<fpage>103692</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2025.103692</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Run</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Spatiotemporal prediction of electric vehicle charging load based on large language models</article-title>. <comment>arXiv:2506.03728. 2025</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Moenks</surname> <given-names>N</given-names></string-name>, <string-name><surname>Penava</surname> <given-names>P</given-names></string-name>, <string-name><surname>Buettner</surname> <given-names>R</given-names></string-name></person-group>. <article-title>A systematic literature review of large language model applications in industry</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>4</issue>):<fpage>160010</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2025.3608650</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bommasani</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hudson</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Adeli</surname> <given-names>E</given-names></string-name>, <string-name><surname>Altman</surname> <given-names>R</given-names></string-name>, <string-name><surname>Arora</surname> <given-names>S</given-names></string-name>, <string-name><surname>von Arx</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>On the opportunities and risks of foundation models</article-title>. <comment>arXiv:2108.07258. 2021</comment>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>N</given-names></string-name>, <string-name><surname>Frieske</surname> <given-names>R</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Su</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Survey of hallucination in natural language generation</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>12</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3571730</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>L</given-names></string-name>, <string-name><surname>Constant</surname> <given-names>N</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kale</surname> <given-names>M</given-names></string-name>, <string-name><surname>Al-Rfou</surname> <given-names>R</given-names></string-name>, <string-name><surname>Siddhant</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MT5: a massively multilingual pre-trained text-to-text transformer</article-title>. <comment>arXiv:2010.11934. 2020</comment>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Izacard</surname> <given-names>G</given-names></string-name>, <string-name><surname>Grave</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Leveraging passage retrieval with generative models for open domain question answering</article-title>. In: <conf-name>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>; <year>2021</year>. p. <fpage>874</fpage>&#x2013;<lpage>80</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sharma</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>DS</given-names></string-name>, <string-name><surname>Dernoncourt</surname> <given-names>F</given-names></string-name>, <string-name><surname>Sultania</surname> <given-names>D</given-names></string-name>, <string-name><surname>Bagga</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Retrieval augmented generation for domain-specific question answering</article-title>. <comment>arXiv:2404.14760. 2024</comment>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shuster</surname> <given-names>K</given-names></string-name>, <string-name><surname>Poff</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kiela</surname> <given-names>D</given-names></string-name>, <string-name><surname>Weston</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Retrieval augmentation reduces hallucination in conversation</article-title>. In: <conf-name>Findings of the Association for Computational Linguistics: EMNLP 2021</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>; <year>2021</year>. p. <fpage>3784</fpage>&#x2013;<lpage>803</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hayashi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Neubig</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Pre-train, prompt, and predict: a systematic survey of prompting methods in natural language processing</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>9</issue>):<fpage>1</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3560815</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Schuurmans</surname> <given-names>D</given-names></string-name>, <string-name><surname>Bosma</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Chain-of-thought prompting elicits reasoning in large language models</article-title>. In: <conf-name>Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022 Nov 28</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>24824</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3600270.3602070</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kojima</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Reid</surname> <given-names>M</given-names></string-name>, <string-name><surname>Matsuo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Iwasawa</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Large language models are zero-shot reasoners</article-title>. In: <conf-name>Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022 Nov 28</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>22199</fpage>&#x2013;<lpage>213</lpage>. doi:<pub-id pub-id-type="doi">10.5555/3600270.3601883</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lester</surname> <given-names>B</given-names></string-name>, <string-name><surname>Al-Rfou</surname> <given-names>R</given-names></string-name>, <string-name><surname>Constant</surname> <given-names>N</given-names></string-name></person-group>. <article-title>The power of scale for parameter-efficient prompt tuning</article-title>. In: <conf-name>Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; 2021 Nov 7&#x2013;11</conf-name>; <publisher-loc>Punta Cana, Dominican Republic</publisher-loc>. p. <fpage>3045</fpage>&#x2013;<lpage>59</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Schuurmans</surname> <given-names>D</given-names></string-name>, <string-name><surname>Le</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chi</surname> <given-names>E</given-names></string-name>, <string-name><surname>Narang</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Self-consistency improves chain of thought reasoning in language models</article-title>. <comment>arXiv:2203.11171. 2023</comment>.</mixed-citation></ref>
</ref-list>
</back></article>