<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">IASC</journal-id>
<journal-id journal-id-type="nlm-ta">IASC</journal-id>
<journal-id journal-id-type="publisher-id">IASC</journal-id>
<journal-title-group>
<journal-title>Intelligent Automation &#x0026; Soft Computing</journal-title>
</journal-title-group>
<issn pub-type="epub">2326-005X</issn>
<issn pub-type="ppub">1079-8587</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">38543</article-id>
<article-id pub-id-type="doi">10.32604/iasc.2023.038543</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Importance-Weighted Transfer Learning for Fault Classification under Covariate Shift</article-title>
<alt-title alt-title-type="left-running-head">Importance-Weighted Transfer Learning for Fault Classification under Covariate Shift</alt-title>
<alt-title alt-title-type="right-running-head">Importance-Weighted Transfer Learning for Fault Classification under Covariate Shift</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Pan</surname><given-names>Yi</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Xie</surname><given-names>Lei</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>leix@iipc.zju.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Su</surname><given-names>Hongye</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Institute of Cyber-Systems and Control, Zhejiang University</institution>, <addr-line>Hangzhou, 310027</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University</institution>, <addr-line>Hangzhou, 310027</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Lei Xie. Email: <email>leix@iipc.zju.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>06</day><month>09</month><year>2024</year></pub-date>
<volume>39</volume>
<issue>4</issue>
<fpage>683</fpage>
<lpage>696</lpage>
<history>
<date date-type="received">
<day>17</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>2</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_IASC_38543.pdf"></self-uri>
<abstract>
<p>In the process of fault detection and classification, the operation mode usually drifts over time, which brings great challenges to the algorithms. Because traditional machine learning based fault classification cannot dynamically update the trained model according to the probability distribution of the testing dataset, the accuracy of these traditional methods usually drops significantly in the case of covariate shift. In this paper, an importance-weighted transfer learning method is proposed for fault classification in the nonlinear multi-mode industrial process. It effectively alters the drift between the training and testing dataset. Firstly, the mutual information method is utilized to perform feature selection on the original data, and a number of characteristic parameters associated with fault classification are selected according to their mutual information. Then, the importance-weighted least-squares probabilistic classifier (IWLSPC) is utilized for binary fault detection and multi-fault classification in covariate shift. Finally, the Tennessee Eastman (TE) benchmark is carried out to confirm the effectiveness of the proposed method. The experimental result shows that the covariate shift adaptation based on importance-weight sampling is superior to the traditional machine learning fault classification algorithms. Moreover, IWLSPC can not only be used for binary fault classification, but also can be applied to the multi-classification target in the process of fault diagnosis.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Covariate shift adaption</kwd>
<kwd>nonlinear multi-mode process</kwd>
<kwd>importance weight sampling</kwd>
<kwd>multi-fault classification</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Data-focused fault detection in industrial process has been heavily researched in the past few years, especially for plant-wide nonlinear chemical processes [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-9">9</xref>]. However, most of the related papers are focused on fault detection in a single working condition. Since traditional machine learning algorithms cannot cope with unknown input data, these methods often perform very poorly with different test samples, so there is a great need in industry for learning models that can learn from many different datasets for different working conditions [<xref ref-type="bibr" rid="ref-10">10</xref>]. This drawback restricted the effectiveness of fault detection algorithms for nonlinear multi-mode processes. It is essential to propose some new methods to better solve the problem of model transfer under unknown operating conditions or drift of operating conditions over time.</p>
<p>The actual industrial process usually runs in multiple operating modes [<xref ref-type="bibr" rid="ref-11">11</xref>] because of raw material fluctuations, seasonal changes in market demand, etc. For large-scale multi-mode industrial fault detection, the traditional machine learning-based fault detection algorithms cannot detect the drift of the process working conditions in time, which may lead effect and performance of the pre-designed fault detection algorithm to continually decline over time. On the other hand, fault detection in multi-mode chemical processes is usually achieved based on model recognition or model division. A mathematical model is established based on historical process data to obtain the monitoring indicators of each model of the multi-model process, and then determine whether the fault has occurred based on the monitoring indicators, thus issuing alerts in time to avoid serious accidents.</p>
<p>In the existing literature, it is assumed that there is no significant drift in operating conditions, and the fluctuation characteristics of process variables are stationary sequences (such as Gaussian distribution, etc.). However, in the actual industrial production environment, the conditions frequently change during the accumulated long-term operation, which would lead to certain drifts in the probable distributions of processing variables in the process. On the other hand, since the dynamic characteristics are determined by the mechanism of the actual production process, the nonlinear dynamic connection between the process variables and the system output usually does not change significantly. It is assumed that the training and testing data have the same input distribution, and a majority of traditional machine learning based methods could not work effectively under covariate shift. It could be solved by reducing the differentiation of the distribution of the training and testing data [<xref ref-type="bibr" rid="ref-12">12</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>The existing literature regarding fault classification in nonlinear multi-mode processes usually does not consider the system operating conditions or the time-varying drift in terms of fault characteristics. The operation-condition ranges of the process system plants are assumed to be within the historical operation-condition ranges. In other words, they assume that there are no unknown operation conditions during the verification of fault classification in multi-mode processes. However, these assumptions are unreasonable for the actual industrial chemical process, because the data sample distributions of the training and testing dataset may not be perfectly consistent, which is due to the change of working conditions. And data characteristics can be changed due to fluctuations of raw materials. In addition, conventional data-driven methods work under the assumption that the dataset used for model training is sufficient, but the effective information or knowledge available in the actual production process usually is not sufficient to train the fault classification model. Therefore, it is significant to introduce the transfer learning [<xref ref-type="bibr" rid="ref-20">20</xref>] to the fault classification in the nonlinear multi-mode process. Transfer learning is a reliable method to solve the issue of deficiency data, and it utilizes the existing data or knowledge in historical scenes to assist the training process of the model [<xref ref-type="bibr" rid="ref-21">21</xref>]. In this study, the concept of covariate shift [<xref ref-type="bibr" rid="ref-22">22</xref>] in the transfer learning is introduced, and the importance-weight sampling is carried out. The contribution in this study can be summarized as:</p>
<p>1) This paper focuses on the application of covariate shifts for the industrial-process fault detection and classification. We introduced a practical reweighting method (importance weighting) [<xref ref-type="bibr" rid="ref-23">23</xref>], which can be formally proven that this approach is theoretically reliable [<xref ref-type="bibr" rid="ref-24">24</xref>].</p>
<p>2) The importance-weighted technique is introduced to address the problem of covariate shift. In this case, a detailed technical analysis and implementation based on such a method is introduced for a typical chemical process.</p>
<p>3) A comprehensive experiment to validate the effectiveness of the proposed method is carried out, in which, the performance analysis of multi-fault classification and binary fault classification under covariate shift is introduced.</p>
<p>The rest of this paper is organized as following: we provide some basic preliminaries related to this work in <xref ref-type="sec" rid="s2">Section 2</xref>; <xref ref-type="sec" rid="s3">Section 3</xref> gives the problem formulation and the detailed algorithm of our methodologies; we implement the algorithm for a benchmark industrial process in <xref ref-type="sec" rid="s4">Section 4</xref>; the discussion of results is presented in <xref ref-type="sec" rid="s5">Section 5</xref>; <xref ref-type="sec" rid="s6">Section 6</xref> conclude the research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Preliminaries</title>
<sec id="s2_1">
<label>2.1</label>
<title>Mutual Information (MI)</title>
<p>According to the probability theory and information theory, the mutual information (MI) of two random variables is a measurement of the mutual dependence between them, which is used to evaluate the amount of information contributed by the appearance of one random variable to the appearance of another random variable. Different from the correlation coefficient, mutual information is not limited to real-valued random variables, and generally determines the similarity of the joint distribution <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and the product of the decomposed marginal distribution <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In fact, MI is a measure of mutual dependence between two sets of events.</p>
<p>Formally, the MI of two discrete random variables <italic>X</italic> and <italic>Y</italic> can be defined as:</p>
<p><disp-formula id="eqn-1">
<label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Y</mml:mi></mml:mrow></mml:munder><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>X</mml:mi></mml:mrow></mml:munder><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the joint probability distribution function of <italic>X</italic> and <italic>Y</italic>, and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are the marginal probability distribution functions of <italic>X</italic> and <italic>Y</italic>, respectively.</p>
<p>Intuitively, MI measures the information shared by <italic>X</italic> and <italic>Y</italic>: it measures the degree to which one of these two variables is known to reduce the uncertainty of the other. For example, if <italic>X</italic> and <italic>Y</italic> are independent of each other, we can conclude that <italic>X</italic> does not provide any information to <italic>Y</italic>, and vice versa, so their MI<inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. At the other extreme, if <italic>X</italic> is a deterministic function of <italic>Y</italic>, and <italic>Y</italic> is also a deterministic function of <italic>X</italic>, then all the information passed is shared by <italic>X</italic> and <italic>Y</italic>: the value of <italic>X</italic> determines the value of <italic>Y</italic>, and vice versa. Therefore, in this case, the MI and <italic>Y</italic> (or <italic>X</italic>) alone contain the same uncertainty, which is called the entropy of <italic>Y</italic> (or <italic>X</italic>). Moreover, this MI is the same as the entropy of <italic>X</italic> and the entropy of <italic>Y</italic>. Indeed, a very special case of this situation is when <italic>X</italic> and <italic>Y</italic> are actually the same random variable.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Importance-Weight Sampling with Covariate Shift</title>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Definition of Covariate Shift</title>
<p>Covariate shift is first proposed in an article in the field of statistics [<xref ref-type="bibr" rid="ref-24">24</xref>]. It is defined as the condition in which the input data (i.e., the training and testing dataset) apply different distributions, and the conditional distribution of the output of a given input data remain unchanged, and such a condition is defined covariate shift.</p>
<p>It is assumed that the input space of the source and target domain are both <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>X</mml:mi></mml:math></inline-formula>, and the output space are both <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>Y</mml:mi></mml:math></inline-formula>. The marginal distribution of the source domain <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is inconsistent with the joint distribution of the target domain <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (i.e., <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2260;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>), but the conditional distributions of the two domains are consistent (i.e., <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>).</p>
<p>The introduction of covariate shift aims to use labeled source data and unlabeled target data to learn a model for labeling target data [<xref ref-type="bibr" rid="ref-25">25</xref>]. A common method of covariate shift adaptation is to compute density ratio weights from unlabeled source data and target data, and then to learn the final hypothesis by directly minimizing the weighted loss [<xref ref-type="bibr" rid="ref-23">23</xref>].</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Overview of Importance Weight Sampling Technique</title>
<p>Given the input variable <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="bold-italic">X</mml:mi><mml:mo>&#x2282;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> , in which <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>d</mml:mi></mml:math></inline-formula> denotes the input dimensionality, and the output variable <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>Y</mml:mi><mml:mo>&#x225C;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> that is a set of categories for classification. In supervised learning, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> is usually assumed to be independently drawn from an input probability distribution with density <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>y</mml:mi></mml:math></inline-formula> is independently concluded from a conditional probability distribution with density<inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
<p>However, with covariate shift, the input distribution of the source domain is inconsistent with that of the target domain, meanwhile the conditional distribution of the two domains remains unchanged. It is supposed that the labelled training samples are defined as <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>, where the independent and identically distributed training input samples <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are drawn from a probability distribution with strictly positive density <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and the training output samples <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> are drawn from a conditional probability distribution with density <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In contrast, the unlabelled test input samples are in general unlabelled, then we suppose that <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mo>&#x2282;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is concluded independently from a probability distribution with density <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Importance sampling aims to solve the problem of covariate shift, the <italic>importance weight</italic> is defined as:</p>
<p><disp-formula id="eqn-2">
<label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:mi>w</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2261;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula>with the defined importance weight, the different training and testing input distributions can be altered. Then, for any function <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, we can derive the following equation:</p>
<p><disp-formula id="eqn-3">
<label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:mo>&#x222B;</mml:mo><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x222B;</mml:mo><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula>with the above expressions, we can conclude that the input distribution (i.e., the input sample weights) can be altered in the training process of any prescribed machine learning model according to the input distribution of the testing dataset. Therefore, the importance weighting also utilizes the labelled data of the source domain and the unlabelled data of the target domain to guide the knowledge transfer.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Fault Classification Model</title>
<sec id="s3_1">
<label>3.1</label>
<title>Feature Selection Using Mutual Information</title>
<p>In this study, there is a hypothesis that samples that are frequently grouped into a category, usually have greater MI with this category. Then we introduce the general feature selection method based on MI. The steps of the MI feature selection procedure: 1) divide the dataset; 2) sort these features according to their MI values; 3) select the top <italic>n</italic> features and adopt prescribed machine learning model to train; 4) evaluate the error rate of feature subset on the testing dataset. Through the MI-based feature selection, these features related to the samples from the process can be sorted by correlation, but how many features need to be selected depends on prior knowledge.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Importance-Weighted Least-Squares Probabilistic Classifier</title>
<p>The multi-mode process probabilistic fault classifier is used to estimate the class-posterior probability <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (i.e., to predict a class label <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> for a test input point <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>), which can be further utilized in the fault detection of nonlinear multi-mode processes. The class-posterior probability <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is formulated as:</p>
<p><disp-formula id="eqn-4">
<label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2261;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>-dimensional parameter vector and <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a kernel function, which is used to extend <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> to a higher-dimensional state space, thus making the simple linear LSPC classifier applicable in the original low-dimensional probabilistic classifier <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Then, the commonly used Gaussian kernel is defined as:</p>
<p><disp-formula id="eqn-5">
<label>(5)</label>
<mml:math id="mml-eqn-5" display="block"><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:msup><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> denotes the Gaussian kernel width. The goal is to minimize the following squared error performance index <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, by determining the LSPC parameter <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>:</p>
<p><disp-formula id="eqn-6">
<label>(6)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2261;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mo>&#x222B;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mo>&#x222B;</mml:mo><mml:mi>p</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mo>&#x222B;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>+</mml:mo><mml:mi>C</mml:mi></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:msubsup><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:mi mathvariant="bold-italic">Q</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">q</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>C</mml:mi></mml:math></inline-formula> is a constant, and the elements of the <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> matrix <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi mathvariant="bold-italic">Q</mml:mi></mml:math></inline-formula> and the <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>-dimensional vector <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi mathvariant="bold-italic">q</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>are defined as:</p>
<p><disp-formula id="eqn-7">
<label>(7)</label>
<mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mstyle mathsize="0.85em"><mml:mrow><mml:mo>&#x222B;</mml:mo></mml:mrow></mml:mstyle></mml:mrow><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-8">
<label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2261;</mml:mo><mml:mo>&#x222B;</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Then, we can formally approximate <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi mathvariant="bold-italic">Q</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi mathvariant="bold-italic">q</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, using the aforementioned importance weight sampling method, i.e., the importance weight <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. As a matter of fact, <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi mathvariant="bold-italic">Q</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi mathvariant="bold-italic">q</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can be expressed as follows:</p>
<p><disp-formula id="eqn-9">
<label>(9)</label>
<mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x222B;</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-10">
<label>(10)</label>
<mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mo>&#x222B;</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x222B;</mml:mo><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the training input density for a certain class <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>y</mml:mi></mml:math></inline-formula>. Then, using the training samples <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi mathvariant="bold-italic">Q</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi mathvariant="bold-italic">q</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can be estimated in accordance with <xref ref-type="disp-formula" rid="eqn-9">Eqs. (9)</xref> and <xref ref-type="disp-formula" rid="eqn-10">(10)</xref>:</p>
<p><disp-formula id="eqn-11">
<label>(11)</label>
<mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>Q</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup></mml:mrow></mml:msub><mml:mo>&#x2261;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mo>&#x00A0;</mml:mo><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>w</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03BD;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="eqn-12">
<label>(12)</label>
<mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>q</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2261;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>&#x003A;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:munder><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>w</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03BD;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>where the class-prior probability <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is estimated by <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msubsup><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the number of training samples with label <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>y</mml:mi></mml:math></inline-formula>, and the flattening parameter <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03BD;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. As argued in [<xref ref-type="bibr" rid="ref-21">21</xref>], there exists the problem of bias-variance trade-off (i.e., the importance weights influence the bias and variance of the IWLSPC model). If <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>&#x03BD;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, the bias gets smaller, but the variance tends to be larger; yet if <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>&#x03BD;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, then the bias is larger, but the variance is smaller. The problem can be formally described as:</p>
<p><disp-formula id="eqn-13">
<label>(13)</label>
<mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2261;</mml:mo><mml:munder><mml:mrow><mml:mi>arg</mml:mi><mml:mspace width="thinmathspace" /><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:munder><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msubsup><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">Q</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">q</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:msubsup><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a regularization term with the parameter <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> avoiding over-fitting. Then, the IWLSPC solution is given analytically as:</p>
<p><disp-formula id="eqn-14">
<label>(14)</label>
<mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">Q</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">q</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denotes the <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>-dimensional identity matrix. Considering that the class-posterior probability is non-negative by its definition, then the solution <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> can be modified as:</p>
<p><disp-formula id="eqn-15">
<label>(15)</label>
<mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2261;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>Z</mml:mi></mml:mfrac><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p>Finally, using the learned class-posterior probability <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, we can predict the class label <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of a new test sample <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> by</p>
<p><disp-formula id="eqn-16">
<label>(16)</label>
<mml:math id="mml-eqn-16" display="block"><mml:msup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2261;</mml:mo><mml:munder><mml:mrow><mml:mi>arg</mml:mi><mml:mspace width="thinmathspace" /><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mi>y</mml:mi></mml:munder><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p>If there does not exist covariate shift during a steady-state operation of the nonlinear process, i.e., <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, the IWLSPC is actually equivalent to the traditional LSPC method.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Fault Classification in TE Process</title>
<sec id="s4_1">
<label>4.1</label>
<title>TE Process</title>
<p>In the field of process system engineering, Tennessee Eastman (TE) process is a widely used Benchmark for process control [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>] and fault monitoring [<xref ref-type="bibr" rid="ref-29">29</xref>&#x2013;<xref ref-type="bibr" rid="ref-31">31</xref>]. The TE process is an open simulation platform in chemical engineering and process industry, and it&#x2019;s developed by Eastman Company from the United States. With the support to the dynamic simulation of a chemical reaction process, various features in the complex industrial process system can be well simulated with TE process. The data generated by the TE has complex characteristics (e.g., time-varying, strong coupling, and high nonlinearity). Therefore, it is widely adopted in fields such as system optimization, process monitoring and fault diagnosis, and it also can be used to verify the control and fault diagnosis for industrial chemical processes.</p>
<p>The TE process consists of five units: 1) stripper reboiler, 2) condenser, 3) flash separator, 4) two-phase exothermic reactor and 5) circulating compressor (as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>). The process contains a total of 52 variables, in which 22 are process variables, 18 are component variables and the other 12 are manipulated variables.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The control structure and flow diagram of TE process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_38543-fig-1.tif"/>
</fig>
<p>The TE process contains four gaseous raw materials (<italic>A</italic>, <italic>C</italic>, <italic>D</italic> and <italic>E</italic>), two liquid products (<italic>G</italic> and <italic>H</italic>), and by-products (<italic>F</italic>) and inert gas (<italic>B</italic>). The irreversible, exothermic chemical reaction in TE process is as follows:</p>
<p><disp-formula id="eqn-17">
<label>(17)</label>
<mml:math id="mml-eqn-17" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>2</mml:mn><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>g</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The gas components (<italic>A</italic>, <italic>C</italic>, <italic>D</italic> and <italic>E</italic>) and the inert component (<italic>B</italic>) are fed into the reactor, thus forming the liquid products (<italic>G</italic> and <italic>H</italic>), where the reaction rate obeys the Arrhenius function of the reaction kinetics. Coming out of the reactor, the product steam is condensed to liquid and transferred to a liquid separator, from which a string of steam would be resent to the reactor through the compressor. Some recirculation shunt should be dismissed to avoid accumulating the by-products and inert components in the reaction. The condensed products from the separator (Steam 10) are sent to the stripper. Steam 4 is applied to remove the remaining reagents, in Steam 10, mixed with the recirculating steam from Steam 5. The products (G and H) are sent to the downstream process. Most of the by-products are evacuated as gas in the vapor-liquid separator. According to the different component mass ratios of <italic>G</italic> and <italic>H</italic> in the product, the TE process supports 7 operation modes/working conditions. The simulation platform can set 21 fault modes, e.g., the fault modes of step type, random change type, slow drift type, valve sticky type, valve stuck type and unknown types.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experiment Setup</title>
<p>The TE process supports a variety of operation modes and a total of 28 prescribed fault types are given in detail. The simulated multi-mode data can be obtained from the TE process. A number of typical fault conditions in TE simulation dataset are selected, i.e., the TE process operation data without fault, and fault status #4, #10, #13, #16, #18 and #28. Then, the implementation process of the algorithm is described. We continuously extract a group pf time-series samples in different conditions (including different conditions with 6 faults and the normal condition). As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, two different operation modes are set up in this experiment.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Operation conditions in two different operation modes (mode 1 is set for training, and mode 2 is set for testing)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Settings</th>
<th>Mode 1</th>
<th>Mode 2</th>
</tr>
</thead>
<tbody>
<tr>
<td>Production setpoint</td>
<td>22.89</td>
<td>20.2</td>
</tr>
<tr>
<td>Strip setpoint</td>
<td>50</td>
<td>50</td>
</tr>
<tr>
<td>Sep level setpoint</td>
<td>50</td>
<td>50</td>
</tr>
<tr>
<td>Reactor level setpoint</td>
<td>65</td>
<td>65</td>
</tr>
<tr>
<td>Reactor press setpoint</td>
<td>2800</td>
<td>2800</td>
</tr>
<tr>
<td>Mole % G setpoint</td>
<td>53.8</td>
<td>90.07</td>
</tr>
<tr>
<td><italic>y</italic><sub><italic>A</italic></sub> setpoint</td>
<td>63.14</td>
<td>61.47</td>
</tr>
<tr>
<td><italic>y</italic><sub><italic>AC</italic></sub> setpoint</td>
<td>51</td>
<td>48.79</td>
</tr>
<tr>
<td>Reactor temp setpoint</td>
<td>122.9</td>
<td>123</td>
</tr>
<tr>
<td>Recycle valve position</td>
<td>0</td>
<td>71.17</td>
</tr>
<tr>
<td>Steam valve position</td>
<td>0</td>
<td>1</td>
</tr>
<tr>
<td>Agitator setting</td>
<td>100</td>
<td>100</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The training dataset contains the samples obtained in mode 1, in which the data in normal condition and in fault condition is 2400, respectively. Meanwhile the testing dataset contains samples obtained in mode 2, in which the data in normal condition and in fault condition is 200, respectively. The data generated in mode 1 is used to train the binary fault classification model, and the data generated in mode 2 is used to verify the TE process fault classification (i.e., classification accuracy). Since the training and testing datasets are obtained in different operation modes in TE process, it can be considered that the probability distribution of input samples in training datasets is different from that of the testing dataset. There exists the covariate shift phenomenon between the training dataset in operation mode 1 and the testing dataset in operation mode 2. To reduce the amount of data and increase the efficiency of the method, it is necessary to select some variables from all process variables, which have strong correlations with the fault variables. We introduce the mutual information for such a task. The MI is equivalent to dimensionality-reduction operation on the original data.</p>
<p>In the experiment, the MI-based variable selection works as follows: 1) re-arrange the variables in the training and testing dataset in the order of MI value; 2) choose 25 variables with largest MI value in the training and testing dataset, respectively; 3) observe the selected 25 variables, and select all the common variables among them; 4) calculate the MI value between the selected common variables, respectively, and remove these variables with low MI value (there is weak correlation or negative correlation, this operation is to prevent the occurrence of negative transfer); 5) delete other variables, the remaining variables are exactly the process variables we need, and these variables selected can be used to train the model of fault classification.</p>
<p>The MI-based pre-processing can not only reduce the computational workload, but also make the selected data more effective. On the other hand, in addition to data redundancy, not all the data are valid for fault classification and fault detection model. Some data have stronger correlation with the fault, and that the correlation between process variable and fault variable is not always the same. Besides, there is some data that cannot indicate whether there exist some faults in the process system, and the data is invalid for fault detection. Therefore, with the MI-based feature selection, we can extract more useful information for fault classification model and form the reorganized input dataset for training process. In this way, the data pre-processing can greatly eliminate the invalid samples and variables, and retain the most relevant variables for fault detection and classification.</p>
<p>The training and testing dataset processed by MI feature selection is input to IWLSPC for fault detection and classification. Specifically, we perform the testing with binary fault classification and multiple fault classification. In addition, the IWLSPC is compared with other existing methods in detail.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Result</title>
<sec id="s5_1">
<label>5.1</label>
<title>Comparison of Binary Fault Classification Performances under Covariate Shift</title>
<p>Firstly, we apply the proposed method to the binary fault detection and classification of TE process. In fact, if the binary fault situation refers to the two categories of having some specific faults or no fault, then binary fault classification is equivalent to the judgement of the process fault, i.e., exist or not exist. Therefore, to improve the comparability of different binary fault classification, we classify all the fault categories with the case without fault, that is, the designed algorithms should be able to distinguish whether there exist some faults or not in the TE process under covariate shift. The detailed performance comparison of binary fault classification is shown in <xref ref-type="table" rid="table-2">Table 2</xref>. Taking no fault and specific fault as the two labels of binary fault classification, the results presented are the accuracy of fault classification with the testing dataset. In addition, the results marked as <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msup><mml:mo>&#x00A0;</mml:mo><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> in the table are the results of data-driven fault classification model trained by the validation dataset with particularly few samples (only 200 samples), the results are used to demonstrate the effectiveness of importance-weighted transfer learning under covariate shift.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of binary fault classification performances under covariate shift (no fault <italic>vs</italic>. specific fault)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>SVM<sup>&#x002A;</sup></th>
<th>SVM</th>
<th>MI &#x002B; SVM<sup>&#x002A;</sup></th>
<th>MI &#x002B; SVM</th>
<th>LSPC<sup>&#x002A;</sup></th>
<th>LSPC</th>
<th>IWLSPC</th>
<th>MI &#x002B; LSPC<sup>&#x002A;</sup></th>
<th>MI &#x002B; LSPC</th>
<th>MI &#x002B; IWLSPC</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>100</td>
<td>91</td>
<td><bold>100</bold></td>
<td>95.5</td>
<td>86</td>
<td>91.25</td>
<td>59.75</td>
<td><bold>100</bold></td>
<td>96</td>
<td>97</td>
</tr>
<tr>
<td>2</td>
<td>67.5</td>
<td>56.25</td>
<td>74.5</td>
<td>80</td>
<td>75</td>
<td>75.75</td>
<td>74</td>
<td>87</td>
<td>86.75</td>
<td><bold>95.71</bold></td>
</tr>
<tr>
<td>3</td>
<td>71.5</td>
<td>81.5</td>
<td>79.5</td>
<td>82</td>
<td>41</td>
<td>79.25</td>
<td>77.75</td>
<td>83.5</td>
<td>81</td>
<td><bold>84.25</bold></td>
</tr>
<tr>
<td>4</td>
<td>58</td>
<td>98.25</td>
<td><bold>100</bold></td>
<td>100</td>
<td>69</td>
<td>96.25</td>
<td>95.75</td>
<td>73.5</td>
<td>100</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>5</td>
<td>99.5</td>
<td>83.25</td>
<td><bold>100</bold></td>
<td>99.75</td>
<td><bold>100</bold></td>
<td>64.75</td>
<td>75.75</td>
<td>100</td>
<td>80.25</td>
<td>87.5</td>
</tr>
<tr>
<td>7</td>
<td>99.5</td>
<td>89.5</td>
<td><bold>100</bold></td>
<td>100</td>
<td><bold>100</bold></td>
<td>81.5</td>
<td>93</td>
<td><bold>100</bold></td>
<td>91.5</td>
<td><bold>100</bold></td>
</tr>
<tr>
<td>8</td>
<td>53.5</td>
<td>83</td>
<td>83</td>
<td>87</td>
<td>90</td>
<td>85.5</td>
<td>72.5</td>
<td><bold>99.5</bold></td>
<td>88</td>
<td>99.25</td>
</tr>
<tr>
<td>9</td>
<td>60.5</td>
<td>59.75</td>
<td>68</td>
<td>65</td>
<td>52</td>
<td>60.25</td>
<td>57.25</td>
<td>64.5</td>
<td>72</td>
<td><bold>85</bold></td>
</tr>
<tr>
<td>10</td>
<td>45.5</td>
<td>62.75</td>
<td><bold>100</bold></td>
<td>88</td>
<td>40.5</td>
<td>56</td>
<td>58.5</td>
<td>90</td>
<td>90.25</td>
<td>90.25</td>
</tr>
<tr>
<td>11</td>
<td>44.5</td>
<td>78.5</td>
<td>61.5</td>
<td>94</td>
<td>47.5</td>
<td>80</td>
<td>74.5</td>
<td>61</td>
<td>94.25</td>
<td><bold>94.5</bold></td>
</tr>
<tr>
<td>12</td>
<td>95</td>
<td>89</td>
<td><bold>100</bold></td>
<td>95</td>
<td>71</td>
<td>81.25</td>
<td>64.75</td>
<td>85</td>
<td>85.75</td>
<td>91.75</td>
</tr>
<tr>
<td>13</td>
<td>61.5</td>
<td>48.25</td>
<td>64</td>
<td>68.75</td>
<td>64</td>
<td>68.75</td>
<td>45.75</td>
<td>75</td>
<td>83</td>
<td><bold>83.75</bold></td>
</tr>
<tr>
<td>14</td>
<td>49.5</td>
<td>59.25</td>
<td>72.5</td>
<td>81.75</td>
<td>72.5</td>
<td>51.75</td>
<td>54.75</td>
<td>75</td>
<td>83</td>
<td><bold>95.5</bold></td>
</tr>
<tr>
<td>15</td>
<td>47</td>
<td>34.25</td>
<td>69.5</td>
<td>68.75</td>
<td>47</td>
<td>27.75</td>
<td>47.75</td>
<td>63</td>
<td>50</td>
<td><bold>68.5</bold></td>
</tr>
<tr>
<td>16</td>
<td>39.5</td>
<td>46.5</td>
<td>74.5</td>
<td>58</td>
<td>29.5</td>
<td>50</td>
<td>50</td>
<td>80</td>
<td>50</td>
<td>50</td>
</tr>
<tr>
<td>17</td>
<td>52.5</td>
<td>67.75</td>
<td>81</td>
<td>78</td>
<td>58</td>
<td>70</td>
<td>72.5</td>
<td>74.5</td>
<td>75</td>
<td><bold>93</bold></td>
</tr>
<tr>
<td>18</td>
<td>75.5</td>
<td>87.75</td>
<td>81.5</td>
<td>91.25</td>
<td>79.5</td>
<td>84.25</td>
<td>77.5</td>
<td>95</td>
<td>91.5</td>
<td><bold>99</bold></td>
</tr>
<tr>
<td>19</td>
<td>69.5</td>
<td>84.25</td>
<td>82</td>
<td>95.5</td>
<td>46.5</td>
<td>79.5</td>
<td>61.5</td>
<td>88.5</td>
<td>93.75</td>
<td><bold>95.5</bold></td>
</tr>
<tr>
<td>20</td>
<td>52.5</td>
<td>57.25</td>
<td>74.5</td>
<td>74.25</td>
<td>55</td>
<td>56</td>
<td>54.5</td>
<td>66.5</td>
<td>73.5</td>
<td><bold>82.5</bold></td>
</tr>
<tr>
<td>21</td>
<td>55</td>
<td>46.75</td>
<td><bold>66</bold></td>
<td>51.75</td>
<td>65</td>
<td>41.75</td>
<td>51.25</td>
<td>70.5</td>
<td>53.25</td>
<td>61</td>
</tr>
<tr>
<td>22</td>
<td>45.5</td>
<td>59</td>
<td><bold>87.5</bold></td>
<td>63.25</td>
<td>68</td>
<td>59.5</td>
<td>59</td>
<td>77</td>
<td>63.25</td>
<td>72</td>
</tr>
<tr>
<td>23</td>
<td>60.5</td>
<td>52.25</td>
<td>90</td>
<td>78.5</td>
<td>57</td>
<td>54.25</td>
<td>56</td>
<td>68</td>
<td>90.5</td>
<td><bold>92</bold></td>
</tr>
<tr>
<td>24</td>
<td>78</td>
<td>80.25</td>
<td>79</td>
<td><bold>92.75</bold></td>
<td>31</td>
<td>82.25</td>
<td>80</td>
<td>77</td>
<td>92</td>
<td>87</td>
</tr>
<tr>
<td>25</td>
<td>66.5</td>
<td>47.25</td>
<td><bold>85</bold></td>
<td>56.75</td>
<td>67.5</td>
<td>34</td>
<td>61</td>
<td>78.5</td>
<td>56.5</td>
<td>64.5</td>
</tr>
<tr>
<td>26</td>
<td>50.5</td>
<td>52.25</td>
<td>87</td>
<td>84</td>
<td>67.5</td>
<td>41</td>
<td>42</td>
<td>96.5</td>
<td>86.75</td>
<td><bold>87.75</bold></td>
</tr>
<tr>
<td>27</td>
<td>50</td>
<td>66.75</td>
<td>63</td>
<td>90.5</td>
<td>29</td>
<td>67</td>
<td>66.75</td>
<td>59.5</td>
<td>90</td>
<td><bold>95</bold></td>
</tr>
<tr>
<td>28</td>
<td>52</td>
<td>58</td>
<td>74.5</td>
<td>62.75</td>
<td>43</td>
<td>65.5</td>
<td>64.75</td>
<td>62.5</td>
<td>70.5</td>
<td><bold>73.75</bold></td>
</tr>
<tr>
<td>Avg.</td>
<td>62.98</td>
<td>67.43</td>
<td>81.41</td>
<td>80.84</td>
<td>61.19</td>
<td>66.11</td>
<td>64.76</td>
<td>79.67</td>
<td>80.31</td>
<td><bold>86.15</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>First, we sorted the correlation of the characteristics related to the fault according to their MI values, and then selected the most relevant process variables. The pre-processing can not only reduce the amount of data, but also significantly improve the effectiveness of the data training. We compared the performance of extracting feature variables with/without MI. As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, by introducing the MI to the extraction of variables, the IWLSPC can better detect whether a fault exists or not.</p>

<p>Because the training and testing datasets we use are obtained in different operation modes, the process dynamic characteristics and fault-related information contained in the training and testing datasets are inconsistent. Once the training process completes, the traditional machine learning based classifications cannot adapt to the change of operation modes in the testing dataset according to the sample distribution. The importance weighting used in this study considers the probability distribution of the test input samples during the training process, and it carries out the weighted re-organization of the training input samples based on the importance weights, which can improve the accuracy of fault classification on the testing dataset. As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, for most types of faults, the performance of the MI-based IWLSPC is better than that of the traditional LSPC, and the total average performance of 28 fault types is also better than the MI-based LSPC. The result shows that the IWLSPC can better detect whether there exists a fault under covariate shift. It achieves superior performance to detect whether there exist some specific faults or there do not exist faults for the current state.</p>

<p>Furthermore, we compared the effectiveness of commonly used machine learning based classification algorithms and applied them to the binary fault classification for the TE process. As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, we can see that in most cases, the IWLSPC achieves excellent fault classification performance. Although SVM has some advantages in some cases, IWLSPC achieves better fault classification performances in most of the fault types. However, the reliability of MI &#x002B; IWLSPC can be guaranteed, which is exactly why MI-based IWLSPC outperforms for the total average fault classification accuracy. It is proven that LSPC and IWLSPC can not only be used to binary fault classification, but also can be applied to classify multiple faults simultaneously.</p>

<p>In order to demonstrate the effectiveness of importance-weighted transfer learning under covariate shift. We have also conducted experiments with extremely small samples, in which we directly train the fault classification model on the small dataset. Specifically, we trained the afore-mentioned fault classification models on the validation set with only 200 samples. This kind of small sample training is very practical for nonlinear multi-mode industrial processes, because some of the actual industrial processes usually do not have sufficient historical data under working conditions. The results marked with <sup>&#x002A;</sup> in <xref ref-type="table" rid="table-2">Table 2</xref> show that the performance of directly training on the validation dataset is poor, because there is not sufficient data for the training process. However, with the importance-weighted transfer learning method, the data-driven fault classification model achieves satisfactory performance even on a small sample verification dataset where the covariate shift exists. It confirms the effectiveness of the importance-weighted transfer learning method.</p>

</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Comparison of Multi-Fault Classification Performance with Other Methods</title>
<p>The probabilistic-based classification method for industrial process fault classification works well. In this study, we compared the effectiveness of the BP neural network, traditional LSPC and IWLSPC that is suitable for dealing with covariate shift. It is noted that we used MI to extract feature variables in both methods, thus mainly focusing on the difference among BP neural network, LSPC and IWLSPC, i.e., the effectiveness of the importance-weighted method. The detailed comparison is shown in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of the accuracy of the BP neural network, traditional LSPC and importance-weighted LSPC for multi-fault classification under covariate shift</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Fault 1</th>
<th>Fault 2</th>
<th>Fault 3</th>
<th>Fault 4</th>
<th>Fault 5</th>
<th>Fault 7</th>
<th>Total average</th>
</tr>
</thead>
<tbody>
<tr>
<td>MI &#x002B; BP</td>
<td>64.5</td>
<td>21</td>
<td>0</td>
<td>8</td>
<td>0</td>
<td>100</td>
<td>32.5</td>
</tr>
<tr>
<td>MI &#x002B; LSPC</td>
<td>95.5</td>
<td>39</td>
<td>18.5</td>
<td>100</td>
<td>83</td>
<td>0</td>
<td>56</td>
</tr>
<tr>
<td>MI &#x002B; IWLSPC</td>
<td>92</td>
<td>69.5</td>
<td>89.5</td>
<td>91.5</td>
<td>53.5</td>
<td>38.5</td>
<td>72.42</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In <xref ref-type="table" rid="table-3">Table 3</xref>, the effectiveness of the IWLSPC algorithm for multi-fault classification of the TE process under covariate shift is detailed. According to <xref ref-type="table" rid="table-3">Table 3</xref>, we can conclude that the IWLSPC achieves better performance for the simultaneous classification for multiple faults in the TE process. Especially in some cases, the accuracy of the IWLSPC is much higher than that of the traditional LSPC, which confirms the effectiveness of IWLSPC. In addition, it is noted that although the BP neural network is extremely effective in the classification of certain types of faults, the classification performance of this method under different conditions has serious fluctuations, resulting in a very poor average performance.</p>

</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>In this paper, the importance-weighted sampling is introduced to address the problem of inconsistent data distribution in training and testing dataset, and it is designed for the fault classification in the multi-mode industrial process, especially for the nonlinear multi-mode industrial process. In the case of covariate shift, it can smoothly adapt to the multi-mode dataset and achieve good performances for fault classification. Besides, we have also confirmed that the introduction of MI for feature selection significantly improves the performance of the fault classifier. This is because the MI can select specific features closely related to these faults. It not only reduces the computational workload, but also gets rid of those variables that are irrelevant to fault classification. In order to address the limitations of the present method, in the future, we will take the ratio of posterior density into consideration, which is calculated in line with the importance-weight estimation and can theoretically break through the assumption that the posterior probability does not change between the training and testing phases. In order to improve the efficiency of feature selection, we will focus on those methods, which can effectively extract feature combinations to improve the performance of negative migration-enhanced classifiers. Besides, in the future, we will try other tools to further improve the efficiency of the optimization.</p>
</sec>
</body>
<back>
<ack><p>None.</p></ack>
<sec><title>Funding Statement</title>
<p>The authors received no funding for this study.</p>
</sec>
<sec><title>Author Contributions</title>
<p>Conceptualization, Methodology, Software, Data Curation, Writing&#x2014;Original Draft Preparation: Yi Pan; Investigation, Validation, Writing&#x2014;Reviewing: Lei Xie; Supervision: Hongye Su. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are available from the corresponding author, Lei Xie, upon reasonable request.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Ge</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Plant-wide industrial process monitoring: A distributed modeling framework</article-title>,&#x201D; <source>IEEE Trans. Ind. Inform.</source>, vol. <volume>12</volume>, no. <issue>1</issue>, pp. <fpage>310</fpage>&#x2013;<lpage>321</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1109/TII.2015.2509247</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Yao</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Ge</surname></string-name></person-group>, &#x201C;<article-title>Industrial big data modeling and monitoring framework for plant-wide processes</article-title>,&#x201D; <source>IEEE Trans. Ind. Inf.</source>, vol. <volume>17</volume>, no. <issue>9</issue>, pp. <fpage>6399</fpage>&#x2013;<lpage>6408</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TII.2020.3010562</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Ge</surname></string-name></person-group>, &#x201C;<article-title>Review on data-driven modeling and monitoring for plant-wide industrial processes</article-title>,&#x201D; <source>Chemometr. Intell. Lab. Syst.</source>, vol. <volume>171</volume>, no. <issue>2</issue>, pp. <fpage>16</fpage>&#x2013;<lpage>25</lpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1016/j.chemolab.2017.09.021</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. H.</given-names> <surname>Raveendran</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>Process monitoring using a generalized probabilistic linear latent variable model</article-title>,&#x201D; <source>Automatica</source>, vol. <volume>96</volume>, no. <issue>7</issue>, pp. <fpage>73</fpage>&#x2013;<lpage>83</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.automatica.2018.06.029</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yan</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>Review and perspectives of data-driven distributed monitoring for industrial plant-wide processes</article-title>,&#x201D; <source>Ind. Eng. Chem. Res.</source>, vol. <volume>58</volume>, no. <issue>29</issue>, pp. <fpage>12899</fpage>&#x2013;<lpage>12912</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1021/acs.iecr.9b02391</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Ge</surname></string-name></person-group>, &#x201C;<article-title>Monitoring and prediction of big process data with deep latent variable models and parallel computing</article-title>,&#x201D; <source>J. Process Control</source>, vol. <volume>92</volume>, no. <issue>11</issue>, pp. <fpage>19</fpage>&#x2013;<lpage>34</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1016/j.jprocont.2020.05.010</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Su</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Mao</surname></string-name></person-group>, &#x201C;<article-title>A probabilistic model with spike-and-slab regularization for inferential fault detection and isolation of industrial processes</article-title>,&#x201D; <source>J. Taiwan Inst. Chem. Eng.</source>, vol. <volume>123</volume>, no. <issue>1</issue>, pp. <fpage>68</fpage>&#x2013;<lpage>78</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1016/j.jtice.2021.05.047</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Deep neural network-embedded stochastic nonlinear state-space models and their applications to process monitoring</article-title>,&#x201D; <source>IEEE Trans. Neural Netw. Learn. Syst.</source>, vol. <volume>33</volume>, no. <issue>12</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yin</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Steven</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Xie</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Luo</surname></string-name></person-group>, &#x201C;<article-title>A review on basic data-driven approaches for industrial process monitoring</article-title>,&#x201D; <source>IEEE Trans. Ind. Electron.</source>, vol. <volume>61</volume>, no. <issue>11</issue>, pp. <fpage>6418</fpage>&#x2013;<lpage>6428</lpage>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.1109/TIE.2014.2301773</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Xie</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Su</surname></string-name></person-group>, &#x201C;<article-title>Deep learning with tensor factorization layers for sequential fault diagnosis and industrial process monitoring</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>105494</fpage>&#x2013;<lpage>105506</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3000004</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Zhong</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Nonlinear multimode industrial process fault detection using modified kernel principal component analysis</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>5</volume>, pp. <fpage>23121</fpage>&#x2013;<lpage>23132</lpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2017.2764518</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Sugiyama</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Kawanabe</surname></string-name></person-group>, <source>Machine Learning in Non-Stationary Environments: Introduction to Covariate Shift Adaptation</source>. <publisher-name>The MIT Press</publisher-name>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. J.</given-names> <surname>Song</surname></string-name> and <string-name><given-names>S. B.</given-names> <surname>Park</surname></string-name></person-group>, &#x201C;<article-title>An adapted surrogate kernel for classification under covariate shift</article-title>,&#x201D; <source>Appl. Soft Comput.</source>, vol. <volume>69</volume>, no. <issue>3</issue>, pp. <fpage>435</fpage>&#x2013;<lpage>442</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.asoc.2018.04.060</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. T.</given-names> <surname>Amin</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Imtiaz</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Ahmed</surname></string-name></person-group>, &#x201C;<article-title>Robust process monitoring methodology for detection and diagnosis of unobservable faults</article-title>,&#x201D; <source>Ind. Eng. Chem. Res.</source>, vol. <volume>58</volume>, no. <issue>41</issue>, pp. <fpage>19149</fpage>&#x2013;<lpage>19165</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1021/acs.iecr.9b03406</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Qi</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A correlation graph based approach for personalized and compatible web APIs recommendation in mobile APP development</article-title>,&#x201D; <source>IEEE Trans. Knowl. Data Eng.</source>, vol. <volume>35</volume>, no. <issue>6</issue>, pp. <fpage>5444</fpage>&#x2013;<lpage>5457</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TKDE.2022.3168611</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. H.</given-names> <surname>Yang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>ASTREAM: Data-stream-driven scalable anomaly detection with accuracy guarantee in IIoT environment</article-title>,&#x201D; <source>IEEE Trans. Netw. Sci. Eng.</source>, vol. <volume>10</volume>, no. <issue>5</issue>, pp. <fpage>3007</fpage>&#x2013;<lpage>3016</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TNSE.2022.3157730</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Privacy-aware traffic flow prediction based on multi-party sensor data with zero trust in smart city</article-title>,&#x201D; <source>ACM Trans. Internet Technol.</source>, vol. <volume>23</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>19</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L. Y.</given-names> <surname>Qi</surname></string-name>, <string-name><given-names>Y. H.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>X. K.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Rafique</surname></string-name>, and <string-name><given-names>J. H.</given-names> <surname>Ma</surname></string-name></person-group>, &#x201C;<article-title>Fast anomaly identification based on multi-aspect data streams for intelligent intrusion detection toward secure Industry 4.0</article-title>,&#x201D; <source>IEEE Trans. Ind. Inform.</source>, vol. <volume>18</volume>, no. <issue>9</issue>, pp. <fpage>6503</fpage>&#x2013;<lpage>6511</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TII.2021.3139363</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. Q.</given-names> <surname>Wu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Popularity-aware and diverse web APIs recommendation based on correlation graph</article-title>,&#x201D; <source>IEEE Trans. Comput. Soc. Syst.</source>, vol. <volume>10</volume>, no. <issue>2</issue>, pp. <fpage>771</fpage>&#x2013;<lpage>782</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TCSS.2022.3168595</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. J.</given-names> <surname>Pan</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>A survey on transfer learning</article-title>,&#x201D; <source>IEEE Trans. Knowl. Data Eng.</source>, vol. <volume>22</volume>, no. <issue>10</issue>, pp. <fpage>1345</fpage>&#x2013;<lpage>1359</lpage>, <year>2010</year>. doi: <pub-id pub-id-type="doi">10.1109/TKDE.2009.191</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Qiao</surname></string-name></person-group>, &#x201C;<article-title>Knowledge-data-driven model predictive control for a class of nonlinear systems</article-title>,&#x201D; <source>IEEE Trans. Syst. Man Cybernet.: Syst.</source>, vol. <volume>51</volume>, no. <issue>7</issue>, pp. <fpage>4492</fpage>&#x2013;<lpage>4504</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Fabric defect detection in textile manufacturing: A survey of the state of the art</article-title>,&#x201D; <source>Secur. Commun. Netw.</source>, vol. <volume>2021</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1155/2024/5034640</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Tailoring density ratio weight for covariate shift adaptation</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>333</volume>, no. <issue>2</issue>, pp. <fpage>135</fpage>&#x2013;<lpage>144</lpage>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2018.11.082</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Shimodaira</surname></string-name></person-group>, &#x201C;<article-title>Improving predictive inference under covariate shift by weighting the log-likelihood function</article-title>,&#x201D; <source>J. Stat. Plan. Infer.</source>, vol. <volume>90</volume>, no. <issue>2</issue>, pp. <fpage>227</fpage>&#x2013;<lpage>244</lpage>, <year>2000</year>. doi: <pub-id pub-id-type="doi">10.1016/S0378-3758(00)00115-4</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Sugiyama</surname></string-name></person-group>, &#x201C;<chapter-title>Learning under non-stationarity: Covariate shift adaptation by importance weighting</chapter-title>,&#x201D; in <source>Handbook of Computational Statistics: Concepts and Methods</source>, <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>, <year>2012</year>, pp. <fpage>927</fpage>&#x2013;<lpage>952</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Tsuboi</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Kashima</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Hido</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bickel</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Sugiyama</surname></string-name></person-group>, &#x201C;<article-title>Direct density ratio estimation for large-scale covariate shift adaptation</article-title>,&#x201D; <source>Inf. Med. Technol.</source>, vol. <volume>4</volume>, no. <issue>2</issue>, pp. <fpage>529</fpage>&#x2013;<lpage>546</lpage>, <year>2009</year>. doi: <pub-id pub-id-type="doi">10.2197/ipsjjip.17.138</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Lu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Zhu etal, HFENet: A lightweight hand-crafted feature enhanced CNN for ceramic tile surface defect detection</article-title>,&#x201D; <source>Int. J. Intell. Syst.</source>, vol. <volume>37</volume>, no. <issue>12</issue>, pp. <fpage>10670</fpage>&#x2013;<lpage>10693</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1002/int.22935</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Niu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Guo</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>A generative adversarial network-based fault detection approach for photovoltaic panel</article-title>,&#x201D; <source>Appl. Sci.</source>, vol. <volume>12</volume>, no. <issue>4</issue>, pp. <fpage>1789</fpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.3390/app12041789</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Ge</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Song</surname></string-name></person-group>, &#x201C;<article-title>Process monitoring based on independent component analysis&#x2212;principal component analysis (ICA&#x2212;PCA) and similarity factors</article-title>,&#x201D; <source>Ind. Eng. Chem. Res.</source>, vol. <volume>46</volume>, no. <issue>7</issue>, pp. <fpage>2054</fpage>&#x2013;<lpage>2063</lpage>, <year>2007</year>. doi: <pub-id pub-id-type="doi">10.1021/ie061083g</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Lin</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Zeng</surname></string-name></person-group>, &#x201C;<article-title>Shrinking principal component analysis for enhanced process monitoring and fault isolation</article-title>,&#x201D; <source>Ind. Eng. Chem. Res.</source>, vol. <volume>52</volume>, no. <issue>49</issue>, pp. <fpage>17475</fpage>&#x2013;<lpage>17486</lpage>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1021/ie401030t</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Wen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Ge</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Song</surname></string-name></person-group>, &#x201C;<article-title>Multimode dynamic process monitoring based on mixture canonical variate analysis model</article-title>,&#x201D; <source>Ind. Eng. Chem. Res.</source>, vol. <volume>54</volume>, no. <issue>5</issue>, pp. <fpage>1605</fpage>&#x2013;<lpage>1614</lpage>, <year>2015</year>. doi: <pub-id pub-id-type="doi">10.1021/ie503324g</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>