<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">57606</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.057606</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Loss Aware Feature Attention Mechanism for Class and Feature Imbalance Issue</article-title>
<alt-title alt-title-type="left-running-head">Loss Aware Feature Attention Mechanism for Class and Feature Imbalance Issue</alt-title>
<alt-title alt-title-type="right-running-head">Loss Aware Feature Attention Mechanism for Class and Feature Imbalance Issue</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Wu</surname><given-names>Yuewei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Fu</surname><given-names>Ruiling</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Xing</surname><given-names>Tongtong</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yin</surname><given-names>Fulian</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><email>yinfulian@cuc.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>College of Information and Communication Engineering, Communication University of China</institution>, <addr-line>Beijing, 100024</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>State Key Laboratory of Media Convergence and Communication, Communication University of China</institution>, <addr-line>Beijing, 100024</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Fulian Yin. Email: <email>yinfulian@cuc.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>03</day><month>1</month><year>2025</year>
</pub-date>
<volume>82</volume>
<issue>1</issue>
<fpage>751</fpage>
<lpage>775</lpage>
<history>
<date date-type="received">
<day>22</day>
<month>8</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>10</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_57606.pdf"></self-uri>
<abstract>
<p>In the Internet era, recommendation systems play a crucial role in helping users find relevant information from large datasets. Class imbalance is known to severely affect data quality, and therefore reduce the performance of recommendation systems. Due to the imbalance, machine learning algorithms tend to classify inputs into the positive (majority) class every time to achieve high prediction accuracy. Imbalance can be categorized such as by features and classes, but most studies consider only class imbalance. In this paper, we propose a recommendation system that can integrate multiple networks to adapt to a large number of imbalanced features and can deal with highly skewed and imbalanced datasets through a loss function. We propose a loss aware feature attention mechanism (LAFAM) to solve the issue of feature imbalance. The network incorporates an attention mechanism and uses multiple sub-networks to classify and learn features. For better results, the network can learn the weights of sub-networks and assign higher weights to important features. We propose suppression loss to address class imbalance, which favors negative loss by penalizing positive loss, and pays more attention to sample points near the decision boundary. Experiments on two large-scale datasets verify that the performance of the proposed system is greatly improved compared to baseline methods.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Imbalanced data</kwd>
<kwd>deep learning</kwd>
<kwd>e-commerce recommendation</kwd>
<kwd>loss function</kwd>
<kwd>big data analysis</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Key Research and Development Program of China</funding-source>
<award-id>2021YFF0901705</award-id>
<award-id>2021YFF0901700</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Recommendation systems (RSs) are increasingly utilized by e-commerce websites to assist consumers in discovering products of interest [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>]. By providing personalized recommendations, RSs can significantly enhance customer engagement and subsequently drive sales [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. An effective recommendation algorithm can increase the profit of an e-commerce platform by up to 20% [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>]. E-commerce websites usually place fast-selling products in an advertising column to promote sales. Consequently, with an equivalent exposure rate, a more accurate prediction will generate higher sales, as well as profits from commission collection methods such as cost per click (CPC) and cost per sale (CPS). Therefore, accurately predicting fast-selling products is important for e-commerce platforms.</p>
<p>Since only 10&#x2013;50 products can be displayed on a webpage, imbalances can occur when an RS is used to predict potential fast-sellers from billions of available products [<xref ref-type="bibr" rid="ref-9">9</xref>]. The prediction accuracy of a traditional RS decreases greatly when the distribution of classes is highly skewed [<xref ref-type="bibr" rid="ref-10">10</xref>]. Because the number of instances in one class can be much smaller than that in another, an instance in a minority class has a strong bias to be classified in the majority class [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>]. Since the overall number of products typically far exceeds the number of hot sale products, these fast-sellers are often overlooked, resulting in high overall accuracy but low precision and F-measure scores.</p>
<p>Based on the input dataset, imbalance in RSs can usually be classified as either class or feature imbalance, and both can diminish performance [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. Class imbalance stems from significant inequality among the number of examples in different classes, which skews predicted results [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>]. Feature imbalance occurs when few features have a significant impact on the result, which dilutes the contributions of important features to the output [<xref ref-type="bibr" rid="ref-17">17</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>]. In other words, the large number of features results in numerous invalid operations, which not only wastes computing resources but also diverts the algorithm&#x2019;s focus from the most critical features.</p>
<p>In the class imbalance issue, the sample can be subdivided into four categories based on the distances of negative and positive instances to the decision boundary: hard negative, easy negative, easy positive, and hard positive, as shown in <xref ref-type="table" rid="table-1">Table 1</xref> [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>]. The sample points are divided into a large positive category and a small negative category. According to the confidence level of the network output (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>y</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula>, greater deviation implies lower confidence), they are divided by whether they are easy or hard to judge. Combining these two dimensions, sample points can be divided into four categories. Hard negative samples (potential fast-selling products we tend to predict) are the most difficult to judge in traditional networks but are of the most concern in an RS. Conversely, easy positive sample points often contribute the most to the loss, since they constitute a large proportion of the training dataset, which diminishes the algorithm&#x2019;s ability to focus on learning the negative sample. Traditional machine learning methods often use preprocessing (including upsampling and downsampling) to address class imbalance [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]. Cost-sensitive algorithms that assume higher costs for misclassification of the minority class [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>] gain more attention since they do not change the original data distribution. However, cost-sensitive algorithms do not consider the confidence level, which represents the degree to which sample points are correctly classified. Sample points that are easily classified correctly (i.e., far from the decision boundary) are more likely to be correctly classified in any network, but those that are difficult to judge (which are close to the decision boundary) are more likely to be misclassified in general networks. Approaches such as focal loss [<xref ref-type="bibr" rid="ref-20">20</xref>] and shrinkage loss [<xref ref-type="bibr" rid="ref-21">21</xref>] have been proposed to solve the two dimensions (number of sample points and confidence level). Shrinkage loss has excellent accuracy, but training stability and computing speed are not considered. To solve these problems, we propose <bold>suppression loss</bold>, which can solve the class imbalance in two dimensions at the same time while also providing a faster training speed and smoother training process.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Classification of sample points in class imbalance issue. <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi mathvariant="bold-italic">f</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi mathvariant="bold-italic">p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is suppression loss, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi mathvariant="bold-italic">&#x03B3;</mml:mi></mml:math></inline-formula> is a penalty parameter of suppression loss</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Hard</th>
<th>Easy</th>
</tr>
</thead>
<tbody>
<tr>
<td>Negative (Minority)</td>
<td><bold>Hard negative</bold></td>
<td>Easy negative (<inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula>)</td>
</tr>
<tr>
<td>Positive (Majority)</td>
<td>Hard positive (<inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>)</td>
<td>Easy positive (<inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula>)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Feature imbalance is usually solved through feature selection and fusion [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>], which are used to select useful features and discard those less helpful for classification [<xref ref-type="bibr" rid="ref-28">28</xref>]. However, features with smaller contributions should also be considered. Discarding some features will affect prediction accuracy to a certain extent [<xref ref-type="bibr" rid="ref-29">29</xref>]. We combine the attention mechanism and a mixture-of-experts (MoE) [<xref ref-type="bibr" rid="ref-30">30</xref>] framework in a <bold>loss-aware feature attention mechanism (LAFAM)</bold>, which can adjust the proportion of each sub-network (with different types of features) by calculating the output confidence level and letting more important features have higher weights, thereby solving the issue of feature imbalance.</p>
<p>We summarize the three major contributions of this paper:
<list list-type="bullet">
<list-item>
<p>We propose suppression loss to address the class imbalance. It can greatly improve the precision of network prediction by penalizing easy and positive classes. Due to the simple form of the loss function, training is stable and fast, and gradient explosion and disappearance occur less than with other loss functions.</p></list-item>
<list-item>
<p>We propose LAFAM for feature imbalance issues. LAFAM fuses multiple sub-networks to learn weights for different features so that important features can contribute more to prediction results without discarding features. To avoid dependence of the network on just one sub-network with excellent performance, total loss weights the loss of each sub-network, and the greater the sub-network loss the greater the contribution to the total loss.</p></list-item>
<list-item>
<p>We propose an RS to recommend potential hot sale products, which solves highly skewed class imbalance and serious feature imbalance issues by applying the proposed suppression loss and LAFAM. A feature preprocessing module sorts and classifies features. We tested the performance of LAFAM on two real datasets covering billions of products with a large amount of data (10 GB, 453 features), and the prediction effect was about 10% higher than that of an existing algorithm.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>We investigate the impact of imbalance on deep learning algorithms and their solutions in <xref ref-type="sec" rid="s2_1">Section 2.1</xref>. In <xref ref-type="sec" rid="s2_2">Section 2.2</xref>, the common structure of RS and the evolution of the algorithms utilized for LAFAM and suppression loss are described.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Imbalance Issue</title>
<p>The issue of class imbalance exists in many areas [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>], and has a great impact on the prediction performance of machine learning algorithms [<xref ref-type="bibr" rid="ref-31">31</xref>&#x2013;<xref ref-type="bibr" rid="ref-33">33</xref>] and a nontrivial impact on the RS [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>]. Class and feature imbalance in e-commerce has attracted the attention of many experts [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>].</p>
<p>Regarding class imbalance, traditional machine learning algorithms often apply preprocessing to increase or reduce the number of positive or negative instances, respectively [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]. To address the problem of sample redundancy or outliers, Wei et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] presented an improved stochastic synthetic minority oversampling technique that applies ascending operations to rank the majority class of samples and assigns weights through kernel density estimation. Hoyos-Osorio et al. [<xref ref-type="bibr" rid="ref-39">39</xref>] proposed an undersampling method for imbalanced data classification based on information-theoretic learning, which selects the most relevant examples from the majority class to enhance classification performance in imbalanced data scenarios. Preprocessing methods, such as oversampling and undersampling, increase computational complexity and cause overfitting by replicating minority classes, or lose information by reducing the size of the majority class [<xref ref-type="bibr" rid="ref-40">40</xref>]. Therefore, Lin et al. [<xref ref-type="bibr" rid="ref-41">41</xref>] investigated the effect of hybrid combinations of undersampling and oversampling methods of different order on 44 different class-imbalanced datasets. Cost-sensitive algorithms that assume higher costs for misclassification of minority classes are gaining attention since they do not change the original data distribution [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>]. Cost-sensitive deep neural networks that learn weights for different classes [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-42">42</xref>&#x2013;<xref ref-type="bibr" rid="ref-44">44</xref>] or employ new loss functions [<xref ref-type="bibr" rid="ref-45">45</xref>] have been proposed. However, current cost-sensitive algorithms do not consider the confidence level, which represents the accuracy that a sample point is correctly classified. Lin et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed focal loss to solve the dense object detection problem, and were the first to mention penalization of easy samples. Lu et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed shrinkage loss, which penalizes easy samples without losing hard samples. However, the complexity of the calculation and stability of the training process should also be considered. We focus on suppression loss to solve the imbalance issue, with faster calculation and stabler training under the premise of better results.</p>
<p>Traditional feature selection and feature fusion are common processing methods to address feature imbalance [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>]. However, feature selection requires the discarding of some features with a low contribution to the result, which will affect the integrity of features [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]. LAFAM can well solve the feature imbalance issue, as discussed in <xref ref-type="sec" rid="s3">Section 3</xref>.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Recommendation Systems</title>
<p>Almost all e-commerce platforms use product recommendation systems [<xref ref-type="bibr" rid="ref-46">46</xref>&#x2013;<xref ref-type="bibr" rid="ref-48">48</xref>], which can often help customers find items of interest and thereby contribute to boosting sales. Traditional recommendation systems can be categorized as content-based recommendation systems (CBRSs), collaborative filtering recommendation systems (CFRSs), and hybrids [<xref ref-type="bibr" rid="ref-49">49</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>]. A CBRS generates preferences based on a user&#x2019;s profile and product features [<xref ref-type="bibr" rid="ref-51">51</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>]. A CFRS uses a similarity matrix to generate preferences based on the rating of neighbor users and items [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-53">53</xref>]. The hybrid framework combines CBRS and CFRS to achieve precise performance by reducing the drawbacks of conventional techniques [<xref ref-type="bibr" rid="ref-54">54</xref>,<xref ref-type="bibr" rid="ref-55">55</xref>]. Although classical RS methods have achieved remarkable success, they suffer from issues such as cold starts and data sparsity [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-56">56</xref>].</p>
<p>With recent deep learning achievements in applications such as natural language processing (NLP), machine translation, and computer vision (CV), machine learning models have been exploited for RSs, bringing more capabilities by addressing the challenges of traditional RS models. Compared to the traditional recommendation architectures, deep learning-based RS models provide better representation learning of user-item interactions. Multi-layer perceptron (MLP) can model data with simple correlation to enhance nonlinear transformation, but high complexity and slow convergence limit its performance [<xref ref-type="bibr" rid="ref-56">56</xref>]. Convolutional neural network (CNN) is powerful for feature extraction of contextual information, but it requires high parameterization tuning [<xref ref-type="bibr" rid="ref-57">57</xref>]. Recurrent neural network (RNN) is specifically used to model sequential data [<xref ref-type="bibr" rid="ref-58">58</xref>]. However, it suffers from the exploding or vanishing gradient, which makes it difficult to train when incorporating temporal layers to capture sequential information. As deep learning models are increasingly adopted in recommendation systems, explainability has become a critical concern for researchers and practitioners [<xref ref-type="bibr" rid="ref-59">59</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>]. Explainable artificial intelligence (XAI) [<xref ref-type="bibr" rid="ref-61">61</xref>] is being progressively integrated into recommendation models to assist users and developers in better understanding how predictions are made.</p>
<p>Since imbalance has a great impact on the prediction results of machine learning algorithms (as described in <xref ref-type="sec" rid="s2_1">Section 2.1</xref>), it should be considered when designing an RS. Hence, we propose LAFAM as a recommendation system, which focuses on solving the impact of imbalance on machine learning algorithms and RSs.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methods</title>
<p>The framework of the proposed RS is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, which has three stages: feature preprocessing, LAFAM training, and classification. The feature preprocessing stage completes feature sorting and classification by three algorithms, replacing feature selection and fusion in traditional methods. LAFAM training solves the feature imbalance issue, allocates suitable networks for different types of features, and completes the self-learning of sub-network weights and feature weights through dynamic adjustment of total loss, enabling important features to make greater contributions to prediction results without discarding features; this cannot be achieved by traditional algorithms. To solve the class imbalance issue, we use suppression loss in the appropriate LAFAM sub-network, which can deal with datasets combining highly skewed class imbalance and serious feature imbalance. The classification stage classifies scores output from LAFAM training to obtain the final classification result. The proposed RS can perform well on large real datasets.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Framework of proposed RS, which has three main parts: (1) Feature preprocessing: Sort and classify features by importance; (2) LAFAM training (core of RS): Adjust weights through supervision network to solve class and feature imbalance issue; (3) Classification: Categorize output scores</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Feature Preprocessing</title>
<p>We first utilize PS-smart [<xref ref-type="bibr" rid="ref-62">62</xref>], XGBoost [<xref ref-type="bibr" rid="ref-63">63</xref>], and GBDT [<xref ref-type="bibr" rid="ref-64">64</xref>] to determine and rank the feature importance. These algorithms can effectively prioritize features that contribute most to model performance, thereby mitigating the potential negative impact of less informative features that could exacerbate feature imbalance. Based on the rankings provided by these three methods, we then calculate the weighted feature importance scores to produce a final ranked list of features. To optimize the output efficiency of network, we truncate the ranked list to discard features with insignificant contributions. The exact truncation position should be adapted to the specific problem at hand, aiming to maximize efficiency and minimize computational resources and time without compromising output quality. This step is particularly useful for addressing feature imbalance, as it reduces the influence of redundant or noisy features, thus helping the model focus on the most relevant information.</p>
<p>Furthermore, we categorize the features into dense features, sparse features, and sequential features to better align with the network operations. Different types of features often exhibit distinct distributional characteristics, and proper categorization helps optimize the handling of imbalanced features. The detailed classification process and corresponding sub-networks will be explained in <xref ref-type="sec" rid="s4_2">Section 4.2</xref>. It is important to note that for addressing different problems, the classification approach should be adjusted according to the types of subsequent sub-networks used. This targeted feature preprocessing approach ensures that imbalanced features are treated in a way that maximizes efficiency and minimizes the risk of performance degradation due to feature imbalances.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>LAFAM</title>
<p>LAFAM can perform independent sub-network training on important features, which are obtained through feature preprocessing, with learning networks depending on features. A supervision network learns the contribution weight of each sub-network. More important features contribute more weight to the result, which increases their influence and improves accuracy. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows the structure of the model, in which different features are input to <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>N</mml:mi></mml:math></inline-formula> sub-networks, and the supervision network learns their output weights.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Structure of loss aware feature attention mechanism (LAFAM)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-2.tif"/>
</fig>
<p>Specifically, we set up several sub-networks to learn the features, and the number and type of sub-networks can be adjusted according to different problems. The sub-networks shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> are of different types. In our model, the output of each network is a score between 0 and 1, which indicates the sample&#x2019;s probability of becoming a hot sale product learning within each respective network. A score of 0 signifies the lowest potential to be a hot sale product, and a score of 1 signifies the highest potential.</p>
<p>For each sub-network, the larger the loss value the lower the confidence of the current output. It is difficult to judge a sub-network based on features with low confidence, and these are usually discarded by the traditional methods. However, any feature will have an impact on the prediction result, and we should not discard a feature that is difficult to judge or makes a small contribution. For example, in e-commerce recommendations, commission, and rebate characteristics are often difficult to judge by the predictive network, but they are important. In our network, confidence means that the network should strengthen the learning of the sub-network, so its loss value accounts for a larger proportion of the loss of the entire network. Therefore, we calculate the softmax value of the loss value output by the network by <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>, and determine the proportion of its contribution to the overall loss. At the end of the network, the weighted sum of the loss values obtains the overall loss.</p>
<p>For a given <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which is the output loss of sub-networks, we estimate the probability <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for each category <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>j</mml:mi></mml:math></inline-formula> with softmax, i.e., we estimate the probability of each classification result of <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula>. Let <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mrow><mml:mover><mml:mi>&#x03B5;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> be a vector of weights of output values from each network. When we obtain the predicted value through the network, we hope that a result with high confidence will dominate. Therefore, in calculating the predicted value, we use the opposite coefficient <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> as the weight of output from sub-networks:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>&#x22EF;</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:math></disp-formula></p>
<p>Then, we design a supervised network whose output is the ratio of each subnetwork output to the total output. The labels used for training are the weight vectors <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mrow><mml:mover><mml:msup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> about <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> obtained after a softmax calculation based on the loss values. We substitute the parameters shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> in the formula, which simplifies the ratio of each category to:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:math></disp-formula></p>
<p>In the training process, we need to fit <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mrow><mml:mover><mml:msup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mover><mml:mi>&#x03B5;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> to make these two probability distributions infinitely close. Therefore, the final loss function <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>b</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mo>&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mover><mml:mi>&#x03B5;</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>b</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mo>&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:msup><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mo>&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mfrac><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mfrac><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>T</mml:mi></mml:mfrac></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mfrac><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>T</mml:mi></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2225;</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">(</mml:mo></mml:mrow></mml:mstyle><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.047em" minsize="2.047em">)</mml:mo></mml:mrow></mml:mstyle><mml:msup><mml:mo stretchy="false">&#x2225;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula> are network parameters that can be set and adjusted using grid search or other tuning methods; <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the final output of the network overall; <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the label of the original sample; <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>b</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> is the output of the sub-network; <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2192;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is the output of the supervision network; <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of feature groups, which in this task is 4; and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>T</mml:mi></mml:math></inline-formula> is the temperature coefficient used in softmax operation, and it can adjust the difference between the values calculated by softmax.</p>
<p>In general, <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>T</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. The smaller <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>T</mml:mi></mml:math></inline-formula> is, the steeper the softmax curve is, and the bigger the difference between the output values will be. The larger <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>T</mml:mi></mml:math></inline-formula> is, the smoother the softmax curve is, and the difference between the output values will be small. In this way, we control the balance of weighting the sub-network loss <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> by the coefficient <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>T</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Suppression Loss</title>
<p>We propose suppression loss to solve the class imbalance issue, which influences machine learning algorithms to classify inputs in the positive (majority) class every time to achieve high prediction accuracy. If we train the deep learning model with traditional loss functions (e.g., square loss, log loss), the loss value of positive samples will account for a larger proportion of the total loss. However, in practice, we are often more concerned with the accuracy of the negative class (precision or F-measure). For example, in our scenario, hot sale products are a minority, but traditional deep learning-based models tend to neglect the minority class and output general products as the prediction result. Therefore, regarding class imbalance, we want negative classes to contribute more loss value. Hence, we propose suppression loss:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03B2;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>2</mml:mn><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03B2;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03C1;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>2</mml:mn><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03C1;</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>]</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>&#x00D7;</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></inline-formula> is the absolute difference between the estimated probability <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and its true label <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>; <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>&#x03C1;</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> are the parameters of the model, which can be set according to the practical dataset problem through the grid search method.</p>
<p>The suppression loss function consists of three parts. The first part is the suppression of the loss contribution to the large category of samples through a function. The second part is to suppress the easy sample contribution by using the same function and different parameters. The last part is the suppression of the easy sample contribution employing a high-power function that expands the output disparity. The feasibility of suppression loss is explained below, and a more detailed process can be found in the Appendix (<xref ref-type="fig" rid="fig-9">Figs. A1</xref> and <xref ref-type="fig" rid="fig-10">A2</xref>).</p>
<p>Firstly, we need a function <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to meet the following conditions:
<list list-type="bullet">
<list-item>
<p>The value range of <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> or <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> when <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> when <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow></mml:math></inline-formula>.</p></list-item>
<list-item>
<p>The function <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is continuous and differentiable.</p></list-item>
</list></p>
<p>We find that the simple function <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula> has the above properties. This is an S-type function, with similar properties to a sigmoid function. However, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is more concise and less complex, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3a</xref>. The function <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03B2;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>2</mml:mn><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>&#x03B2;</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula> is obtained by linear transformation (translation and compression). Its value range is <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and its slope and the point over 0.5 can be adjusted by changing <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-3">Fig. 3a</xref> is the original graph of <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-3">Fig. 3b</xref> is a function diagram for controlling both the independent and dependent variables within the range of 0&#x2013;1 after compression and translation. <xref ref-type="fig" rid="fig-3">Fig. 3c</xref> is a graph of <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> under different parameters, from which it can be seen that the larger <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is, the larger the slope is, i.e., easily classified data can be better penalized. <xref ref-type="fig" rid="fig-3">Fig. 3d</xref> is a comparative diagram of three loss function adjustment factors. When <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>x</mml:mi></mml:math></inline-formula> is closer to 0, the smaller the suppression loss the stronger the penalty of easily classified data. The larger <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>x</mml:mi></mml:math></inline-formula> is, the closer the function value is to 1. The function hardly punishes sufficiently large <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>x</mml:mi></mml:math></inline-formula>. It can be seen from <xref ref-type="fig" rid="fig-3">Fig. 3d</xref> that focal loss still inhibits large <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>x</mml:mi></mml:math></inline-formula> values. The two-parameter control of shrinkage loss is not as accurate as the three-parameter control of suppression loss. Suppression loss has a simpler functional form than shrinkage loss (<inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is simple, but <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is elementary), which will greatly help its stability. We provide verification results in the following experiment.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Proposed process of suppression loss: (a) The original graph of function <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>; (b) The graph of function after compression and translation; (c) The graph of <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> under different parameters; (d) The comparative diagram of three loss function adjustment factors</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-3.tif"/>
</fig>
<p>When applied to classification problems, it is assumed that the label of the positive class is 0. In this case, the closer the predicted value is to 0, the more its loss is penalized, i.e., this kind of sample contributes less to the overall loss. If the label is 1, then the smaller the predicted value the larger its loss, and the less it is penalized in the function, so this kind of sample contributes greatly to the overall loss.</p>
<p>According to the focal loss [<xref ref-type="bibr" rid="ref-20">20</xref>], we superimpose <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> on <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msup><mml:mi>l</mml:mi><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. In this way, we can penalize the sample points of positive and easy classes. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, easy positive sample points are penalized by <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> at the same time, i.e., they are restrained to the greatest extent. Hard positive sample points are penalized by <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Easy negative sample points are penalized by <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula>, while hard negative sample points are barely penalized. This achieves the goal of making hard negative sample points contribute the most loss. The accuracy of the classification of hard negative sample points determines the final loss value of the network. The more mistakes in hard negative sample points, the greater the loss. The network pays more attention to hard negative sample points. Therefore, the classification rate of negative classes can be improved.</p>

</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>We use the Tmall (Taobao) non-public dataset and the Kaggle public dataset Corporaci&#x00F3;n Favorita Grocery Sales Forecasting (CFGSF) for experimental verification.
<list list-type="bullet">
<list-item>
<p><bold>Tmall:</bold> The dataset is 10.7 GB in size and contains 51,134,193 rows of data, with 453 data features from August 2019 to August 2020. The dataset contains information such as product name, product category, product price, coupon price, historical sales volume, highest single-day sales, and single-day average sales.</p></list-item>
<list-item>
<p><bold>CFGSF:</bold> It provides the sales information of 54 stores in different parts of Ecuador from 01 January 2013, to 31 August 2017, including commodity serial number, sales volume, and category; whether goods are easily corrupted; and whether they are in promotion. The dataset al.so provides the store category, city, total sales, and oil price on the day of the sale.</p></list-item>
</list></p>
<p><xref ref-type="table" rid="table-2">Table 2</xref> presents some examples of sales records from the CFGSF dataset. After merging, extracting, cleaning, and screening to ensure that abnormal data and noise would not affect the classification accuracy of the neural network, the dataset contains 70,205,249 rows of data with a dataset size of 13.93 GB. To put the forecasting problem into practice, we transform the forecasting of commodity sales from a regression problem to a classification problem. After normalizing the sales volume of the training set, we select 0.15 as the threshold to define the two classifications. After classification, the data volumes of categories 0 and 1 are 70,184,766 and 20,483, respectively, and the imbalance ratio is 3426:1.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Some examples of sales records in the CFGSF dataset</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Feature</th>
<th>1</th>
<th>2</th>
<th>3</th>
<th>4</th>
<th>5</th>
</tr>
</thead>
<tbody>
<tr>
<td>Date</td>
<td>2013/4/1</td>
<td>2014/12/24</td>
<td>2015/1/2</td>
<td>2016/8/10</td>
<td>2017/4/12</td>
</tr>
<tr>
<td>Store ID</td>
<td>18</td>
<td>39</td>
<td>1</td>
<td>23</td>
<td>30</td>
</tr>
<tr>
<td>Product ID</td>
<td>315463</td>
<td>1473413</td>
<td>220435</td>
<td>1403464</td>
<td>1098624</td>
</tr>
<tr>
<td>Product category</td>
<td>1236</td>
<td>2032</td>
<td>1080</td>
<td>1080</td>
<td>4114</td>
</tr>
<tr>
<td>Perishable</td>
<td>0</td>
<td>1</td>
<td>0</td>
<td>0</td>
<td>0</td>
</tr>
<tr>
<td>Store city</td>
<td>Quito</td>
<td>Cuenca</td>
<td>Quito</td>
<td>Ambato</td>
<td>Guayaquil</td>
</tr>
<tr>
<td>Store category</td>
<td>B</td>
<td>B</td>
<td>D</td>
<td>D</td>
<td>C</td>
</tr>
<tr>
<td>Transactions</td>
<td>1483</td>
<td>2882</td>
<td>1021</td>
<td>980</td>
<td>528</td>
</tr>
<tr>
<td>Oil price</td>
<td>97.1</td>
<td>55.7</td>
<td>52.72</td>
<td>41.75</td>
<td>53.12</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Structure and Parameter Settings of LAFAM</title>
<p>With a simple MoE structure, the importance weight distribution of the softmax gate learning model on experts of different feature groups is shown in <xref ref-type="fig" rid="fig-4">Fig. 4a</xref>. The weight of time series features is high, followed by the importance of commission and voucher features. This is consistent with the conclusion from the PS-smart that the future sales volume of most samples is highly correlated with the sales volume of N days in recent history. <xref ref-type="fig" rid="fig-4">Fig. 4b</xref> shows the learning result based on LAFAM. It has a high learning fit for a large number of simple samples, with high importance of historical time series features and insufficient learning of other features. We observe that the latter&#x2019;s dependence weight distribution on different features is more even so that other basic feature models can also be fully learned, thereby obtaining better recommendation results.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Importance weight distribution of different feature groups based on: (a) Simple mixture-of-experts (MoE) network; (b) Loss aware feature attention mechanism (LAFAM)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-4.tif"/>
</fig>
<p>Therefore, in the process of realizing the model, we divide the features into four parts. The model application structure diagram is shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The basic feature depicts the basic attributes of commodities (including historical prices, historical sales, and other statistical features). The second and third parts are the coupling and commission features, which are the key factors to describe whether goods can become hot sale products. The fourth part is the sequence feature, i.e., the time series of daily transactions of commodities in the past three months. The features are categorized and fed into four distinct sub-networks.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Application structure of our LAFAM</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-5.tif"/>
</fig>
<p>All features are processed using the deep &#x0026; cross network (DCN), a commonly employed approach in recommendation systems, which includes 3 deep layers with ReLU activation functions and 2 cross layers. The cross layers are specifically designed to learn feature interactions in a more efficient manner, capturing bounded-degree interactions between features at different levels. This structure is highly effective for both vertical recommendations and wide-ranging interest exploration, and it offers high computational efficiency, making it an ideal choice for learning from all features. Coupling and commission features are handled using a 3-layer MLP with ReLU activation function, which can extract high-dimensional features from the samples, making it suitable for recommendation predictions. Sequence features are processed using a 2-layer gated recurrent unit (GRU), each layer with 64 hidden units and using the tanh activation function. The GRU is a variant of the long short-term memory (LSTM) network, itself an evolved form of the RNN, known for its excellent predictive performance. Compared to the LSTM, the GRU has a simpler structure and effectively addresses the long-term dependency problem of RNNs. This results in high computational efficiency and excellent accuracy, making the GRU well-suited for learning from sequence features. For the training process, we employ the Adam optimizer with a learning rate of 0.01 and a batch size of 64. The model is trained for 50 epochs, with the suppression loss used to compute the final loss value. The parameter settings for the LAFAM framework are determined based on our own experimental results to ensure optimal performance for the given dataset and task.</p>
<p>Combined with the theoretical derivation and practical application, we improve and optimize the network as follows:
<list list-type="bullet">
<list-item>
<p>We combine the idea of MoE to build an expert network for different feature groups. We combine the expert results through the softmax gate layer and use the loss aware method for learning. In the training process, the output result <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of each expert is calculated with the true value of loss in the training stage, which is recorded as <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, and the distribution of different losses is obtained through a softmax layer. Assuming that each expert learns well and tries to accurately predict the final value <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the distribution of loss represents the importance of different feature groups to the prediction results. Therefore, the difference in attention results of samples in different feature groups is strengthened in the training stage.</p></list-item>
<list-item>
<p>Because we cannot obtain the loss value of each part in the prediction stage, we add <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the training stage to measure the distance between the output probability distributions at A and B in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. We make the probability distributions of two places similar.</p></list-item>
<list-item>
<p>MoE has disadvantages. It cannot guarantee that every sub-network in the framework will do its best to predict. Simply calculating the prediction results according to the smaller the loss the higher the weight will cause the model to gradually abandon the learning of the expert for the larger loss. This causes each expert to not do its best to predict the result, and it cannot reflect the authenticity of each expert&#x2019;s attention result. Therefore, we propose <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>b</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In the process of fusion, we pay more attention to the learning of the higher part of the loss, and control the balance of the softmax result through the temperature parameter <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>T</mml:mi></mml:math></inline-formula>, so that each expert tries best to learn.</p></list-item>
</list></p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation Process and Metrics</title>
<p>We evaluate the performance of the LAFAM network, suppression loss, and the RS by controlling variables, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. We first assess the performance of LAFAM according to regression problems. Since the output of LAFAM is a score (as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>), which belongs to <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, the higher the score, the more likely a product is to become a hot sale product. Then, we use the parameters obtained through a grid search to verify the performance of suppression loss, and we validate the performance of LAFAM and suppression loss from a classification problem perspective. Finally, we verify the effect of the entire RS, i.e., the fusion of LAFAM and suppression loss.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Recommendation system (RS) evaluation process and result guidance</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-6.tif"/>
</fig>
<p>In the experiments, our proposed LAFAM can calculate the probability of a product being a hot sale product. In regression problems, the performance of LAFAM is assessed with four evaluation metrics, which are Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), Weighted Mean Absolute Percentage Error (WMAPE), and Top 50/100/200 Hit Rate (HR@50/100/200). In sorting problems, we use Precision, Accuracy, and F-measure to evaluate the effectiveness of LAFAM and suppression loss.</p>
<p>RMSE presents the gap between the predicted value by the model and the true value:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:msqrt></mml:math></disp-formula>where <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>n</mml:mi></mml:math></inline-formula> is the number of samples, <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the predicted output, and <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the true label.</p>
<p>MAE directly calculates the absolute value of the error between the predicted value and the true value:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>WMAPE weights the prioritized products so as to bias the prediction error towards those products:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>W</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mi>E</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>HR is a recall-based metric, which is defined as:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>H</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>k</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>G</mml:mi><mml:mi>T</mml:mi></mml:math></inline-formula> denotes the number of test sets, and <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>N</mml:mi><mml:mi>k</mml:mi></mml:math></inline-formula> denotes the sum of test sets in each user&#x2019;s Top <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>k</mml:mi></mml:math></inline-formula> recommended products.</p>
<p>Accuracy indicates the percentage of the number of correctly categorized samples to the total number of samples:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> indicates that the final prediction of the positive sample is positive, <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> demonstrates that the final prediction of the negative sample is negative, <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> means that the negative sample ends up with a positive prediction, <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> implies that the positive sample ends up with a negative prediction.</p>
<p>Precision denotes the proportion of correctly categorized positive samples out of all categorized positive samples:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Recall represents how many of all positive samples are correctly categorized by the model. F-measure is the harmonic mean of Precision and Recall:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>F</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>=</mml:mo><mml:mspace width="thinmathspace" /><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>&#x00D7;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x00D7;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is a coefficient regulating the weight of Precision and Recall, typically taken as 1.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Results</title>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Regression Validation of LAFAM</title>
<p>To verify the performance improvement of the LAFAM network, we first compare the model with the following models widely used in recommendation systems in terms of a regression problem:
<list list-type="bullet">
<list-item>
<p>GBDT (Gradient Boosting Decision Tree) [<xref ref-type="bibr" rid="ref-64">64</xref>]: an ensemble learning method that builds and combines many decision trees sequentially to improve predictive performance by minimizing errors.</p></list-item>
<list-item>
<p>DCN (Deep &#x0026; Cross Network) [<xref ref-type="bibr" rid="ref-65">65</xref>]: it retains the benefits of the deep neural network and introduces a novel cross network to learn certain bounded-degree feature interactions more efficiently.</p></list-item>
<list-item>
<p>MoE (Mixture-of-Experts) [<xref ref-type="bibr" rid="ref-30">30</xref>]: it consists of several feed-forward sub-networks and selects a sparse combination of sub-networks via a trainable gated network.</p></list-item>
</list></p>
<p>In this experiment, the squared loss function is used as the basic backpropagation loss function. For GBDT, we use a learning rate of 0.1, a maximum depth of 6, and 100 boosting iterations. For DCN, we follow a standard configuration with 3 deep layers and 2 cross layers, utilizing ReLU as the activation function. For MoE, we configured the model with 4 experts and 2 gating layers, with each expert being an MLP with 3 hidden layers. These settings are selected based on both previous studies and a hyperparameter tuning process conducted in our own experiments to ensure all models were optimized under the same conditions.</p>
<p>On the model side, MAE can intuitively show the regression error situation, RMSE can better reflect the influence of extreme values on the error, while WMAPE is less affected by extreme values and individuals, and can show the overall prediction of the network more evenly. Therefore, the three metrics together can measure the performance of the regression model more completely. However, since the prediction problem in this paper is a class imbalance problem, the metrics above are all calculated with the same weights for both large and small classes, so they cannot well reflect the effectiveness of the network for the imbalance problem. For this reason, we choose the metric HR@K to judge the performance of the model under imbalance conditions in practical application.</p>
<p>The experimental results as regression problems are shown in <xref ref-type="table" rid="table-3">Table 3</xref>. It shows that LAFAM outperforms other network structures on both datasets. In terms of RMSE, MAE, and WMAPE, the LAFAM network has no obvious disadvantage compared with other networks, and even if it cannot reach the best, it is still very close to the best value. Moreover, MAE and WMAPE parameters even occupy the optimal position in the Tmall dataset. RMSE parameters are reduced because large classes contribute many loss values in highly unbalanced datasets. In practical applications, we often measure the performance of a model by HR@K (K &#x003D; 50, 100, 200), the hit rate of the top K products. It can be seen from <xref ref-type="table" rid="table-2">Table 2</xref> that LAFAM improves stability on the HR@K parameter. For the Tmall dataset, LAFAM increases by 34.8%, 8.4%, and 10.4% on HR@50, HR@100 and HR@200, respectively, compared to MoE. For the CGFSF dataset, the improvement is 22.9%, 3.6%, and 24.4%, respectively. This is a great benefit in practical commercial applications. It indicates that LAFAM can optimize the problem of reduced prediction accuracy caused by feature imbalance and class imbalance, and has certain advantages compared to other recommendation applications.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of different networks in regression problems based on two datasets</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Tmall/CFGSF</th>
<th>RMSE</th>
<th>MAE</th>
<th>WMAPE</th>
<th>HR@50</th>
<th>HR@100</th>
<th>HR@200</th>
</tr>
</thead>
<tbody>
<tr>
<td>GBDT</td>
<td>22.583/<bold>0.539</bold></td>
<td>1.294/0.389</td>
<td>0.447/0.402</td>
<td>0.32/0.44</td>
<td>0.37/0.52</td>
<td>0.415/<bold>0.585</bold></td>
</tr>
<tr>
<td>DCN</td>
<td>23.637/0.566</td>
<td>1.185/<bold>0.328</bold></td>
<td>0.494/0.388</td>
<td>0.40/0.51</td>
<td>0.46/0.55</td>
<td>0.450/0.525</td>
</tr>
<tr>
<td>MoE</td>
<td><bold>22.525</bold>/0.574</td>
<td>1.196/0.375</td>
<td>0.452/<bold>0.375</bold></td>
<td>0.46/0.48</td>
<td>0.48/0.47</td>
<td>0.475/0.450</td>
</tr>
<tr>
<td>LAFAM</td>
<td>22.639/0.543</td>
<td><bold>1.171</bold>/0.337</td>
<td><bold>0.433</bold>/0.379</td>
<td><bold>0.62</bold>/<bold>0.59</bold></td>
<td><bold>0.53</bold>/<bold>0.57</bold></td>
<td><bold>0.515</bold>/0.560</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Classification Validation of LAFAM and Suppression Loss</title>
<p>In practical applications, conspicuous positions on the sales page are limited, so we require high prediction accuracy. We can accept the misclassification of small categories into large categories, but the misclassification of large categories into small categories is costly and unacceptable. Therefore, in the experiment, the higher the precision the better the experimental effect. We also refer to the F-measure and accuracy to comprehensively evaluate the network. The F-measure is a typical parameter to measure the effect of a model in an unbalanced data field.</p>
<p>We determine the parameters of suppression loss function by grid search. <xref ref-type="fig" rid="fig-7">Fig. 7</xref> shows the grid search process of suppression loss, where the red dot position is the selected parameter value, and other experiments are conducted in the same way. Since the practical tuning uses high dimensions for the search, the location of the red point in <xref ref-type="fig" rid="fig-7">Fig. 7b</xref> is not the global optimal solution, but it has optimal performance when combined with other parameters. The optimal parameter points are selected as experimental parameters, whose final values are shown in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Suppression loss grid search image: (a) <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi mathvariant="bold-italic">&#x03BB;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi></mml:math></inline-formula>; (b) <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi mathvariant="bold-italic">&#x03B7;</mml:mi></mml:math></inline-formula></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-7.tif"/>
</fig><table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Parameters used in the experiment</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Source formula</th>
<th>Selected value</th>
<th>Setting range</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref></td>
<td>0.60</td>
<td><inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref></td>
<td>0.25</td>
<td><inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref></td>
<td>0.35</td>
<td><inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref></td>
<td>58</td>
<td><inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref></td>
<td>0.69</td>
<td><inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref></td>
<td>50</td>
<td><inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mi>&#x03C1;</mml:mi></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref></td>
<td>0.37</td>
<td><inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula></td>
<td><xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref></td>
<td>2.1</td>
<td><inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>50</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We should ensure the stability of the function as well as its effectiveness. <xref ref-type="fig" rid="fig-8">Fig. 8</xref> shows the change curves of accuracy and loss during training with shrinkage and suppression loss, from which we can find that acc and loss values of suppression loss basically do not fluctuate after reaching a stable level. In addition, there is obviously less gradient disappearance and explosion of suppression loss during training, so its performance is more stable. Stable performance will have higher credibility in practical application, which also benefits recommendation income.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Comparison of training process between shrinkage loss and suppression loss</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-8.tif"/>
</fig>
<p>Since both LAFAM and suppression loss have excellent performance in the control variable test, we use them together to verify the effectiveness of the RS for class and feature imbalance. We examine focal loss, shrinkage loss, and suppression loss respectively under different networks, and select MLP, GRU, DCN, MoE, and LAFAM networks for comparison. The first four are the sub-networks used in the LAFAM network in this paper and thus are tested separately. The parameter settings of the sub-networks used for comparison are kept consistent with the LAFAM network. In this way, we complete the ablation experiment while baseline comparison, proving that the joint network has enhancement compared to each sub-network.</p>
<p><xref ref-type="table" rid="table-5">Table 5</xref> shows the performance of loss functions in different networks based on the imbalanced dataset Tmall. The imbalance ratios of 1:5, 1:10, and 1:20 are selected to represent different levels of data imbalance commonly observed in real-world recommendation system applications. These ratios simulate different degrees of data skewness, allowing us to evaluate the performance of the LAFAM network and suppression loss under both mild (1:5) and extreme (1:20) imbalance conditions. By testing across these imbalance degrees, we ensure the generalizability of our findings across different levels of class distribution imbalances that are relevant in practical applications.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comparison of different networks and loss functions in classification problems under different imbalances based on the Tmall dataset</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>1:5/1:10/1:20</th>
<th>Precision</th>
<th>Accuracy</th>
<th>F-measure</th>
</tr>
</thead>
<tbody>
<tr>
<td>MLP-Focal</td>
<td>0.1678/0.1002/0.0524</td>
<td>0.1736/0.1860/0.1484</td>
<td>0.2874/0.1820/0.0997</td>
</tr>
<tr>
<td>MLP-Shrinkage</td>
<td>0.2909/0.2460/0.1694</td>
<td>0.6355/0.7796/0.8345</td>
<td>0.4304/0.3619/0.2675</td>
</tr>
<tr>
<td>MLP-Suppression</td>
<td>0.3480/<bold>0.3202</bold>/0.2214</td>
<td>0.7216/0.8560/0.8934</td>
<td>0.4791/0.3899/0.2464</td>
</tr>
<tr>
<td>GRU-Focal</td>
<td>0.1336/0.1158/0.1006</td>
<td>0.1662/0.1703/0.1715</td>
<td>0.2523/0.1639/0.1005</td>
</tr>
<tr>
<td>GRU-shrinkage</td>
<td>0.2269/0.2021/0.1859</td>
<td>0.6012/0.6824/0.7219</td>
<td>0.4017/0.3028/0.2771</td>
</tr>
<tr>
<td>GRU-Suppression</td>
<td>0.3520/0.3114/0.2073</td>
<td>0.7033/<bold>0.8623</bold>/0.8982</td>
<td>0.4663/0.3893/0.2215</td>
</tr>
<tr>
<td>DCN-Focal</td>
<td>0.1562/0.0997/0.0312</td>
<td>0.1527/0.1808/0.1879</td>
<td>0.2413/0.1507/0.0958</td>
</tr>
<tr>
<td>DCN-shrinkage</td>
<td>0.2934/0.2185/0.1523</td>
<td>0.6378/0.7209/0.8660</td>
<td>0.4273/0.3469/0.2801</td>
</tr>
<tr>
<td>DCN-Suppression</td>
<td>0.3478/0.3019/0.2115</td>
<td><bold>0.7829</bold>/0.8602/0.9099</td>
<td>0.4870/0.3772/0.2518</td>
</tr>
<tr>
<td>MoE-Focal</td>
<td>0.1473/0.0843/0.0476</td>
<td>0.1455/0.1033/0.1143</td>
<td>0.2376/0.1102/0.0883</td>
</tr>
<tr>
<td>MoE-shrinkage</td>
<td>0.2933/0.2374/0.1940</td>
<td>0.6222/0.7635/0.8878</td>
<td>0.4355/0.3309/0.2541</td>
</tr>
<tr>
<td>MoE-Suppression</td>
<td><bold>0.3561</bold>/0.2912/0.2012</td>
<td>0.6559/0.8243/0.8841</td>
<td><bold>0.5010</bold>/0.3840/0.2954</td>
</tr>
<tr>
<td>LAFAM-Focal</td>
<td>0.1667/0.0923/0.0505</td>
<td>0.1670/0.1024/0.0569</td>
<td>0.2858/0.1691/0.0961</td>
</tr>
<tr>
<td>LAFAM-shrinkage</td>
<td>0.2934/0.2916/0.2127</td>
<td>0.6436/0.8364/0.9003</td>
<td>0.4306/0.3817/0.2695</td>
</tr>
<tr>
<td>LAFAM-Suppression</td>
<td>0.3373/0.3055/<bold>0.2426</bold></td>
<td>0.6755/0.8460/<bold>0.9192</bold></td>
<td>0.4580/<bold>0.3943</bold>/<bold>0.3028</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Regarding the network, LAFAM shows significant improvements on the dataset with a high imbalance degree compared to the other networks. In the 1:20 imbalance dataset, the LAFAM is in the optimal position for all three parameters, which is about 10% higher than the other networks on average. While in the 1:5 and 1:10 imbalance datasets, it is only about 7% lower than the optimal network on average. Therefore, the LAFAM network is used in the class imbalance and feature imbalance datasets with stable results and obvious advantages. Regarding the loss function, it can be found that suppression loss has better performance in each network, as it can well solve the problems of the unbalanced number of categories and unbalanced sample discriminant confidence. The experimental results of the suppression loss have an average increase of about 15% compared to the other loss functions, which shows that it has good adaptability to unbalanced datasets. The LAFAM-suppression loss, which is combined with the two methods, has the best performance under a 1:20 imbalance degree (which is closest to practical application). The performance superiority of LAFAM-suppression loss continues to increase with the imbalance ratio. Therefore, it can be concluded that LAFAM-suppression loss can improve the recommendation of multi-angle imbalance.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion and Future Work</title>
<p>We proposed an RS to solve the recommendation problem on unbalanced datasets. It includes the LAFAM network framework and suppression loss, which can solve the issues of feature and class imbalance, respectively. Combined, they can well improve the imbalance issue encountered in the traditional machine learning algorithm used in the RS field. Comparative experiments using other networks and loss functions on two datasets show that they can solve the imbalance issue. Furthermore, the results on datasets with a high degree of imbalance show greater improvement than traditional methods.</p>
<p>For future work, our proposed LAFAM network requires background data to make recommendations, and thus cannot cope with the cold-start issue. We can optimize this problem by fast trial or interest migration. Alternatively, supervised learning can be combined with reinforcement learning to promote recommendation accuracy, diversity, and vertical category ratio optimization for unbalanced e-commerce data.</p>
</sec>
</body>
<back>
<ack>
<p>We sincerely thank Zhenyu Zhang of Alibaba Group for providing the dataset for this study.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by the National Key Research and Development Program of China (Grant numbers: 2021YFF0901705, 2021YFF0901700); the State Key Laboratory of Media Convergence and Communication, Communication University of China; the Fundamental Research Funds for the Central Universities; the High-Quality and Cutting-Edge Disciplines Construction Project for Universities in Beijing (Internet Information, Communication University of China).</p>
</sec>
<sec><title>Author Contributions</title>
<p>Conceptualization, Yuewei Wu; Methodology, Yuewei Wu, Ruiling Fu; Data curation, Ruiling Fu, Tongtong Xing; Formal analysis, Ruiling Fu, Tongtong Xing; Investigation, Yuewei Wu, Fuliang Yin; Writing&#x2014;original draft, Yuewei Wu, Ruiling Fu. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>Due to the nature of this research, participants of this study did not agree for their data to be shared publicly, so supporting data is not available.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>He</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>A Survey on accuracy-oriented neural recommendation: From collaborative filtering to information-rich recommendation</article-title>,&#x201D; <source>IEEE Trans. Knowl. Data Eng.</source>, vol. <volume>35</volume>, no. <issue>5</issue>, pp. <fpage>4425</fpage>&#x2013;<lpage>4445</lpage>, <year>May 2023</year>. doi: <pub-id pub-id-type="doi">10.1109/TKDE.2022.3145690</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Ullah</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Polat</surname></string-name></person-group>, &#x201C;<article-title>Categorization of knowledge graph based recommendation methods and benchmark datasets from the perspectives of application scenarios: A comprehensive survey</article-title>,&#x201D; <source>Expert Syst. Appl.</source>, vol. <volume>206</volume>, no. <issue>15</issue>, <year>Nov. 2022</year>, Art. no. 117737. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2022.117737</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Ma</surname></string-name></person-group>, &#x201C;<article-title>A Survey on the Fairness of Recommender Systems</article-title>,&#x201D; <source>ACM Trans. Inf. Syst.</source>, vol. <volume>41</volume>, no. <issue>3</issue>, <year>Feb. 2023</year>, Art. no. <fpage>52</fpage>. doi: <pub-id pub-id-type="doi">10.1145/3547333</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Dhelim</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Aung</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Bouras</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Ning</surname></string-name>, and <string-name><given-names>E.</given-names> <surname>Cambria</surname></string-name></person-group>, &#x201C;<article-title>A survey on personality-aware recommendation systems</article-title>,&#x201D; <source>Artif. Intell. Rev.</source>, vol. <volume>55</volume>, no. <issue>3</issue>, pp. <fpage>2409</fpage>&#x2013;<lpage>2454</lpage>, <year>Mar. 2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s10462-021-10063-7</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. E.</given-names> <surname>Bawack</surname></string-name>, <string-name><given-names>S. F.</given-names> <surname>Wamba</surname></string-name>, <string-name><given-names>K. D. A.</given-names> <surname>Carillo</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Akter</surname></string-name></person-group>, &#x201C;<article-title>Artificial intelligence in e-commerce: A bibliometric study and literature review</article-title>,&#x201D; <source>Electr. Mark.</source>, vol. <volume>32</volume>, pp. <fpage>297</fpage>&#x2013;<lpage>338</lpage>, <year>Mar. 2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s12525-022-00537-z</pub-id>; <pub-id pub-id-type="pmid">35600916</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Bai</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Mapping the evolution of e-commerce research through co-word analysis: 2001&#x2013;2020</article-title>,&#x201D; <source>Electr. Commer. R. A.</source>, vol. <volume>55</volume>, no. <issue>2</issue>, <year>Sep.&#x2013;Oct. 2022</year>, Art. no. 101190. doi: <pub-id pub-id-type="doi">10.1016/j.elerap.2022.101190</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Pei</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Value-aware recommendation based on reinforcement profit maximization</article-title>,&#x201D; in <conf-name>World Wide Web Conf. (WWW&#x2019;19)</conf-name>, <publisher-loc>San Francisco, CA, USA</publisher-loc>, <year>May 13&#x2013;17, 2019</year>, pp. <fpage>3123</fpage>&#x2013;<lpage>3129</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Zhou</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Zou</surname></string-name></person-group>, &#x201C;<article-title>Competing for recommendations: The strategic impact of personalized product recommendations in online marketplaces</article-title>,&#x201D; <source>Market. Sci.</source>, vol. <volume>42</volume>, no. <issue>2</issue>, pp. <fpage>360</fpage>&#x2013;<lpage>376</lpage>, <year>Mar.&#x2013;Apr. 2023</year>. doi: <pub-id pub-id-type="doi">10.1287/mksc.2022.1388</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>G&#x00F3;mez</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Boratto</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Salam&#x00F3;</surname></string-name></person-group>, &#x201C;<article-title>Provider fairness across continents in collaborative recommender systems</article-title>,&#x201D; <source>Inform. Process. Manag.</source>, vol. <volume>59</volume>, no. <issue>1</issue>, <year>Jan. 2022</year>, Art. no. 102719. doi: <pub-id pub-id-type="doi">10.1016/j.ipm.2021.102719</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Aguiar</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Krawczyk</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Cano</surname></string-name></person-group>, &#x201C;<article-title>A survey on learning from imbalanced data streams: Taxonomy, challenges, empirical study, and reproducible experimental framework</article-title>,&#x201D; <source>Mach. Learn.</source>, vol. <volume>113</volume>, pp. <fpage>4165</fpage>&#x2013;<lpage>4243</lpage>, <year>Jun. 2024</year>. doi: <pub-id pub-id-type="doi">10.1007/s10994-023-06353-6</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Geng</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Qiang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yuan</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wu</surname></string-name></person-group>, &#x201C;<article-title>Self-adaptive deep asymmetric network for imbalanced recommendation</article-title>,&#x201D; <source>IEEE Trans. Emerg. Top. Comput. Intell.</source>, vol. <volume>8</volume>, no. <issue>1</issue>, pp. <fpage>968</fpage>&#x2013;<lpage>980</lpage>, <year>Feb. 2024</year>. doi: <pub-id pub-id-type="doi">10.1109/TETCI.2023.3300740</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Guo</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Ji</surname></string-name></person-group>, &#x201C;<article-title>An instance-based learning recommendation algorithm of imbalance handling methods</article-title>,&#x201D; <source>Appl. Math. Comput.</source>, vol. <volume>351</volume>, pp. <fpage>204</fpage>&#x2013;<lpage>218</lpage>, <year>Jun. 2019</year>. doi: <pub-id pub-id-type="doi">10.1016/j.amc.2018.12.020</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Yu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A biased sampling method for imbalanced personalized ranking</article-title>,&#x201D; in <conf-name>Proc. 31st ACM Int. Conf. Inform. Knowl. Manag. (CIKM&#x2019;22)</conf-name>, <publisher-loc>Atlanta, GA, USA</publisher-loc>, <year>Oct. 17&#x2013;21, 2022</year>, pp. <fpage>2393</fpage>&#x2013;<lpage>2402</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>J. H.</given-names> <surname>Ryu</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Meta-learning with adaptive weighted loss for imbalanced cold-start recommendation</article-title>,&#x201D; in <conf-name>Proc. 32nd ACM Int. Conf. Inform. Knowl. Manag. (CIKM&#x2019;23)</conf-name>, <publisher-loc>Birmingham, UK</publisher-loc>, <year>Oct. 21&#x2013;25, 2023</year>, pp. <fpage>1077</fpage>&#x2013;<lpage>1086</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Qiao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Duo</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Lin</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Siamese neural networks for user identity linkage through web browsing</article-title>,&#x201D; <source>IEEE Trans. Neural Netw. Learn. Syst.</source>, vol. <volume>31</volume>, no. <issue>8</issue>, pp. <fpage>2741</fpage>&#x2013;<lpage>2751</lpage>, <year>Aug. 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TNNLS.2019.2929575</pub-id>; <pub-id pub-id-type="pmid">31425058</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Oksuz</surname></string-name>, <string-name><given-names>B. C.</given-names> <surname>Cam</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Kalkan</surname></string-name>, and <string-name><given-names>E.</given-names> <surname>Akbas</surname></string-name></person-group>, &#x201C;<article-title>Imbalance problems in object detection: A review</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>43</volume>, no. <issue>10</issue>, pp. <fpage>3388</fpage>&#x2013;<lpage>3415</lpage>, <year>Oct. 2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2020.2981890</pub-id>; <pub-id pub-id-type="pmid">32191882</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>A feature subset selection algorithm automatic recommendation method</article-title>,&#x201D; <source>J. Artif. Intell. Res.</source>, vol. <volume>47</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>34</lpage>, <year>May 2013</year>. doi: <pub-id pub-id-type="doi">10.1613/jair.3831</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>E. H.</given-names> <surname>Han</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Karypis</surname></string-name></person-group>, &#x201C;<article-title>Feature-based recommendation system</article-title>,&#x201D; in <conf-name>Proc. 14th ACM Int. Conf. Inform. Knowl. Manag. (CIKM&#x2019;05)</conf-name>, <publisher-loc>Bremen, Germany</publisher-loc>, <year>Oct. 31&#x2013;Nov. 05, 2015</year>, pp. <fpage>446</fpage>&#x2013;<lpage>452</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Qiao</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Cheng</surname></string-name></person-group>, &#x201C;<article-title>A multi-clustering algorithm to solve driving cycle prediction problems based on unbalanced data sets: A Chinese case study</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>20</volume>, no. <issue>9</issue>, <year>Apr. 2020</year>, Art. no. 2448. doi: <pub-id pub-id-type="doi">10.3390/s20092448</pub-id>; <pub-id pub-id-type="pmid">32344855</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. Y.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Goyal</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name>, <string-name><given-names>K.</given-names> <surname>He</surname></string-name>, and <string-name><given-names>P.</given-names> <surname>Doll&#x00E1;r</surname></string-name></person-group>, &#x201C;<article-title>Focal loss for dense object detection</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>42</volume>, no. <issue>2</issue>, pp. <fpage>318</fpage>&#x2013;<lpage>327</lpage>, <year>Feb. 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2018.2858826</pub-id>; <pub-id pub-id-type="pmid">30040631</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Ni</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Reid</surname></string-name> and <string-name><given-names>M. H.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Deep regression tracking with shrinkage loss</article-title>,&#x201D; in <conf-name>Comput. Vis.-ECCV 2018: 15th European Conf.</conf-name>, <publisher-loc>Munich, Germany</publisher-loc>, <year>Sep. 08&#x2013;14, 2018</year>, pp. <fpage>369</fpage>&#x2013;<lpage>386</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Hong</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Heo</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yun</surname></string-name>, and <string-name><given-names>J. Y.</given-names> <surname>Choi</surname></string-name></person-group>, &#x201C;<article-title>The majority can help the minority: Context-rich minority oversampling for long-tailed classification</article-title>,&#x201D; in <conf-name>Proc. 2022 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR&#x2019;22)</conf-name>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <year>Jun. 18&#x2013;24, 2022</year>, pp. <fpage>6877</fpage>&#x2013;<lpage>6886</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Goyal</surname></string-name></person-group>, &#x201C;<article-title>Handling class-imbalance with KNN (neighbourhood) under-sampling for software defect prediction</article-title>,&#x201D; <source>Artif. Intell. Rev.</source>, vol. <volume>55</volume>, pp. <fpage>2409</fpage>&#x2013;<lpage>2454</lpage>, <year>Mar. 2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s10462-021-10044-w</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Kaur</surname></string-name>, <string-name><given-names>H. S.</given-names> <surname>Pannu</surname></string-name>, and <string-name><given-names>A. K.</given-names> <surname>Malhi</surname></string-name></person-group>, &#x201C;<article-title>A systematic review on imbalanced data challenges in machine learning: Applications and solutions</article-title>,&#x201D; <source>ACM Comput. Surv.</source>, vol. <volume>52</volume>, no. <issue>4</issue>, <year>Aug. 2019</year>, <comment>Art. no. 79</comment>. doi: <pub-id pub-id-type="doi">10.1145/3343440</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. N.</given-names> <surname>Tarekegn</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Giacobini</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Michalak</surname></string-name></person-group>, &#x201C;<article-title>A review of methods for imbalanced multi-label classification</article-title>,&#x201D; <source>Pattern Recogn.</source>, vol. <volume>118</volume>, <year>Oct. 2021</year>, Art. no. 107965. doi: <pub-id pub-id-type="doi">10.1016/j.patcog.2021.107965</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Meng</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Jing</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Yan</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Pedrycz</surname></string-name></person-group>, &#x201C;<article-title>A survey on machine learning for data fusion</article-title>,&#x201D; <source>Inform. Fusion.</source>, vol. <volume>57</volume>, pp. <fpage>115</fpage>&#x2013;<lpage>129</lpage>, <year>May 2022</year>. doi: <pub-id pub-id-type="doi">10.1016/j.inffus.2019.12.001</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ji</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Zheng</surname></string-name></person-group>, &#x201C;<article-title>Spatio-temporal feature fusion for dynamic taxi route recommendation via deep reinforcement learning</article-title>,&#x201D; <source>Knowl.-Based Syst.</source>, vol. <volume>205</volume>, <year>Oct. 2020</year>, Art. no. 106302. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2020.106302</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y. T.</given-names> <surname>Wu</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Predicting implicit user preferences with multimodal feature fusion for similar user recommendation in social media</article-title>,&#x201D; <source>Appl. Sci.</source>, vol. <volume>11</volume>, no. <issue>3</issue>, <year>Jan. 2021</year>, Art. no. 1064. doi: <pub-id pub-id-type="doi">10.3390/app11031064</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Xu</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Nian</surname></string-name></person-group>, &#x201C;<article-title>FFDNN: Feature fusion depth neural network model of recommendation system</article-title>,&#x201D; in <conf-name>Proc. 2020 Int. Conf. Internet of Things Intell. Appl. (ITIA&#x2019;20)</conf-name>, <publisher-loc>Zhenjiang, China</publisher-loc>, <year>Nov. 27&#x2013;29, 2020</year>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Shazeer</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Outrageously large neural networks: The sparsely-gated mixture-of-experts layer</article-title>,&#x201D; <year>2017</year>, <italic>arXiv:1701.06538</italic>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. M.</given-names> <surname>Johnson</surname></string-name> and <string-name><given-names>T. M.</given-names> <surname>Khoshgoftaar</surname></string-name></person-group>, &#x201C;<article-title>Survey on deep learning with class imbalance</article-title>,&#x201D; <source>J. Big Data</source>, vol. <volume>6</volume>, <year>Mar. 2019</year>, Art. no. <fpage>27</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s40537-019-0192-5</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Bria</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Marrocco</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Tortorella</surname></string-name></person-group>, &#x201C;<article-title>Addressing class imbalance in deep learning for small lesion detection on medical images</article-title>,&#x201D; <source>Comput. Biol. Med.</source>, vol. <volume>120</volume>, <year>May 2020</year>, Art. no. 103735. doi: <pub-id pub-id-type="doi">10.1016/j.compbiomed.2020.103735</pub-id>; <pub-id pub-id-type="pmid">32250861</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Cang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Jujube quality grading using a generative adversarial network with an imbalanced data set</article-title>,&#x201D; <source>Biosyst. Eng.</source>, vol. <volume>236</volume>, pp. <fpage>224</fpage>&#x2013;<lpage>237</lpage>, <year>Dec. 2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.biosystemseng.2023.11.002</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>G&#x00F3;mez</surname></string-name></person-group>, &#x201C;<article-title>Characterizing and mitigating the impact of data imbalance for stakeholders in recommender systems</article-title>,&#x201D; in <conf-name>Proc. 14th ACM Conf. Recommender Syst. (RecSys&#x2019;20)</conf-name>, <publisher-loc>Brazil</publisher-loc>, <year>Sep. 22&#x2013;26, 2020</year>, pp. <fpage>756</fpage>&#x2013;<lpage>757</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Qin</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Huang</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Huangfu</surname></string-name></person-group>, &#x201C;<article-title>MSTIL: Multi-cue shape-aware transferable imbalance learning for effective graphic API recommendation</article-title>,&#x201D; <source>J. Syst. Software.</source>, vol. <volume>200</volume>, <year>Jun. 2023</year>, Art. no. 111650. doi: <pub-id pub-id-type="doi">10.1016/j.jss.2023.111650</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Dhote</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Vichoray</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Pais</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Baskar</surname></string-name>, and <string-name><given-names>P.</given-names> <surname>Mohamed Shakeel</surname></string-name></person-group>, &#x201C;<article-title>Hybrid geometric sampling and AdaBoost based deep learning approach for data imbalance in e-commerce</article-title>,&#x201D; <source>Electr. Commer. Res.</source>, vol. <volume>20</volume>, no. <issue>2</issue>, pp. <fpage>259</fpage>&#x2013;<lpage>274</lpage>, <year>Jun. 2020</year>. doi: <pub-id pub-id-type="doi">10.1007/s10660-019-09383-2</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. G.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>A. L.</given-names> <surname>Yee</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>The use of an internet of things data management system using data mining association algorithm in an e-commerce platform</article-title>,&#x201D; <source>J. Organ. End User Com.</source>, vol. <volume>35</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>19</lpage>, <year>May 2023</year>. doi: <pub-id pub-id-type="doi">10.4018/JOEUC.322553</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Mu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Song</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Dou</surname></string-name></person-group>, &#x201C;<article-title>An improved and random synthetic minority oversampling technique for imbalanced data</article-title>,&#x201D; <source>Knowl.-Based Syst.</source>, vol. <volume>248</volume>, <year>Jul. 2022</year>, Art. no. 108839. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2022.108839</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Hoyos-Osorio</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alvarez-Meza</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Daza-Santacoloma</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Orozco-Gutierrez</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Castellanos-Dominguez</surname></string-name></person-group>, &#x201C;<article-title>Relevant information undersampling to support imbalanced data classification</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>436</volume>, pp. <fpage>136</fpage>&#x2013;<lpage>146</lpage>, <year>May 2021</year>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2021.01.033</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Werner de Vargas</surname></string-name>, <string-name><given-names>J. A.</given-names> <surname>Schneider Aranda</surname></string-name>, <string-name><given-names>R.</given-names> <surname>dos Santos Costa</surname></string-name>, <string-name><given-names>P. R.</given-names> <surname>da Silva Pereira</surname></string-name>, and <string-name><given-names>J. L.</given-names> <surname>Vict&#x00F3;ria Barbosa</surname></string-name></person-group>, &#x201C;<article-title>Imbalanced data preprocessing techniques for machine learning: A systematic mapping study</article-title>,&#x201D; <source>Knowl. Inf. Syst.</source>, vol. <volume>65</volume>, no. <issue>1</issue>, pp. <fpage>31</fpage>&#x2013;<lpage>57</lpage>, <year>Jan. 2023</year>. doi: <pub-id pub-id-type="doi">10.1007/s10115-022-01772-8</pub-id>; <pub-id pub-id-type="pmid">36405957</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>C. F.</given-names> <surname>Tsai</surname></string-name>, and <string-name><given-names>W. C.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>Towards hybrid over-and under-sampling combination methods for class imbalanced datasets: An experimental study</article-title>,&#x201D; <source>Artif. Intell. Rev.</source>, vol. <volume>56</volume>, no. <issue>2</issue>, pp. <fpage>845</fpage>&#x2013;<lpage>863</lpage>, <year>Apr. 2022</year>. doi: <pub-id pub-id-type="doi">10.1007/s10462-022-10186-5</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sze-To</surname></string-name> and <string-name><given-names>A. K. C.</given-names> <surname>Wong</surname></string-name></person-group>, &#x201C;<article-title>A weight-selection strategy on training deep neural networks for imbalanced classification</article-title>,&#x201D; in <conf-name>Image Anal. Recognit.: 14th Int. Conf. (ICIAR&#x2019;2017)</conf-name>, <publisher-loc>Montreal, QC, Canada</publisher-loc>, <year>Jul. 05&#x2013;07, 2017</year>, pp. <fpage>3</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Tran</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Mac</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Tong</surname></string-name>, <string-name><given-names>H. A.</given-names> <surname>Tran</surname></string-name>, and <string-name><given-names>L. G.</given-names> <surname>Nguyen</surname></string-name></person-group>, &#x201C;<article-title>A LSTM based framework for handling multiclass imbalance in DGA botnet detection</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>275</volume>, pp. <fpage>2401</fpage>&#x2013;<lpage>2413</lpage>, <year>Jan. 2018</year>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2017.11.018</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Pang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Shao</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Wen</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>DGA-based botnet detection toward imbalanced multiclass learning</article-title>,&#x201D; <source>Tsinghua Sci. Technol.</source>, vol. <volume>26</volume>, no. <issue>4</issue>, pp. <fpage>387</fpage>&#x2013;<lpage>402</lpage>, <year>Aug. 2021</year>. doi: <pub-id pub-id-type="doi">10.26599/TST.2020.9010021</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H. T.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>Advances in cost-sensitive multiclass and multilabel classification</article-title>,&#x201D; in <conf-name>Proc. 25th ACM SIGKDD Int. Conf. Knowl. Discov. Data Min.</conf-name>, <publisher-loc>Anchorage, AK, USA</publisher-loc>, <year>Aug. 04&#x2013;08, 2019</year>, pp. <fpage>3187</fpage>&#x2013;<lpage>3188</lpage>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>An industrial framework for personalized serendipitous recommendation in E-commerce</article-title>,&#x201D; in <conf-name>Proc. 17th ACM Conf. Recommender Syst. (RecSys&#x2019;23)</conf-name>, <publisher-loc>Singapore</publisher-loc>, <year>Sep. 18&#x2013;22, 2023</year>, pp. <fpage>1015</fpage>&#x2013;<lpage>1018</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Islek</surname></string-name> and <string-name><given-names>S. G.</given-names> <surname>Oguducu</surname></string-name></person-group>, &#x201C;<article-title>A hierarchical recommendation system for E-commerce using online user reviews</article-title>,&#x201D; <source>Electr. Commer. R. A</source>, vol. <volume>52</volume>, <year>Mar.&#x2013;Apr. 2022</year>, Art. no. 101131. doi: <pub-id pub-id-type="doi">10.1016/j.elerap.2022.101131</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>An</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Xiao</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Xiao</surname></string-name></person-group>, &#x201C;<article-title>A recommendation model for e-commerce platforms oriented to explicit information compensation and hidden information mining</article-title>,&#x201D; <source>Knowl.-Based Syst.</source>, vol. <volume>286</volume>, <year>Feb. 2024</year>, Art. no. 111359. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2023.111359</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Movafegh</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Rezapour</surname></string-name></person-group>, &#x201C;<article-title>Improving collaborative recommender system using hybrid clustering and optimized singular value decomposition</article-title>,&#x201D; <source>Eng. Appl. Artif. Intel.</source>, vol. <volume>126</volume>, <year>Nov. 2023</year>, Art. no. 107109. doi: <pub-id pub-id-type="doi">10.1016/j.engappai.2023.107109</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Patel</surname></string-name> and <string-name><given-names>H. B.</given-names> <surname>Patel</surname></string-name></person-group>, &#x201C;<article-title>A state-of-the-art survey on recommendation system and prospective extensions</article-title>,&#x201D; <source>Comput. Electron. Agr.</source>, vol. <volume>178</volume>, <year>Nov. 2020</year>, Art. no. 105779. doi: <pub-id pub-id-type="doi">10.1016/j.compag.2020.105779</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>P&#x00E9;rez-Almaguer</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Yera</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Alzahrani</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Mart&#x00ED;nez</surname></string-name></person-group>, &#x201C;<article-title>Content-based group recommender systems: A general taxonomy and further improvements</article-title>,&#x201D; <source>Expert Syst. Appl.</source>, vol. <volume>184</volume>, <year>Dec. 2021</year>, Art. no. 115444. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2021.115444</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Lu</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Duan</surname></string-name></person-group>, &#x201C;<article-title>Online content-based sequential recommendation considering multimodal contrastive representation and dynamic preferences</article-title>,&#x201D; <source>Neural Comput. Appl.</source>, vol. <volume>36</volume>, no. <issue>13</issue>, pp. <fpage>7085</fpage>&#x2013;<lpage>7103</lpage>, <year>Feb. 2024</year>. doi: <pub-id pub-id-type="doi">10.1007/s00521-024-09447-x</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Koren</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Rendle</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Bell</surname></string-name></person-group>, &#x201C;<article-title>Advances in collaborative filtering</article-title>,&#x201D; in <source>Recommender Syst. Handbook</source>, <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Springer</publisher-name>, <year>Nov. 2021</year>, pp. <fpage>91</fpage>&#x2013;<lpage>142</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-1-0716-2197-4_3</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Sharma</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Rana</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Malhotra</surname></string-name></person-group>, &#x201C;<article-title>Automatic recommendation system based on hybrid filtering algorithm</article-title>,&#x201D; <source>Educ. Inf. Technol.</source>, vol. <volume>27</volume>, pp. <fpage>1523</fpage>&#x2013;<lpage>1538</lpage>, <year>Jul. 2021</year>. doi: <pub-id pub-id-type="doi">10.1007/s10639-021-10643-8</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. Z.</given-names> <surname>Darban</surname></string-name> and <string-name><given-names>M. H.</given-names> <surname>Valipour</surname></string-name></person-group>, &#x201C;<article-title>GHRS: Graph-based hybrid recommendation system with application to movie recommendation</article-title>,&#x201D; <source>Expert Syst. Appl.</source>, vol. <volume>200</volume>, <year>Aug. 2022</year>, Art. no. 116850. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2022.116850</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Da&#x2019;u</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Salim</surname></string-name></person-group>, &#x201C;<article-title>Recommendation system based on deep learning methods: A systematic review and new directions</article-title>,&#x201D; <source>Artif. Intell. Rev.</source>, vol. <volume>53</volume>, pp. <fpage>2709</fpage>&#x2013;<lpage>2748</lpage>, <year>Apr. 2020</year>. doi: <pub-id pub-id-type="doi">10.1007/s10462-019-09744-1</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. M.</given-names> <surname>Islam</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Nooruddin</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Karray</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Muhammad</surname></string-name></person-group>, &#x201C;<article-title>Human activity recognition using tools of convolutional neural networks: A state of the art review, data sets, challenges, and future prospects</article-title>,&#x201D; <source>Comput. Biol. Med.</source>, vol. <volume>149</volume>, <year>Oct. 2022</year>, Art. no. 106060. doi: <pub-id pub-id-type="doi">10.1016/j.compbiomed.2022.106060</pub-id>; <pub-id pub-id-type="pmid">36084382</pub-id></mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Hewamalage</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Bergmeir</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Bandara</surname></string-name></person-group>, &#x201C;<article-title>Recurrent neural networks for time series forecasting: Current status and future directions</article-title>,&#x201D; <source>Int. J. Forecasting.</source>, vol. <volume>37</volume>, no. <issue>1</issue>, pp. <fpage>388</fpage>&#x2013;<lpage>427</lpage>, <year>Jan.&#x2013;Mar. 2021</year>. doi: <pub-id pub-id-type="doi">10.1016/j.ijforecast.2020.06.008</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. A.</given-names> <surname>Chatti</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Guesmi</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Muslim</surname></string-name></person-group>, &#x201C;<article-title>Visualization for recommendation explainability: A survey and new perspectives</article-title>,&#x201D; <source>ACM T. Interact. Intel.</source>, vol. <volume>14</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>40</lpage>, <year>Aug. 2024</year>. doi: <pub-id pub-id-type="doi">10.1145/3672276</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>A. D.</given-names> <surname>Tian</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>When post hoc explanation knocks: Consumer responses to explainable AI recommendations</article-title>,&#x201D; <source>J. Interact. Mark.</source>, vol. <volume>59</volume>, no. <issue>3</issue>, pp. <fpage>234</fpage>&#x2013;<lpage>250</lpage>, <year>Dec. 2023</year>. doi: <pub-id pub-id-type="doi">10.1177/10949968231200221</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Dwivedi</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Explainable AI (XAI): Core ideas, techniques, and solutions</article-title>,&#x201D; <source>ACM Comput. Surv.</source>, vol. <volume>55</volume>, no. <issue>9</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>33</lpage>, <year>Jan. 2023</year>. doi: <pub-id pub-id-type="doi">10.1145/3561048</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Scaling distributed machine learning with the parameter server</article-title>,&#x201D; in <conf-name>Proc. 11th USENIX Conf. Oper. Syst. Design Implementation</conf-name>, <publisher-loc>Broomfield, CO</publisher-loc>, <year>Oct. 06&#x2013;08, 2014</year>, pp. <fpage>583</fpage>&#x2013;<lpage>598</lpage>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Guestrin</surname></string-name></person-group>, &#x201C;<article-title>XGBoost: A scalable tree boosting system</article-title>,&#x201D; in <conf-name>Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discov. Data Min.</conf-name>, <publisher-loc>San Francisco, CA, USA</publisher-loc>, <year>Aug. 13&#x2013;17, 2016</year>, pp. <fpage>785</fpage>&#x2013;<lpage>794</lpage>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. H.</given-names> <surname>Friedman</surname></string-name></person-group>, &#x201C;<article-title>Greedy function approximation: A gradient boosting machine</article-title>,&#x201D; <source>Ann. Stat.</source>, vol. <volume>29</volume>, no. <issue>5</issue>, pp. <fpage>1189</fpage>&#x2013;<lpage>1232</lpage>, <year>Oct. 2001</year>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Fu</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Deep &#x0026; cross network for ad click predictions</article-title>,&#x201D; in <conf-name>Proc. ADKDD&#x2019;17</conf-name>, <publisher-loc>Halifax, NS, Canada</publisher-loc>, <year>Aug. 14, 2017</year>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
</ref-list>
<app-group>
<app id="app-1">
<title>Appendix A. Suppression Loss Feasibility Proof</title>
<p>A loss function must be continuous and differentiable in <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:math></inline-formula>. Therefore, we will prove the continuity and differentiability of <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which requires us to prove that <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula> is continuous and differentiable.</p>
<p><inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is a piecewise function. When <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mi>x</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula>, and <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is obviously continuous and differentiable. When <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula>, and again, <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is obviously continuous and differentiable. Therefore, we only need to prove that <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is continuous and differentiable at <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>.</p>
<p>We first prove that <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is differentiable at point 0.
<disp-formula id="eqn-A1"><label>(A1)</label><mml:math id="mml-eqn-A1" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-A2"><label>(A2)</label><mml:math id="mml-eqn-A2" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-A3"><label>(A3)</label><mml:math id="mml-eqn-A3" display="block"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-A4"><label>(A4)</label><mml:math id="mml-eqn-A4" display="block"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>When <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Therefore, <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is differentiable in the real domain.</p>
<p>We next prove that <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is continuous at 0.
<disp-formula id="eqn-A5"><label>(A5)</label><mml:math id="mml-eqn-A5" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>x</mml:mi></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></disp-formula>
<disp-formula id="eqn-A6"><label>(A6)</label><mml:math id="mml-eqn-A6" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>x</mml:mi></mml:mfrac></mml:mstyle><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></disp-formula>
<disp-formula id="eqn-A7"><label>(A7)</label><mml:math id="mml-eqn-A7" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Therefore, <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is continuous in the domain of <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:math></inline-formula>.</p>
<p>Thus <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is continuous and differentiable in the domain of <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:math></inline-formula>, so <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can be used as a loss function.</p>
</app>
<app id="app-2">
<title>Appendix B. Experiment Supplementary Diagram</title>
<fig id="fig-9">
<label>Figure A1</label>
<caption>
<title>Receiver operating characteristic (ROC) curve of shrinkage and suppression loss. The ROC curve is generally used to reflect the threshold sensitivity and prediction accuracy of a model. The value of AUC represents the area under the ROC curve. The larger the area the better the prediction effect. We can see that the AUC value of the suppression loss is greater than that of shrinkage, indicating that suppression loss has better predictive performance</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-9.tif"/>
</fig><fig id="fig-10">
<label>Figure A2</label>
<caption>
<title>Distribution of predicted score by shrinkage loss and suppression loss. In the scoring process, we hope that the scores are concentrated at both ends, i.e., there are many sample points with low and high scores, but the scores are relatively small. A more concentrated number of scores, at a certain value (such as (a)), means that the algorithm does not distinguish the scores. It demonstrates that suppression loss is better than shrinkage loss and focal loss in classification performance</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-10a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57606-fig-10b.tif"/>
</fig>
</app>
</app-group>
</back></article>