<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">IASC</journal-id>
<journal-id journal-id-type="nlm-ta">IASC</journal-id>
<journal-id journal-id-type="publisher-id">IASC</journal-id>
<journal-title-group>
<journal-title>Intelligent Automation &#x0026; Soft Computing</journal-title>
</journal-title-group>
<issn pub-type="epub">2326-005X</issn>
<issn pub-type="ppub">1079-8587</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">41873</article-id>
<article-id pub-id-type="doi">10.32604/iasc.2023.041873</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Learning Model for News Quality Evaluation Based on Explicit and Implicit Information</article-title>
<alt-title alt-title-type="left-running-head">Deep Learning Model for News Quality Evaluation Based on Explicit and Implicit Information</alt-title>
<alt-title alt-title-type="right-running-head">Deep Learning Model for News Quality Evaluation Based on Explicit and Implicit Information</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Song</surname><given-names>Guohui</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Wang</surname><given-names>Yongbin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>ybwang@cuc.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Jianfei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Hu</surname><given-names>Hongbin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>State Key Laboratory of Media Convergence and Communication, Communication University of China</institution>, <addr-line>Beijing, 100024</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Computer and Cyber Sciences, Communication University of China</institution>, <addr-line>Beijing, 100024</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yongbin Wang. Email: <email>ybwang@cuc.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic"><year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>27</day><month>2</month><year>2024</year></pub-date>
<volume>38</volume>
<issue>3</issue>
<fpage>275</fpage>
<lpage>295</lpage>
<history>
<date date-type="received"><day>09</day><month>5</month><year>2023</year></date>
<date date-type="accepted"><day>19</day><month>7</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Song et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Song et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_IASC_41873.pdf"></self-uri>
<abstract>
<p>Recommending high-quality news to users is vital in improving user stickiness and news platforms&#x2019; reputation. However, existing news quality evaluation methods, such as clickbait detection and popularity prediction, are challenging to reflect news quality comprehensively and concisely. This paper defines news quality as the ability of news articles to elicit clicks and comments from users, which represents whether the news article can attract widespread attention and discussion. Based on the above definition, this paper first presents a straightforward method to measure news quality based on the comments and clicks of news and defines four news quality indicators. Then, the dataset can be labeled automatically by the method. Next, this paper proposes a deep learning model that integrates explicit and implicit news information for news quality evaluation (EINQ). The explicit information includes the headline, source, and publishing time of the news, which attracts users to click. The implicit information refers to the news article&#x2019;s content which attracts users to comment. The implicit and explicit information affect users&#x2019; click and comment behavior differently. For modeling explicit information, the typical convolution neural network (CNN) is used to get news headline semantic representation. For modeling implicit information, a hierarchical attention network (HAN) is exploited to extract news content semantic representation while using the latent Dirichlet allocation (LDA) model to get the subject distribution of news as a semantic supplement. Considering the different roles of explicit and implicit information for quality evaluation, the EINQ exploits an attention layer to fuse them dynamically. The proposed model yields the Accuracy of 82.31&#x0025; and the F-Score of 80.51&#x0025; on the real-world dataset from Toutiao, which shows the effectiveness of explicit and implicit information dynamic fusion and demonstrates performance improvements over a variety of baseline models in news quality evaluation. This work provides empirical evidence for explicit and implicit factors in news quality evaluation and a new idea for news quality evaluation.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Deep learning</kwd>
<kwd>news quality</kwd>
<kwd>communication studies</kwd>
<kwd>classification</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Fundamental Research Funds for the Central Universities</funding-source>
<award-id>CUC230B008</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>With the rapid development of internet technology, news platforms based on personalized recommendation algorithms, such as Google News and Toutiao, have gradually emerged. These platforms aggregate the news content provided by online content service providers, rapidly change the way of information dissemination, and progressively become the primary way for users to obtain information. Given the growing amount of Internet traffic, news platforms often rely on personalized recommendation methods and popularity prediction methods to allocate resources better to meet the evolving information needs of readers. Therefore, how to identify high-quality news that attracts users has become a hot spot in current research.</p>
<p>At present, deep learning technology has achieved great success in information recommendation and has been widely used in various fields, such as medicine [<xref ref-type="bibr" rid="ref-1">1</xref>], banking [<xref ref-type="bibr" rid="ref-2">2</xref>], law [<xref ref-type="bibr" rid="ref-3">3</xref>], education [<xref ref-type="bibr" rid="ref-4">4</xref>], biology [<xref ref-type="bibr" rid="ref-5">5</xref>], and insurance [<xref ref-type="bibr" rid="ref-6">6</xref>]. In news communication, news platforms often use deep learning techniques to recommend potentially widespread information to users. These methods are mainly based on the user&#x2019;s historical behavior to cater to the preferences of the user [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-9">9</xref>] or push potentially high-click news articles to the user [<xref ref-type="bibr" rid="ref-10">10</xref>&#x2013;<xref ref-type="bibr" rid="ref-12">12</xref>]. However, these studies focus on optimizing news clicks or catering to users&#x2019; preferences while ignoring the quality of the articles they recommend, which can lead to under-recommended content news. Human editors or automated quality inspection systems may be unable to detect some low-quality information on news platforms. These low-quality news articles may not gain users&#x2019; attention and discussion, and regularly recommending them seriously damage the user experience and the competitiveness of online news platforms. Therefore, news platforms need a quality evaluation mechanism that can evaluate the quality of news before it is published, which helps the news platform push information that can attract users&#x2019; attention and increase the competitiveness and engagement of the news platform.</p>
<p>News quality evaluation methods can divide into news popularity prediction [<xref ref-type="bibr" rid="ref-10">10</xref>&#x2013;<xref ref-type="bibr" rid="ref-12">12</xref>] and clickbait detection [<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-16">16</xref>]. Clickbait detection is often considered a classification problem to determine whether a headline is misleading. News articles with misleading headlines are considered low-quality because their contents do not satisfy users&#x2019; expectations. However, clickbait detection methods cannot detect high-quality news because not all non-clickbait news articles are of high quality since their content may not catch users&#x2019; attention. The researches focus on news popularity prediction mainly to predict the popularity of news based on the number of news readers have viewed. Most of these studies considered news popularity prediction a classification or regression problem. Unpopular news articles are regarded as low quality since their inability to attract users to click. However, high-popularity news may be due to catchy headlines whose content does not attract attention. Therefore, neither clickbait detection nor popularity prediction can effectively assess the quality of news. In addition, some studies consider the dwell time of news reading as an indicator of the quality of the news [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]. But the dwell time of news reading may be closely related to the length of the news content. At the same time, neither news popularity prediction nor clickbait detection can reflect whether the news has received widespread public attention and discussion. Therefore, defining the quality of news is an open issue that needs to be combined with practical application scenarios. In addition, there are many difficulties in assessing the quality of news. First, there are currently no datasets with news-quality labels. Secondly, given the labeled data, how to use the various features of news to design models to predict the quality of news is another vital issue accurately. Thirdly, the research object of news quality evaluation is mainly aimed at news headlines, measuring the relationship between news headlines and contents. The role of news headlines, content, and other features in evaluating news quality has not been well studied.</p>
<p>To address the above limitations of existing research, this study proposes a deep learning model for online news quality evaluation, named EINQ, based on explicit and implicit news information. This study proposes a straightforward method to measure news quality based on the comments and clicks of each piece of news. The number of clicks represents the breadth of news dissemination, while the number of comments is more reflective of the publicity effect of the news. Therefore, the quality of the news in this study mainly considers the ability of the news headlines, content, and other features to attract users to click and comment. According to this method, this paper defines four quality indicators to evaluate news quality. The explicit information includes the headline, source, and publishing time of the news, which users can see before they click. The implicit information refers to the news article&#x2019;s content and is the main factor that triggers user comments. The implicit and explicit information affect users&#x2019; click and comment behavior differently. After obtaining the embedded representation of these two types of information, we dynamically fuse them through the attention mechanism for evaluating news quality.</p>
<p>In summary, the main contributions of this paper are as follows:</p>
<p>(1) This work proposes the concepts and indicators of news quality. To the best of our knowledge, this is the first attempt to incorporate the comments and clicks of news in news quality evaluation to evaluate the ability of news articles to elicit clicks and comments from users.</p>
<p>(2) This work proposes a deep learning model, EINQ, that integrates explicit and implicit news information for news quality evaluation. More importantly, EINQ incorporates an attention unit based on the implicit and explicit information of the news article, which can adaptively learn their combinations for classification tasks.</p>
<p>(3) This study verifies the effectiveness of EINQ as a whole, provides empirical evidence for explicit and implicit factors in news quality evaluation, and provides a new idea for news quality evaluation.</p>
<p>The rest of this paper is organized as follows. The next section summarizes the current work related to this study. <xref ref-type="sec" rid="s3">Section 3</xref> defines four quality indicators to evaluate news quality and describes ways to automate data labeling. <xref ref-type="sec" rid="s4">Section 4</xref> formally introduces the EINQ model for evaluating news quality. <xref ref-type="sec" rid="s5">Section 5</xref> presents the experimental results and detailed analysis. <xref ref-type="sec" rid="s6">Section 6</xref> finally concludes this paper.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Work</title>
<p>News quality evaluation is widely used in clickbait detection, and its original research intention is to detect some online news outlets trying to attract users to click through gimmicky headlines. Clickbait detection is generally defined as a binary classification problem that aims to determine whether a news headline is misleading. There is a large amount of research related to clickbait detection in academia. Kaur et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] proposed a two-phase hybrid convolutional neural network (CNN) and long short-time memory network (LSTM) models for modeling short topic content. The experiment showed that clickbait such as Shocking/Unbelievable, Hypothesis/Guess, and Reaction are the highest in numbers among the other clickbait headlines published online. Indurthi et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] modeled clickbait strength prediction as a regression problem. While previous methods have relied on traditional machine learning or recurrent neural networks (RNN), this study rigorously investigates using a transformer regression model for clickbait strength prediction. The experiment results with a benchmark dataset result in a new state-of-the-art for the clickbait intensity prediction task. Since previous researches usually focus on the semantic information of the English clickbait corpus, some studies have begun to focus on the Chinese clickbait problem. Liu et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] constructed a Chinese WeChat clickbait dataset. They proposed an effective deep method by integrating semantic, syntactic, and auxiliary information, which respectively use Bidirectional Encoder Representation from Transformers (BERT) and Bidirectional Long Short-Term Memory (Bi-LSTM) network with an attention mechanism to encode title semantics. In addition, they propose an improved Graph Attention Network (GAT) to aggregate local syntactic structures of titles and use an attention mechanism to capture valuable structures. The experiment results prove that model performance is better than compared baseline methods. Pujahari et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed a hybrid categorization technique for separating clickbait and non-clickbait articles by combining features, sentence structure, and clustering. The experimental results indicate that the proposed hybrid model is more robust, reliable, and efficient than any particular categorization technique. Some studies suggested that the use of clickbait can enhance the effectiveness of official propaganda. Meanwhile, some works focused on detecting fake news, considered low-quality due to the spread of false information. Zhang et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] proposed a deep learning-based fast fake news detection model for cyber-physical social services, which adopted a convolution-based neural computing framework to extract feature representation for Chinese news texts. Such a design can ensure both processing speed and detection ability in scenes of Chinese short texts. The experiment results show that the proposal has a lower training time cost and higher classification accuracy than baseline methods. Some studies use the credibility of news to rate the quality of news. Romanou et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed a methodology for creating scientific news article representations by modeling the directed graph between the scientific news articles and the cited scientific publications. The network used for the experiments comprises scientific news articles, their topic, the cited research literature, and corresponding authors. The experiment results show promising applications of graph neural network approaches in knowledge tracing and scientific news credibility assessment. Mosallanezhad et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] incorporated auxiliary information (e.g., user comments and user-news interactions) into a reinforcement learning-based model to address the diverse nature of news domains and expensive annotation costs. Extensive experiments on real-world datasets illustrate the effectiveness of the proposed model, especially when limited labeled data is available in the target domain.</p>
<p>On the other hand, some studies predict the quality of news before the news is released based on news headlines or content. The study mainly indicates the popularity of news based on the number of clicks, and news with low clicks is considered low-quality. Xiong et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] believed that attractiveness and timeliness of news are two prominent drivers of news clicks and proposed a deep news click prediction method to integrate attractiveness, timeliness, and text of news for news click prediction. The experiment results show the effectiveness of attractiveness and timeliness in click-through prediction. Saeed et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] predicted the popularity of a news item published on a specific website by exploiting the initial tweeting behavior of the news item on Twitter. The temporal characteristics of a news item are exploited as the news propagates via tweets. The experiment results show the effectiveness of temporal propagation patterns in predicting news popularity. Sun et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] proposed an explicit Time embedding-based Cascade Attention Network (TCAN) as a novel popularity prediction architecture for large-scale information networks, which integrates temporal attributes into node features via an available time embedding approach to learn the representation of cascade graphs and cascade sequences. The experiment results show that TCAN outperforms other representative baselines while maintaining good interpretability. Wu et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed an effective news quality evaluation method based on the distributions of users&#x2019; reading dwell time on the news, which incorporate news quality information into user interest modeling by designing a content-quality attention network to select clicked news based on news semantics and qualities. Extensive experiments on two real-world datasets show that the model can effectively improve the overall quality of recommended news and reduce the recommendation of low-quality news. Omidvar et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] believed that the quality of headlines is determined by the number of clicks on the news and the time that users stay on the news content, and proposed a deep learning model for evaluating the quality of news headlines by four indicators. Stokowiec et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] believed that the quality of news headlines is the primary factor that triggers user clicks, so they only used Bi-LSTM to predict the number of clicks on the news. The experiment results show that using pre-trained word vectors in the embedding layer improves the results of LSTM models, especially when the training set is small. Weissburg et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed an interpretable attention-based model to assess the impact of a post&#x2019;s title on its popularity while controlling for the time of posting. The experiment results show that the title&#x2019;s emotion, narrative ability, and political power are essential in promoting the post&#x2019;s popularity.</p>
<p>In addition, some studies attempted to assess news quality from the feature extraction perspective. Rieis et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] extracted features from the content of 69,907 news articles to find ways to improve the quality of headlines. The findings suggest that positive or negative headlines attract more clicks, while neutral headlines get fewer clicks. Kim et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed a novel generative model, the headline click-based topic model (HCTM), that extends latent Dirichlet allocation (LDA) to reveal the effect of topical context on the click-value of words in headlines. The experiment results show that by jointly taking topics and clicks into account, the model can detect changes in user interests within topics. Wu et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] designed an encoder-decoder model using a large-scale training corpus to generate high-quality news headlines by reducing the importance of easy-to-classify sentences. In this part of the study, the characteristics affecting the quality of news headlines are identified by multiple feature extraction methods. Experimental results show that this method significantly improves the quality of headline editing compared to previous methods. Zhou et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] proposed a joint predictor for both users&#x2019; click and dwell time in session-based settings, which adopts a novel three-layered RNN structure that includes a fast-slow layer for very short sessions and an attention layer for noisy sessions. Experiments demonstrate that this model outperforms state-of-the-art user clicks and dwell time prediction methods. Romanou et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] constructed a novel system for evaluating the quality of news articles which automatically collects contextual information about news articles in real time and provides quality indicators about their validity and trustworthiness. The news&#x2019; quality indicators include social media discussions, news content, and source. Choi et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] employed a traditional social science survey of over 7800 news audiences and implemented natural language processing for 1500 news articles concerning public affairs. Results suggest believability, depth, and diversity are more important in predicting news quality than readability, objectivity, factuality, and sensationalism. Alam et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] proposed a score-based technique to quantify the quality of online news articles through news article metadata, content analysis, and epidemiological entity extraction. The experiment results show significant enhancement in the data quality by filtering irrelevant news.</p>
<p>From the above-related works, we can see different definitions of news quality in practical applications. Clickbait and fake news detection methods focus on identifying low-quality news headlines that are gimmicky or misleading. However, these methods fail to identify high-quality news that attracts users. News popularity prediction methods mainly indicate the popularity of news based on the number of clicks, and news with low clicks is considered low-quality. However, user click behavior is mainly triggered by the attractiveness of news headlines. The news popularity prediction methodology does not consider whether the news content attracts users and sparks widespread discussion among users. The methods based on feature extraction are mainly used to identify the features that affect the quality of news, such as words, sentences, emotions, etc. However, the interaction between news headlines, content, and other metadata features in evaluating news quality has not been thoroughly studied. To address the limitations of existing research, this paper defines news quality as the ability of news articles to elicit clicks and comments from users, which represents whether the article can attract widespread attention and discussion. Based on the above definition, this paper proposes a deep learning model that integrates explicit and implicit news information for news quality evaluation.</p>
</sec>
<sec id="s3"><label>3</label><title>Materials and Methods</title>
<sec id="s3_1"><label>3.1</label><title>Materials</title>
<sec id="s3_1_1"><label>3.1.1</label><title>Data Collection</title>
<p>This work collected an online news articles dataset from a widely used Chinese news platform Toutiao (<ext-link ext-link-type="uri" xlink:href="https://www.toutiao.com">www.toutiao.com</ext-link>). Compared with other news platforms, Toutiao&#x2019;s users are more active, with more click and comment behavior. The news articles come from the five items at the top of the homepage, and the publish date is between November 05, 2020, and June 15, 2022. The collected data consists of multiple-dimensional information such as news headlines, news sources, publish time, news content, the number of clicks and comments. The dataset comprises 12108 news articles. Since these news articles are on the homepage and aimed at all users, they are more reflective of the quality of the news. At the same time, the research content of this paper focuses on evaluating whether news can get widespread attention and discussion rather than fake news detection. The top news in the dataset comes from different official media organizations, which reduces the possibility of fake news and clickbait, which is suitable for this study.</p>
</sec>
<sec id="s3_1_2"><label>3.1.2</label><title>Data Pre-Processing</title>
<p>This work uses features of the news headline and content for text semantic representation learning. During text pre-processing, this work leverages a popular Chinese word segmentation tool, Jieba3<xref ref-type="fn" rid="fn1"><sup>1</sup></xref><fn id="fn1"><label>1</label><p><ext-link ext-link-type="uri" xlink:href="https://github.com/fxsjy/jieba">https://github.com/fxsjy/jieba</ext-link></p></fn>, to segment the text, calculate the occurrence frequency of all terms in the dataset, and remove those that occurred fewer than five times or stopwords. Since 98.2&#x0025; of the news headlines in the dataset are less than 20 words after removing stopwords, the maximum length of the headlines is set to 20. The news content&#x2019;s maximum length is set to 500. Meanwhile, this work pads shorter or truncated longer news headlines and content sequences.</p>
</sec>
</sec>
<sec id="s3_2"><label>3.2</label><title>The Method of Data Labeling</title>
<p>News quality evaluation is a typical classification task, while classification problems usually use supervised learning methods and rely on labeled data. In existing studies, some studies leverage users&#x2019; dwell time on news content to assess users&#x2019; satisfaction with news [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>], but the length of news content may affect the dwell time, so it is inaccurate to use dwell time directly to label news quality. During the user&#x2019;s reading process, explicit information dramatically impacts the number of news clicks, and the number of user comments on the news indicates the headline&#x2019;s relevance to the content and the user&#x2019;s attention to the information. In other words, the number of clicks represents the breadth of news dissemination, while the number of comments is more reflective of the publicity effect of the news. Therefore, this section classifies the news quality by the number of clicks and comments. In this way, the collected data can be labeled automatically. Supposing that the number of clicks and comments on the <italic>i</italic>-th news article in the dataset during its lifetime is <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Then, we scale the values of <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to [0, 1] by the normalization method as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent the maximum and minimum number of comments, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent the maximum and minimum number of clicks.</p>
<p>As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, this work uses the number of comments and clicks on the news articles to define four quality indicators. The meaning of each indicator is described below:</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>Four news quality indicators</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-1.tif"/></fig>
<p>(1) Indicator <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>: High comment count but low click count. News articles close to this indicator are attractive to users due to their high comment count, but their headlines are not interesting enough to trigger users&#x2019; click behavior. However, the content of the news sparked discussion among users.</p>
<p>(2) Indicator <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>: High comment count and high click count. News articles close to this indicator show that high-quality news headlines attract users, and users widely discuss news content.</p>
<p>(3) Indicator <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>: High click count but low comment count. News articles close to this indicator show that their headlines interest users, but their content are not.</p>
<p>(4) Indicator <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>: High comment count and high click count. News articles close to this indicator show that neither the headlines nor the content of these news articles effectively attracts users.</p>
<p>Each news article&#x2019;s proximity to each indicator can be measured by calculating the Euclidean distance to each indicator and finally converting the space into the probability distribution by the softmax function. The calculation method is shown in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msqrt><mml:mn>2</mml:mn></mml:msqrt><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msqrt><mml:mn>2</mml:mn></mml:msqrt><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msqrt><mml:mn>2</mml:mn></mml:msqrt><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:msqrt></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msqrt><mml:mn>2</mml:mn></mml:msqrt><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:msqrt></mml:mtd></mml:mtr></mml:mtable><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s4"><label>4</label><title>The Proposed Model</title>
<sec id="s4_1"><label>4.1</label><title>Problem Statement</title>
<p>This paper defines the news quality evaluation problem as a multi-classification task; that is, the study outputs news quality class according to the explicit and implicit information of the news. Suppose there are <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>N</mml:mi></mml:math></inline-formula> news articles in dataset <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. The explicit information of <italic>i</italic>-th news article contains news headline <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, publish time <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and news source <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> that users can see before making their decisions to click, while the implicit information refers to the news article&#x2019;s content <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Obviously, explicit information has a more significant impact on the users&#x2019; click behavior. The news quality evaluation is to solve the conditional probability <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and predict the class of news articles by determining the model parameters <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>. The calculation method is shown in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">argma</mml:mtext></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s4_2"><label>4.2</label><title>The Model Architecture</title>
<p>This section presents a deep learning model for online news quality evaluation based on explicit and implicit information. In the model, we first construct embedded representations of explicit and implicit information, and then an attention unit is employed to learn their combinations adaptively. Finally, the model outputs a probability distribution of news quality. The workflow of the proposed model is illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>The proposed model for evaluating news quality based on explicit and implicit information</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-2.tif"/></fig>
<sec id="s4_2_1"><label>4.2.1</label><title>Explicit Information Representation Module</title>
<p>As shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, the news&#x2019;s headline, source, and publish time are inputs to the explicit information representation module. The architecture of the explicit information representation module is shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The first component is embedding news sources and publishing time. We consider them because the source and credibility of news have an essential impact on the quality of news [<xref ref-type="bibr" rid="ref-29">29</xref>]. Meanwhile, the publish time measures different time effects on users&#x2019; click behavior. Therefore, we consider the news source and publish time in the explicit information representation. Since the news source and publish time are discrete features, we embed these features into homologous dense vectors. Then we concatenate embedding vectors as <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext></mml:math></inline-formula> and apply fully connected (FC) layers to combine all features to <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msubsup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2032;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mo>&#x2032;</mml:mo></mml:msubsup><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are learnable parameters.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>The architecture of explicit information representation module</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-3.tif"/></fig>
<p>The second component is embedding news headlines with a typical one-dimensional CNN structure [<xref ref-type="bibr" rid="ref-31">31</xref>], which has a wide range of short-text semantic extraction applications. We exploit embedding techniques for embedding news headlines into word embedding matrix <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the <italic>i</italic>-th word in the headline, <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>n</mml:mi></mml:math></inline-formula> represents the number of words, and <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>d</mml:mi></mml:math></inline-formula> represents the embedding dimension. For each sub-matrix <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>:</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mi>q</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, we can get a feature representation <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> by convolution operation:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>:</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mi>q</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>H</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is convolution kernel; <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the window size; <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>H</mml:mi><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>:</mml:mo><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mi>q</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represents the convolution operation; <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>f</mml:mi></mml:math></inline-formula> is the Rectified Linear Unit (ReLu) activation function. After all the convolution operations, the feature sequences can be expressed as follows:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>o</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>q</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents <italic>i</italic>-th result of filter <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>H</mml:mi></mml:math></inline-formula>. Then, we adopt the max-pooling layer to identify the most significant features further:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mover><mml:mi>o</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>q</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Then, this work exploits <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>m</mml:mi></mml:math></inline-formula> convolution filters with the same window size for convolution and max-pooling as above to learn complementary features from the same window. As a result, all the parts are concatenated to get the final representation of the news headlines:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mi>o</mml:mi></mml:mrow><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mi>o</mml:mi></mml:mrow><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mi>o</mml:mi></mml:mrow><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mi>o</mml:mi></mml:mrow><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>Finally, the vectors <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mrow><mml:msubsup><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mo>&#x0027;</mml:mo></mml:msubsup></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are concatenated and fed into a one-layer FC layer to get the final explicit information representation <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. The FC layer employs ReLU as an activation function.
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo> <mml:mrow><mml:msubsup><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mo>&#x0027;</mml:mo></mml:msubsup><mml:mtext>&#x2009;</mml:mtext><mml:mo>;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow> <mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s4_2_2"><label>4.2.2</label><title>Implicit Information Representation Module</title>
<p>In the workflow shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, news content is used as input to the implicit information representation module. The news content is usually long text that determines the user&#x2019;s comments behavior. To fully extract semantic information of news content, this study divides the semantic representation of news content into two parts. The framework of the implicit information representation module is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>The architecture of implicit information representation</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-4.tif"/></fig>
<p>First, this work uses the hierarchical attention networks (HAN) model [<xref ref-type="bibr" rid="ref-32">32</xref>] to obtain the semantics of news content. The HAN model has two characteristics: the first feature is that this hierarchy corresponds to the hierarchy of the document, and the second feature is that it has two attention mechanisms, word level, and sentence level, which allows the network to distinguish essential parts of the document, to generate document representation better. Word and sentence-level encoders are bidirectional gated recurrent units (Bi-GRU) networks. Let <italic>A</italic> denotes the news article&#x2019;s content. A news article&#x2019;s content semantics can be expressed as:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mi>A</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Second, we exploit the latent Dirichlet allocation (LDA) model [<xref ref-type="bibr" rid="ref-33">33</xref>] to obtain the subject distribution of news content. The LDA topic model can represent each news content as a probability distribution composed of some topics, which can be used as part of the model input to supplement the semantic representation of news content. The calculation process is as follows:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mi>D</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Finally, the vectors <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are concatenated and fed into the FC layer to get the final content semantic representation. The FC layer employs ReLu as an activation function.
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>By using two models to represent documents separately, not only is the document&#x2019;s hierarchy considered, but it is more conducive to identifying essential parts of the news content. At the same time, better document representation results can be obtained due to the inclusion of the subject distribution of documents. In the semantic representation of news content, the LDA topic model can be trained in advance without the additional computational overhead.</p>
</sec>
<sec id="s4_2_3"><label>4.2.3</label><title>Attention Unit Module</title>
<p>Since explicit and implicit information has different effects on user behavior, news headlines, publish time, and news sources have high visibility to users, significantly impacting user click behavior. In contrast, as implicit information, news content significantly affects user comment behavior after clicking. At the same time, the headline semantics are related to the content semantics and jointly affect user behavior. Therefore, this section introduces an attention mechanism to learn the weights of both parts automatically. The calculation process is shown as follows:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mrow><mml:mover><mml:mi>a</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msubsup><mml:mrow><mml:mrow><mml:mover><mml:mi>a</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the attention weight assigned to explicit and implicit information. <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are learnable parameters.</p>
<p>Finally, the news embedding is calculated as follows:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">merge</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
<sec id="s4_2_4"><label>4.2.4</label><title>The Model Output</title>
<p>In the model output phase, we input <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">merge</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mtext>&#x00A0;</mml:mtext></mml:math></inline-formula> into the FC layer, followed by the probability distribution of news quality output <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mtext>&#x00A0;</mml:mtext></mml:math></inline-formula> by the <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mrow><mml:mtext mathvariant="italic">softmax</mml:mtext></mml:mrow></mml:math></inline-formula> function. Finally, the <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mrow><mml:mtext mathvariant="italic">argmax</mml:mtext></mml:mrow></mml:math></inline-formula> function is used to output that news falls into the category of news quality. The calculation process is as follows:
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">merge</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">argmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext></mml:math></inline-formula> and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext></mml:math></inline-formula> are learnable parameters.</p>
</sec>
</sec>
<sec id="s4_3"><label>4.3</label><title>The Loss Function</title>
<p>In the training stage, this model employs the cross-entropy loss function to measure the classification effect of the model, and the calculation method is shown in <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref>.
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Due to the imbalance problem of each class in the dataset, the CNN model trained with the imbalance data performs poorly in the weak class [<xref ref-type="bibr" rid="ref-34">34</xref>]. Since this work uses the CNN model to represent news headlines, the imbalanced data impacts the model&#x2019;s performance. Currently, the standard methods to deal with sample imbalance include resampling and undersampling. In using CNN for representation learning, resampling may introduce many duplicate samples, increasing training time and being prone to overfitting. However, undersampling may discard essential samples. Therefore, this model uses the re-weighting method proposed by Cui et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] to increase the loss weight of the class with a small sample and reduce the loss weight of the class with a large sample. By weighing the losses of each class in the dataset, the problem of uneven data distribution can be alleviated. In the re-weighting method, the effective size of a class is defined as:
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>n</mml:mi></mml:math></inline-formula> is the actual number of samples in a class and <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is a hyperparameter whose value is usually close to 1 (for example, 0.99, 0.999, and so on). Thus, the loss function <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> considering the class weights, can be defined as:
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:mfrac><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The EINQ model proposed in this paper consists of three parts: explicit information representation module, implicit information representation module, and attention fusion unit module, briefly described in Algorithm 1.
</p>
<fig id="fig-9">
<graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-9.tif"/>
</fig>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Experiments</title>
<p>This section examines the effectiveness of EINQ by comparing it with several competitive baselines and conducting extensive experiments to evaluate the impact of different modules on classification performance.</p>
<sec id="s5_1"><label>5.1</label><title>Dataset Partitioning</title>
<p>The EINQ model uses the proposed data labeling method in <xref ref-type="sec" rid="s3">Section 3</xref> to automatically label news quality in the dataset and divide the data into four classes. In the experiment, the number of training sets, validation sets, and test sets is divided into 60&#x0025;, 20&#x0025;, and 20&#x0025;, and the distribution of each class is shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Dataset partitioning</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-5.tif"/></fig>
</sec>
<sec id="s5_2"><label>5.2</label><title>Experiment Setup</title>
<p>In the explicit information representation phase, this work sets the convolution kernel window <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>q</mml:mi></mml:math></inline-formula> as 3 and sets the number of convolution kernels <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>m</mml:mi><mml:mtext>&#x00A0;</mml:mtext></mml:math></inline-formula> as 60 for feature extraction in the CNN model. In the implicit information embedding stage, this work starts with pre-trained HAN and LDA for news content representation and fine-tuning parameters with the dataset. The embedding layer adopts a 100-dimensional word vector pre-trained by Word2vec [<xref ref-type="bibr" rid="ref-36">36</xref>]. All of the FC layers have a suitable hidden size of 512. Through multiple experiments, the value of <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is set as 0.999 in the loss function. To reduce dimensionality and avoid overfitting, this works adopts the Adam optimizer for parameters learning, and the learning rate is 0.0001. The intensity of dropout applied to FC layers is 0.2. This model is trained for 20 epochs with a batch size of 32. Training is based on Pytorch 1.11.0 framework, and one NVIDIA A10 is used for calculation.</p>
</sec>
<sec id="s5_3"><label>5.3</label><title>Evaluation Metrics</title>
<p>Since this study is a typical classification problem, the widely used accuracy and macro averaged F-Score are employed to evaluate model performance. Accuracy is the ratio of correctly predicted samples to the total samples. Macro averaged F-Score is the mean of F-scores of each class which can evaluate the mode&#x2019;s overall performance in a global sense. The formulas of macro averaged F-Score are given below:
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mi>F</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>C</mml:mi></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msubsup><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the percentage of all news identified as positive samples that are indeed positive ones in class <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>c</mml:mi></mml:math></inline-formula>; <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the percentage of positive predictions that are correctly recognized in class <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>c</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="s5_4"><label>5.4</label><title>Baseline Models</title>
<p>This work compares the proposed EINQ method with the following baselines: Bi-LSTM [<xref ref-type="bibr" rid="ref-22">22</xref>], HAN [<xref ref-type="bibr" rid="ref-32">32</xref>], CNN [<xref ref-type="bibr" rid="ref-31">31</xref>], recurrent convolutional neural networks (RCNN) [<xref ref-type="bibr" rid="ref-37">37</xref>], and determining the quality of news headlines model (DQNH) [<xref ref-type="bibr" rid="ref-18">18</xref>]. We use these models because they are often used for text classification or news quality evaluation. The details of these baseline models used in this work are as follows:</p>
<p><bold>Bi-LSTM:</bold> The news headline and content are embedded as 100-dimensional vectors with Word2vec, and the headline and content are entered into the Bi-LSTM module, respectively. In the output stage, the last hidden state of the two Bi-LSTM modules is concatenated and input into the FC layer.</p>
<p><bold>HAN:</bold> Since the HAN superimposes the Bi-GRU network structure and attention module, it can better capture the essential words in the sentence. In the comparative experiment, the news headline is input into Bi-GRU, the word sequence of the news content is input into HAN as the input feature, the output of the two models is connected, and the classification result is output through the FC layer.</p>
<p><bold>CNN:</bold> In the comparative experiment, the news headline and content are processed according to the method introduced in <xref ref-type="sec" rid="s4_2">Section 4.2</xref>, and the semantic representations of the headline and content are obtained. Finally, the semantic representation of news headlines and content are connected, and the classification result is output through the FC layer.</p>
<p><bold>RCNN:</bold> The model adds an RNN layer to process the intermediate results of the CNN, and the model input is similar to the CNN.</p>
<p><bold>DQNH:</bold> The model is a method for evaluating the quality of news headlines. It contains three submodules. CNNs capture the semantic relationship between headline and content in the similarity module. In the semantic module, the Bert model represents the underlying features of the title and article. In the topic module, non-negative matrix factorization (NNMF) describes the thematic characteristics of the headline and content. The headline quality classification result is obtained by inputting the three-module output stitching into the FC layer.</p>
</sec>
<sec id="s5_5"><label>5.5</label><title>Performance Analysis</title>
<p>It is evident from <xref ref-type="table" rid="table-1">Table 1</xref> that the EINQ model yields better results than the baseline models in terms of Accuracy and F-Score. (1) CNN model is the worst-performing of all comparison models. The function of CNN convolution can effectively identify the overall structure of the news headline, but for longer news content, CNN can only process the information in the current window, and the convolution of the latter layer can only fuse the information of adjacent windows. Therefore, CNN significantly depends on parameters such as convolution windows and moving steps when processing news content, resulting in poor results in processing long articles such as news content. (2) Bi-LSTM uses two independent LSTM networks, one processing text from left to right and the other from right to left. It captures more comprehensive contextual information and is more effective when working with news content. Therefore, its news quality evaluation effect is better. (3) Compared with CNNs, RCNN has much improvement in two performance indicators and performs better than Bi-LSTM. Compared with CNN, the two-way loop structure of RCNN can obtain the context information of news content more effectively and retain word order on a large scale when learning text representation. Compared with Bi-LSTM, RCNN can automatically determine the importance of features by extracting text information using the maximum pooling layer. (4) HAN works better than RCNN. Since HAN takes advantage of the hierarchical characteristics of news content and uses attention mechanisms from the sentence and document levels, the importance of words and sentences that influence the final classification decision can be identified. Experiments show that the bidirectional loop structure of RCNN has a relatively poor ability to capture information compared with HAN. (5) Compared with other comparison models, DQNH obtains complete headline and content semantics through multiple semantic representations compared with other comparison models, which significantly improves performance. (6) Compared with DQNH, the EINQ model proposed in this paper considers the factors affecting user clicks and comments and assigns attention weight to explicit information and implicit information so that the fusion of various information is more flexible than other baseline models, and the model design is more in line with the behavioral characteristics of users, so it has the best performance. Given the applicable scenarios and experimental results of various models, this work uses CNN for explicit news headline representation, and it is reasonable for HAN to express implicit news content.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Performance comparison of the proposed model with the baseline models</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Accuracy (&#x0025;)</th>
<th align="left">F-Score (&#x0025;)</th>
<th align="left">Model size (M)</th>
<th align="left">Inference time (ms)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Bi-LSTM</td>
<td align="left">67.23</td>
<td align="left">61.45</td>
<td align="left">3.5</td>
<td align="left">0.78</td>
</tr>
<tr>
<td align="left">HAN</td>
<td align="left">71.76</td>
<td align="left">69.19</td>
<td align="left">5.6</td>
<td align="left">1.07</td>
</tr>
<tr>
<td align="left">CNN</td>
<td align="left">63.65</td>
<td align="left">60.38</td>
<td align="left">4.8</td>
<td align="left">0.84</td>
</tr>
<tr>
<td align="left">RCNN</td>
<td align="left">69.43</td>
<td align="left">67.93</td>
<td align="left">5.1</td>
<td align="left">1.12</td>
</tr>
<tr>
<td align="left">DQNH</td>
<td align="left">79.29</td>
<td align="left">77.42</td>
<td align="left">9.1</td>
<td align="left">1.78</td>
</tr>
<tr>
<td align="left">EINQ (ours)</td>
<td align="left"><bold>82.31</bold></td>
<td align="left"><bold>80.51</bold></td>
<td align="left"><bold>7.9</bold></td>
<td align="left"><bold>1.47</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Meanwhile, this study analyzes the model size and inference time. As for the deep neural network, the model size is the storage of the weight matrix. Concretely, the total number of parameters under our experiment setting is 7.9&#x2005;M. To compare the time cost of the proposed model with baseline models, we conduct one-step forward prediction experiments with batch size 32. The EINQ model takes 1.47&#x2005;ms and does not significantly increase the inference time.</p>
</sec>
<sec id="s5_6"><label>5.6</label><title>Ablation Experiments</title>
<p>This work develops several ablation models with different features to verify the influence of the effectiveness and impact of the model component. As a result, EINQ and the ablation models with different inputs are shown as follows:</p>
<p><bold>EINQ-H<sup>&#x002A;</sup>:</bold> In the process of explicit information, only the news headline is considered, and the source and publishing time of the news are not included. The design aims to analyze the impact of news sources and publishing time on the model.</p>
<p><bold>EINQ-LDA<sup>&#x002A;</sup>:</bold> The news content representation uses only the LDA topic model in the implicit information representation process.</p>
<p><bold>EINQ-HAN<sup>&#x002A;</sup>:</bold> The news content representation uses only the HAN model in the implicit information representation process.</p>
<p><bold>EINQ&#x2013;AM</bold>^<bold>:</bold> In the model output stage, the implicit and explicit information is concatenated directly and does not assign attention weights.</p>
<p>It can be seen from <xref ref-type="table" rid="table-2">Table 2</xref> that the effect of using only news headlines in the explicit information representation process is poor, indicating that the publish time and source of news are important factors that trigger user click and comment behavior. Compared with the EINQ-LDA<sup>&#x002A;</sup>, the performance of EINQ-HAN<sup>&#x002A;</sup> is relatively good, which shows that the semantics of news content has a more significant impact on the quality of news than the topic. Compared with the EINQ-LDA<sup>&#x002A;</sup> and EINQ-HAN<sup>&#x002A;</sup>, the model performance of EIQ is improved, mainly because HAN focuses on the structure of news content and LDA focuses on the topic of news content, and the combination of the HAN and LDA makes the semantic representation richer. EINQ-AM^ does not consider the attention weight of explicit and implicit information in the fusion stage, and its performance indicators lag behind other combined modules, indicating that attention weight plays an indispensable role in explicit and implicit information fusion. It can be seen from the results of the ablation experiments that each module has a particular impact on the performance of the model, and the design of explicit and implicit information significantly improves the classification effect compared with the baseline models. The performance of the baseline and ablation models are shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Ablation model performance</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Accuracy (&#x0025;)</th>
<th align="left">F-Score (&#x0025;)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">EINQ-H<sup>&#x002A;</sup></td>
<td align="left">79.45</td>
<td align="left">77.89</td>
</tr>
<tr>
<td align="left">EINQ-LDA<sup>&#x002A;</sup></td>
<td align="left">80.35</td>
<td align="left">78.32</td>
</tr>
<tr>
<td align="left">EINQ-HAN<sup>&#x002A;</sup></td>
<td align="left">80.95</td>
<td align="left">79.32</td>
</tr>
<tr>
<td align="left">EINQ-AM^</td>
<td align="left">79.23</td>
<td align="left">77.78</td>
</tr>
<tr>
<td align="left">EINQ (ours)</td>
<td align="left"><bold>82.31</bold></td>
<td align="left"><bold>80.51</bold></td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-6"><label>Figure 6</label><caption><title>The performance comparison of ablation and baseline models</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-6.tif"/></fig>
</sec>
<sec id="s5_7"><label>5.7</label><title>Hyperparameters Analysis</title>
<sec id="s5_7_1"><label>5.7.1</label><title>The Effective Size in the Loss Function</title>
<p>In the loss function, the hyperparameter <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is used to calculate the effective size of the class, whose value is a number close to 1. It can be seen from <xref ref-type="fig" rid="fig-7">Fig. 7</xref> that when <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>&#x2009;&#x003D;&#x2009;0.999, the model performs best on both performance indicators. It should be noted that when <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>&#x2009;&#x003D;&#x2009;0, the loss function degrades the traditional cross-entropy loss function. Experimental results show that increasing the category weight in the loss function affects performance improvement for unbalanced datasets.</p>
<fig id="fig-7"><label>Figure 7</label><caption><title>The effect of <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> on model performance</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-7.tif"/></fig>
</sec>
<sec id="s5_7_2"><label>5.7.2</label><title>Investigating the Influence of Convolution Kernels</title>
<p>Since EINQ uses CNN to extract the semantic features of news headlines in the explicit information representation stage, the feature processing at this stage is more sensitive to the window size <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>q</mml:mi></mml:math></inline-formula> of the convolution kernel and the number <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>m</mml:mi></mml:math></inline-formula> of convolution kernels, so different sizes and numbers of convolution kernels are used for verification in the experiment. This paper conducted multiple experiments with different <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>q</mml:mi><mml:mo>&#x2208;</mml:mo></mml:math></inline-formula> &#x007B;2, 3, 4, 5&#x007D; and <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo></mml:math></inline-formula> &#x007B;20, 30, 40, 50, 60, 79, 80, 90, 100&#x007D; values to find the appropriate window size and the number of convolution kernels. The experimental results show that the model works best when <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>m</mml:mi></mml:math></inline-formula>&#x2009;&#x003D;&#x2009;60. In the case of <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula>, the value of <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>m</mml:mi></mml:math></inline-formula> is adjusted in the experiment to observe the effect of the number of convolution kernels on the model&#x2019;s performance. As shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the optimal model performance can be achieved at <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>m</mml:mi></mml:math></inline-formula>&#x2009;&#x003D;&#x2009;60.</p>
<fig id="fig-8"><label>Figure 8</label><caption><title>Model performance under different numbers of convolution kernels when <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula></title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="IASC_41873-fig-8.tif"/></fig>
<p>As shown in <xref ref-type="table" rid="table-3">Table 3</xref>, different window sizes of the convolution kernel cause inevitable fluctuations in model performance. In the case of <italic>m</italic> = 60, the CNN module in the explicit information representation can better consider the relationship between the current convolution window and the information of adjacent windows and more effectively identify the overall structure of news headlines when <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula>.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>Model performance under different window sizes when <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>60</mml:mn></mml:math></inline-formula></title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Kernel size (<inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>q</mml:mi></mml:math></inline-formula>)</th>
<th align="left">Accuracy (&#x0025;)</th>
<th align="left">F-Score (&#x0025;)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">2</td>
<td align="left">81.78</td>
<td align="left">80.89</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left"><bold>82.31</bold></td>
<td align="left"><bold>80.51</bold></td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">81.67</td>
<td align="left">80.03</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">81.02</td>
<td align="left">79.67</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s6"><label>6</label><title>Conclusion and Future Work</title>
<p>Recommending high-quality news to users is vital in improving user stickiness and news platforms&#x2019; reputation. This study proposes a deep learning model based on explicit and implicit information for news quality evaluation. An empirical evaluation that involved real-world news datasets demonstrates the significant positive impact of the proposed explicit and implicit information on classification performance. This study makes several research contributions. First, this work defines four news quality indicators representing whether the article can attract widespread attention and discussion. Second, this paper presents a deep learning model named EINQ for news quality evaluation, which divides the features that determine the quality of news into explicit and implicit information. Third, this study conducts various experiments to verify the effectiveness of explicit and implicit information dynamic fusion. The results demonstrate performance improvements over various baseline models in news quality evaluation. Meanwhile, this work verifies the rationality of module combination and parameter design through ablation experiments and parameter analysis. Last, this research&#x2019;s findings imply that combining text and meta-features (e.g., news source and publishing time) would benefit evaluating news quality.</p>
<p>This work has some limitations that offer potential opportunities for future study. First, this paper evaluates whether news articles can get widespread attention and discussion rather than clickbait or fake news detection. Meanwhile, the data are collected from Toutiao, and the news sources are different official media organizations of China, which reduces the possibility of fake news. However, news articles with false information might get many comments and views in practical applications. Therefore, It is beneficial to enhance the ability to identify fake news through feature extraction [<xref ref-type="bibr" rid="ref-19">19</xref>] and auxiliary information (e.g., news comments and user-news interactions) [<xref ref-type="bibr" rid="ref-21">21</xref>] on different datasets in future work. Second, this study focuses on the task of categorizing news quality. However, how to generate high-quality news is also an important issue. We plan to study the generation methods of high-quality news from the perspectives of word attractiveness [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>], sentence importance [<xref ref-type="bibr" rid="ref-26">26</xref>], and news timeliness [<xref ref-type="bibr" rid="ref-10">10</xref>]. Third, this study uses classical natural language processing (NLP) models such as CNN, HAN, and LDA because they have been widely used in research. But using other NLP models, such as BERT [<xref ref-type="bibr" rid="ref-38">38</xref>], may lead to different model performances. Therefore, we plan to explore the impact of the BERT models for semantic representation on news quality evaluation.</p>
</sec>
</body>
<back>
<ack>
<p>We thank the anonymous reviewers for their valuable and helpful comments, which helped us improve this paper&#x2019;s content and presentation.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by the Fundamental Research Funds for the Central Universities (CUC230B008).</p></sec>
<sec><title>Author Contributions</title>
<p>Conceptualization, Y. W.; methodology, G. S., and Y. W.; model design, G. S., J. L., and H. H.; formal analysis, Y. W.; writing&#x2014;original draft preparation, G. S.; writing&#x2014;review and editing, J. L., and H. H.; visualization, G. S., J. L., and H. H.; project administration, Y. W.; funding acquisition, Y. W. All authors have read and agreed to the published version of the manuscript.</p></sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data for this study is extracted from <ext-link ext-link-type="uri" xlink:href="https://www.toutiao.com">toutiao.com</ext-link>, and can be requested from the corresponding author upon request.</p></sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p></sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>&#x00D6;zdemir</surname></string-name> and <string-name><given-names>N. N.</given-names> <surname>Arslan</surname></string-name></person-group>, &#x201C;<article-title>Analysis of deep transfer learning methods for early diagnosis of the COVID-19 disease with chest X-ray images</article-title>,&#x201D; <source>Duzce &#x00FC;niversity Journal of Science &#x0026; Technology</source>, vol. <volume>10</volume>, no. <issue>2</issue>, pp. <fpage>628</fpage>&#x2013;<lpage>640</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Ala&#x2019;raj</surname></string-name>, <string-name><given-names>M. F.</given-names> <surname>Abbod</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Majdalawieh</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Jum&#x2019;a</surname></string-name></person-group>, &#x201C;<article-title>A deep learning model for behavioral credit scoring in banks</article-title>,&#x201D; <source>Neural Computing and Applications</source>, vol. <volume>34</volume>, pp. <fpage>5839</fpage>&#x2013;<lpage>5866</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. H.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>J. P.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Lu</surname></string-name> and <string-name><given-names>H. J.</given-names> <surname>Ding</surname></string-name></person-group>, &#x201C;<article-title>A comparative study of automated legal text classification using random forests and deep learning</article-title>,&#x201D; <source>Information Processing &#x0026; Management</source>, vol. <volume>59</volume>, no. <issue>2</issue>, pp. <fpage>102798</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. A.</given-names> <surname>El-Demerdash</surname></string-name>, <string-name><given-names>S. E.</given-names> <surname>Hussein</surname></string-name> and <string-name><given-names>J. F. W.</given-names> <surname>Zaki</surname></string-name></person-group>, &#x201C;<article-title>Course evaluation based on deep learning and ssa hyperparameters optimization</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>71</volume>, no. <issue>1</issue>, pp. <fpage>941</fpage>&#x2013;<lpage>959</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>&#x00D6;zdemir</surname></string-name> and <string-name><given-names>M. S.</given-names> <surname>Kunduraci</surname></string-name></person-group>, &#x201C;<article-title>Comparison of deep learning techniques for classification of the insects in order level with mobile software application</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>10</volume>, pp. <fpage>35675</fpage>&#x2013;<lpage>35684</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Liang</surname></string-name>, <string-name><given-names>Z. Q.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>X. L.</given-names> <surname>Li</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Uncovering insurance fraud conspiracy with network learning</article-title>,&#x201D; in <conf-name>Proc. of the 42nd Int. ACM SIGIR Conf. on Research and Development in Information Retrieval</conf-name>, <conf-loc>Paris, France</conf-loc>, pp. <fpage>1181</fpage>&#x2013;<lpage>1184</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X. Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>D. L.</given-names> <surname>Xu</surname></string-name></person-group>, &#x201C;<article-title>Combining explicit entity graph with implicit text information for news recommendation</article-title>,&#x201D; in <conf-name>Companion Proc. of the Web Conf. 2021</conf-name>, <conf-loc>Ljubljana, Slovenia</conf-loc>, pp. <fpage>412</fpage>&#x2013;<lpage>416</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C. H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>F. Z.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Qi</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Tian</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Feedrec: News feed recommendation with various user feedbacks</article-title>,&#x201D; in <conf-name>Proc. of the ACM Web Conf. 2022</conf-name>, <conf-loc>Lyon, France</conf-loc>, pp. <fpage>2088</fpage>&#x2013;<lpage>2097</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C. H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>F. Z.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Qi</surname></string-name>, <string-name><given-names>C. L.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>Y. F.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>Is news recommendation a sequential recommendation task</article-title>,&#x201D; in <conf-name>Proc. of the 45th Int. ACM SIGIR Conf. on Research and Development in Information Retrieval</conf-name>, <conf-loc>Madrid, Spain</conf-loc>, pp. <fpage>2382</fpage>&#x2013;<lpage>2386</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Xiong</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>D. S.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>Y. F.</given-names> <surname>Leng</surname></string-name></person-group>, &#x201C;<article-title>DNCP: An attention-based deep learning approach enhanced with attractiveness and timeliness of news for online news click prediction</article-title>,&#x201D; <source>Information &#x0026; Management</source>, vol. <volume>58</volume>, no. <issue>2</issue>, pp. <fpage>103428</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Saeed</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Abbas</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Asif</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Rubab</surname></string-name>, <string-name><given-names>M. M.</given-names> <surname>Khan</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>A framework to predict early news popularity using deep temporal propagation patterns</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>195</volume>, pp. <fpage>116496</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X. G.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>J. Y.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>W. Q.</given-names> <surname>Wei</surname></string-name></person-group>, &#x201C;<article-title>Explicit time embedding based cascade attention network for information popularity prediction</article-title>,&#x201D; <source>Information Processing &#x0026; Management</source>, vol. <volume>60</volume>, no. <issue>3</issue>, pp. <fpage>103278</fpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kaur</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Kumaraguru</surname></string-name></person-group>, &#x201C;<article-title>Detecting clickbaits using two-phase hybrid CNN-LSTM biterm model</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>151</volume>, pp. <fpage>11350</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Indurthi</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Syed</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Gupta</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Varma</surname></string-name></person-group>, &#x201C;<article-title>Predicting clickbait strength in online social media</article-title>,&#x201D; in <conf-name>Proc. of the 28th Int. Conf. on Computational Linguistics</conf-name>, <conf-loc>Barcelona, Spain</conf-loc>, pp. <fpage>4835</fpage>&#x2013;<lpage>4846</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X. Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhou</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Clickbait detection on WeChat: A deep model integrating semantic and syntactic information</article-title>,&#x201D; <source>Knowledge-Based Systems</source>, vol. <volume>245</volume>, pp. <fpage>108605</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Pujahari</surname></string-name> and <string-name><given-names>D. S.</given-names> <surname>Sisodia</surname></string-name></person-group>, &#x201C;<article-title>Clickbait detection using multiple categorisation techniques</article-title>,&#x201D; <source>Journal of Information Science</source>, vol. <volume>47</volume>, no. <issue>1</issue>, pp. <fpage>118</fpage>&#x2013;<lpage>128</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>C. H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>F. Z.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Qi</surname></string-name> and <string-name><given-names>Y. F.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>Quality-aware news recommendation</article-title>,&#x201D; <comment>arXiv:2202.13605</comment>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Omidvar</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Pourmodheji</surname></string-name>, <string-name><given-names>A.</given-names> <surname>An</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Edall</surname></string-name></person-group>, &#x201C;<article-title>A novel approach to determining the quality of news headlines</article-title>,&#x201D; in <conf-name>Proc. of Natural Language Processing in Artificial Intelligence</conf-name>, <conf-loc>Berlin, German</conf-loc>, pp. <fpage>227</fpage>&#x2013;<lpage>245</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z. W.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Y. Y.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Vijayakumar</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Castiglione</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>A deep learning-based fast fake news detection model for cyber-physical social services</article-title>,&#x201D; <source>Pattern Recognition Letters</source>, vol. <volume>168</volume>, pp. <fpage>31</fpage>&#x2013;<lpage>38</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Romanou</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Smeros</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Aberer</surname></string-name></person-group>, &#x201C;<article-title>On representation learning for scientific news articles using heterogeneous knowledge graphs</article-title>,&#x201D; in <conf-name>Companion Proc. of the Web Conf. 2021</conf-name>, <conf-loc>Ljubljana, Slovenia</conf-loc>, pp. <fpage>422</fpage>&#x2013;<lpage>425</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Mosallanezhad</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Karami</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Shu</surname></string-name>, <string-name><given-names>M. V.</given-names> <surname>Mancenido</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Domain adaptive fake news detection via reinforcement learning</article-title>,&#x201D; in <conf-name>Proc. of the ACM Web Conf. 2022</conf-name>, <conf-loc>Lyon, France</conf-loc>, pp. <fpage>3632</fpage>&#x2013;<lpage>3640</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Stokowiec</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Trzcinski</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Wo&#x0142;k</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Marasek</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Rokita</surname></string-name></person-group>, &#x201C;<article-title>Shallow reading with deep learning: Predicting popularity of online content using only its title</article-title>,&#x201D; in <conf-name>Foundations of Intelligent Systems: 23rd Int. Symp.</conf-name>, <conf-loc>Warsaw, Poland</conf-loc>, pp. <fpage>136</fpage>&#x2013;<lpage>145</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Weissburg</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>P. S.</given-names> <surname>Dhillon</surname></string-name></person-group>, &#x201C;<article-title>Judging a book by its cover: Predicting the marginal impact of title on reddit post popularity</article-title>,&#x201D; in <conf-name>Proc. of the Int. AAAI Conf. on Web and Social Media</conf-name>, <conf-loc>Atlanta, USA</conf-loc>, pp. <fpage>1098</fpage>&#x2013;<lpage>1108</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Rieis</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Souza</surname></string-name>, <string-name><given-names>P. V.</given-names> <surname>Melo</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Partes</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Kwak</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Breaking the news: First impressions matter on online news</article-title>,&#x201D; in <conf-name>Proc. of the Int. AAAI Conf. on Web and Social Media</conf-name>, <conf-loc>Oxford, UK</conf-loc>, pp. <fpage>357</fpage>&#x2013;<lpage>366</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J. H.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Mantrach</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jaimes</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Oh</surname></string-name></person-group>, &#x201C;<article-title>How to compete online for news audience: Modeling words that attract clicks</article-title>,&#x201D; in <conf-name>Proc. of the 22nd ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining</conf-name>, <conf-loc>San Francisco, USA</conf-loc>, pp. <fpage>1645</fpage>&#x2013;<lpage>1654</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Q. Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zeng</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>Importance-aware learning for neural headline editing</article-title>,&#x201D; in <conf-name>Proc. of the AAAI Conf. on Artificial Intelligence</conf-name>, <conf-loc>New York, USA</conf-loc>, pp. <fpage>9282</fpage>&#x2013;<lpage>9289</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T. F.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Qian</surname></string-name>, <string-name><given-names>Z. B.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>C. W.</given-names> <surname>Wang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Jump: A joint predictor for user click and dwell time</article-title>,&#x201D; in <conf-name>Proc. of the 27th Int. Joint Conf. on Artificial Intelligence</conf-name>, <conf-loc>Stockholm, Sweden</conf-loc>, pp. <fpage>3704</fpage>&#x2013;<lpage>3710</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Romanou</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Smeros</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Castillo</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Aberer</surname></string-name></person-group>, &#x201C;<article-title>Scilens news platform: A system for real-time evaluation of news articles</article-title>,&#x201D; <source>Proceedings of the Vldb Endowment</source>, vol. <volume>13</volume>, no. <issue>12</issue>, pp. <fpage>2969</fpage>&#x2013;<lpage>2972</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Choi</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Shin</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Kang</surname></string-name></person-group>, &#x201C;<article-title>Predicting audience-rated news quality: Using survey, text mining, and neural network methods</article-title>,&#x201D; <source>Digital Journalism</source>, vol. <volume>9</volume>, no. <issue>1</issue>, pp. <fpage>84</fpage>&#x2013;<lpage>105</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. M.</given-names> <surname>Alam</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Asevska</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Roche</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Teisseire</surname></string-name></person-group>, &#x201C;<article-title>A data-driven score model to assess online news articles in event-based surveillance system</article-title>,&#x201D; in <conf-name>Annual Int. Conf. on Information Management and Big Data</conf-name>, Virtual Event, pp. <fpage>264</fpage>&#x2013;<lpage>280</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="thesis"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Convolutional neural network for sentence classification</article-title>,&#x201D; <comment>M.S. dissertation</comment>, <publisher-name>University of Waterloo</publisher-name>, <publisher-loc>Waterloo, Canada</publisher-loc>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Dyer</surname></string-name>, <string-name><given-names>X.</given-names> <surname>He</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Smola</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Hierarchical attention networks for document classification</article-title>,&#x201D; in <conf-name>2016 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</conf-name>, <conf-loc>San Diego, USA</conf-loc>, pp. <fpage>1480</fpage>&#x2013;<lpage>1489</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. M.</given-names> <surname>Blei</surname></string-name>, <string-name><given-names>A. Y.</given-names> <surname>Ng</surname></string-name> and <string-name><given-names>M. I.</given-names> <surname>Jordan</surname></string-name></person-group>, &#x201C;<article-title>Latent dirichlet allocation</article-title>,&#x201D; <source>Journal of Machine Learning Research</source>, vol. <volume>3</volume>, pp. <fpage>993</fpage>&#x2013;<lpage>1022</lpage>, <year>2003</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Buda</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Maki</surname></string-name> and <string-name><given-names>M. A.</given-names> <surname>Mazurowski</surname></string-name></person-group>, &#x201C;<article-title>A systematic study of the class imbalance problem in convolutional neural networks</article-title>,&#x201D; <source>Neural Networks</source>, vol. <volume>106</volume>, pp. <fpage>246</fpage>&#x2013;<lpage>259</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Cui</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Jia</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Song</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Belongie</surname></string-name></person-group>, &#x201C;<article-title>Class-balanced loss based on effective number of samples</article-title>,&#x201D; in <conf-name>Proc. of CVPR</conf-name>, <conf-loc>Long Beach, CA, USA</conf-loc>, pp. <fpage>9268</fpage>&#x2013;<lpage>9277</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Liu</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Analogical reasoning on Chinese morphological and semantic relations</article-title>,&#x201D; in <conf-name>56th Annual Meeting of the Association for Computational Linguistics (ACL)</conf-name>, <conf-loc>Melbourne, Australia</conf-loc>, pp. <fpage>138</fpage>&#x2013;<lpage>143</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. W.</given-names> <surname>Lai</surname></string-name>, <string-name><given-names>L. H.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Zhao</surname></string-name></person-group>, &#x201C;<article-title>Recurrent convolutional neural networks for text classification</article-title>,&#x201D; in <conf-name>Twenty-Ninth AAAI Conf. on Artificial Intelligence</conf-name>, <conf-loc>Austin, USA</conf-loc>, pp. <fpage>2267</fpage>&#x2013;<lpage>2273</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Devlin</surname></string-name>, <string-name><given-names>M. W.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Lee</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Toutanova</surname></string-name></person-group>, &#x201C;<article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>,&#x201D; in <conf-name>Proc. of the 2019 Annual Conf. of the North American Chapter of the Association for Computational Linguistics</conf-name>, <conf-loc>Minneapolis, USA</conf-loc>, pp. <fpage>4171</fpage>&#x2013;<lpage>4186</lpage>, <year>2019</year>.</mixed-citation></ref>
</ref-list>
</back></article>