<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">72064</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.072064</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Research on the Classification of Digital Cultural Texts Based on ASSC-TextRCNN Algorithm</article-title>
<alt-title alt-title-type="left-running-head">Research on the Classification of Digital Cultural Texts Based on ASSC-TextRCNN Algorithm</alt-title>
<alt-title alt-title-type="right-running-head">Research on the Classification of Digital Cultural Texts Based on ASSC-TextRCNN Algorithm</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Guo</surname><given-names>Zixuan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Houbin</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kumar</surname><given-names>Sameer</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>sameer@um.edu.my</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Yuanfang</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Asia-Europe Institute, Universiti Malaya</institution>, <addr-line>Kuala Lumpur, 50603</addr-line>, <country>Malaysia</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Resources and Environmental Engineering, Ludong University</institution>, <addr-line>Yantai, 264025</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>Faculty of Business and Economics, Universiti Malaya</institution>, <addr-line>Kuala Lumpur, 50603</addr-line>, <country>Malaysia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Sameer Kumar. Email: <email>sameer@um.edu.my</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>12</day><month>1</month><year>2026</year>
</pub-date>
<volume>86</volume>
<issue>3</issue>
<elocation-id>91</elocation-id>
<history>
<date date-type="received">
<day>18</day>
<month>08</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>02</day>
<month>12</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_72064.pdf"></self-uri>
<abstract>
<p>With the rapid development of digital culture, a large number of cultural texts are presented in the form of digital and network. These texts have significant characteristics such as sparsity, real-time and non-standard expression, which bring serious challenges to traditional classification methods. In order to cope with the above problems, this paper proposes a new ASSC (ALBERT, SVD, Self-Attention and Cross-Entropy)-TextRCNN digital cultural text classification model. Based on the framework of TextRCNN, the Albert pre-training language model is introduced to improve the depth and accuracy of semantic embedding. Combined with the dual attention mechanism, the model&#x2019;s ability to capture and model potential key information in short texts is strengthened. The Singular Value Decomposition (SVD) was used to replace the traditional Max pooling operation, which effectively reduced the feature loss rate and retained more key semantic information. The cross-entropy loss function was used to optimize the prediction results, making the model more robust in class distribution learning. The experimental results indicate that, in the digital cultural text classification task, as compared to the baseline model, the proposed ASSC-TextRCNN method achieves an 11.85% relative improvement in accuracy and an 11.97% relative increase in the F1 score. Meanwhile, the relative error rate decreases by 53.18%. This achievement not only validates the effectiveness and advanced nature of the proposed approach but also offers a novel technical route and methodological underpinnings for the intelligent analysis and dissemination of digital cultural texts. It holds great significance for promoting the in-depth exploration and value realization of digital culture.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Text classification</kwd>
<kwd>natural language processing</kwd>
<kwd>TextRCNN model</kwd>
<kwd>albert pre-training</kwd>
<kwd>singular value decomposition</kwd>
<kwd>cross-entropy loss function</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>China National Innovation and Entrepreneurship Project Fund Innovation Training Program</funding-source>
<award-id>202410451009</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the rapid advancement of information technology and the accelerating pace of digital transformation, the digitization of cultural resources has emerged as a global trend [<xref ref-type="bibr" rid="ref-1">1</xref>]. As vital carriers of cultural information, digital cultural texts are widely distributed across news media, social networks, digital archives, online publications, and various cultural dissemination platforms, exhibiting exponential growth in volume [<xref ref-type="bibr" rid="ref-2">2</xref>]. This trend not only drives profound transformations in the production and dissemination of culture but also opens up new avenues for cultural preservation and value creation. However, how to efficiently and accurately classify and analyze massive volumes of digital cultural texts&#x2014;and quickly retrieve the desired information&#x2014;has become a key research focus [<xref ref-type="bibr" rid="ref-3">3</xref>]. The explosive increase and disordered distribution of textual data, including vast amounts of short cultural texts such as microblogs, news headlines, and product reviews, have significantly impacted the efficiency and effectiveness of information retrieval, thereby attracting growing attention to digital cultural text classification [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>As a foundational task in text mining, the performance of text classification directly influences the quality and efficiency of downstream applications. In the context of digital culture, text classification faces unique challenges distinct from those in traditional domains [<xref ref-type="bibr" rid="ref-5">5</xref>]. Digital cultural texts often exhibit high sparsity and real-time generation, with imbalanced information distribution and dispersed semantic features [<xref ref-type="bibr" rid="ref-6">6</xref>]. Short texts dominate in digital media environments, and due to their brevity, they often suffer from incomplete semantic expression and insufficient information content. Moreover, the networked nature of expression introduces linguistic informality and diversity, further complicating feature extraction and semantic modeling [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Against this backdrop, recent methods&#x2014;ranging from classical bag-of-words pipelines to transformer-based classifiers&#x2014;still struggle with several practical and methodological constraints in digital-culture scenarios: short texts amplify sparsity and out-of-vocabulary issues; rapidly evolving vernacular and topic drift degrade model calibration; class imbalance and noisy labels hinder stable optimization; and common architectural choices such as max-pooling or shallow context windows often discard subtle but decisive cues. Moreover, many state-of-the-art transformers exact substantial computational cost yet provide limited gains on brevity-dominated corpora where contextualization, feature preservation, and salient-token selection must be jointly optimized. To address these limitations, we propose ASSC-TextRCNN, which augments the TextRCNN backbone with ALBERT for parameter-efficient deep semantic embeddings, enhancing generalization under data sparsity; a dual self-attention mechanism that adaptively reweights contextual signals to surface lexically sparse but semantically pivotal tokens; and Singular Value Decomposition in lieu of max-pooling to retain low-rank but information-bearing structures and mitigate feature loss, while optimizing with standard cross-entropy for stable class-distribution learning. Empirically, ASSC-TextRCNN delivers substantial gains on digital cultural text classification, achieving an 11.85% relative improvement in accuracy and an 11.97% relative increase in F1 with a 53.18% reduction in error rate over strong baselines, and exhibits a more compact, well-separated confusion structure&#x2014;underscoring that the proposed integration of pre-trained representations, attention, and SVD-based feature preservation provides an effective and efficient solution for short-text classification in modern digital media environments.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Research</title>
<p>As one of the foundational tasks in natural language processing, text classification has long been a central research focus in both academia and industry. Its core objective is to leverage algorithms and models to automatically assign large volumes of unstructured text to predefined categories, thereby enabling efficient information organization and precise retrieval [<xref ref-type="bibr" rid="ref-8">8</xref>]. With the explosive growth of text under the digital-culture paradigm and the escalating breadth of application demands, text classification has assumed an increasingly prominent role in domains such as information filtering, opinion mining, intelligent recommendation, digital archive management, and cultural resource discovery [<xref ref-type="bibr" rid="ref-9">9</xref>]. Continuous advances and innovation in text classification methods thus carry not only methodological significance, but also practical import for upgrading the cultural industry and advancing data-driven social governance.</p>
<p>Common approaches to text classification can be broadly grouped into three categories: rule-based systems, machine-learning-based systems, and deep-learning-based systems [<xref ref-type="bibr" rid="ref-10">10</xref>]. Rule-based systems are essentially akin to decision trees; while they can achieve high accuracy, they typically rely on small test sets and generalize poorly [<xref ref-type="bibr" rid="ref-11">11</xref>]. Compared with rule-based methods, machine-learning-based systems offer stronger generalization, but they require manual feature engineering, and their performance can be biased by the composition of the training data [<xref ref-type="bibr" rid="ref-12">12</xref>]. With the advent of the AI era, deep learning has been widely applied across computer vision, automatic speech recognition (ASR), and natural language processing (NLP) [<xref ref-type="bibr" rid="ref-13">13</xref>]. Deep learning obviates manual feature design and can exploit much larger training sets, albeit at the cost of reduced model interpretability [<xref ref-type="bibr" rid="ref-14">14</xref>]. Given that our task targets fine-grained classification of news texts, interpretability is not a stringent requirement. Although Chinese is among the most widely used languages globally, research on Chinese text classification remains relatively scarce [<xref ref-type="bibr" rid="ref-15">15</xref>]. This is due in part to the greater linguistic complexity of Chinese compared with English and, in part, to the limited availability of large-scale Chinese corpora&#x2014;both of which constrain progress in Chinese text classification.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Traditional Text Classification Methods</title>
<p>Early research primarily relied on statistical approaches and hybrids of rule-based and traditional machine learning methods. Naithani proposed an analytical framework that combines natural language processing (NLP) with a support vector machine (SVM) classifier to extract events from messages [<xref ref-type="bibr" rid="ref-16">16</xref>]. However, the model&#x2019;s performance was limited by substantial noise in the data. Sarker and Gonzalez trained on a combined corpus drawn from three different sources, extracting a rich set of features to train an SVM [<xref ref-type="bibr" rid="ref-17">17</xref>], experiments showed that leveraging features compatible across multiple corpora can significantly improve classification performance. In the PSB social media mining shared task, Chen et al. enhanced model expressiveness and robustness by incorporating multimodal information such as text, images, and knowledge bases [<xref ref-type="bibr" rid="ref-18">18</xref>]. Alizadeh et al. trained SVM and logistic regression classifiers using a rich feature space&#x2014;including lexicon-based, sentiment, semantic features, and word embeddings&#x2014;and reported higher accuracy than convolutional neural networks, with sentiment and embedding features contributing most [<xref ref-type="bibr" rid="ref-19">19</xref>]. Bashiri et al. trained an SVM with extensive textual, sentiment, and domain-specific features, demonstrating that the most influential features were general-domain word embeddings, domain-specific embeddings, and domain terminology [<xref ref-type="bibr" rid="ref-20">20</xref>]. To address issues such as feature sparsity and weak semantic signals, Zheng et al. have used topic vectors from the LDA (Latent Dirichlet Allocation) topic&#x2013;word distribution matrix to reduce dimensionality for short-text classification, while others fused lexical category features with semantics [<xref ref-type="bibr" rid="ref-21">21</xref>]. Although these approaches outperform traditional baselines, they rely on the LDA topic model&#x2014;an unsupervised method that is relatively slow and sensitive to the choice of the number of topics, thus requiring continual tuning. Another line of work extracts document keywords as text features and achieves competitive results [<xref ref-type="bibr" rid="ref-22">22</xref>]; yet the keyword set grows with the number of documents, inflating the dimensionality of the vector-space matrix and increasing computation. Despite steady gains, these lines of work share constraints that are acute in digital-culture corpora: they depend heavily on hand-crafted or corpus-specific features that are brittle under vernacular drift and platform noise; SVM/logistic pipelines struggle to capture long-range context and compositional semantics, while LDA-based representations are slow, topic-number-sensitive, and require continual retuning; keyword and feature inventories expand with data scale, inflating dimensionality and computation yet still discarding subtle cues via coarse pooling or sparse vectors; and multimodal add-ons often improve coverage but at the cost of complex feature fusion and limited generalization.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Machine Learning and Neural Network Classification Methods</title>
<p>In recent years, a wide range of machine learning methods has been applied to text classification. Machine learning, a subfield of artificial intelligence, aims to enable computer systems to process data through automatic learning and performance improvement without explicit programming instructions [<xref ref-type="bibr" rid="ref-23">23</xref>]. Nevertheless, text exhibits weaker regularity, greater arbitrariness, and high complexity, leading to diverse forms and structures; compared with vision, deep learning in the text domain remains relatively less mature [<xref ref-type="bibr" rid="ref-24">24</xref>]. With large-scale data and modern GPU clusters, however, theoretical and computational advances have greatly strengthened the foundation for applying deep learning to text [<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
<p>Within neural architectures, TextCNN, TextRNN, and TextRCNN are widely used models in NLP, each built on distinct neural backbones&#x2014;convolutional neural networks (CNNs), recurrent neural networks (RNNs), or their combinations&#x2014;to suit different task requirements [<xref ref-type="bibr" rid="ref-26">26</xref>]. Soni et al. proposed a CNN-based text classifier that uses pre-trained word embeddings as input and applies one-dimensional convolutions to capture word-order information, thereby enriching sentence-level representations [<xref ref-type="bibr" rid="ref-27">27</xref>]. However, CNNs are relatively weak at modeling long-range dependencies, and pooling may discard positional information, risking information loss. Akpatsa et al. first explored joint learning with a hybrid CNN&#x2013;LSTM architecture for text classification [<xref ref-type="bibr" rid="ref-28">28</xref>], subsequent work by Zhang et al. also fused CNNs and RNNs for short-text classification, leveraging CNNs for local feature extraction and RNNs for long-distance dependencies [<xref ref-type="bibr" rid="ref-29">29</xref>]. To address limitations of pure CNNs and RNNs, Long et al. introduced TextRCNN, which uses a recurrent structure to capture bidirectional context and a max-pooling layer to distill salient features&#x2014;effectively combining the strengths of RNNs and CNNs while mitigating their respective weaknesses [<xref ref-type="bibr" rid="ref-30">30</xref>]. Despite their progress, these CNN/RNN hybrids still face gaps salient in short, noisy digital-cultural texts: CNNs under-capture long-range semantics; RNNs incur gradient and efficiency costs; and TextRCNN&#x2019;s max-pooling can discard position- and context-sensitive cues. Many variants also depend on static or task-specific features and grow parameters when stacking or ensembling, which hampers robustness under slang, topic drift, and sparsity.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Improvement Directions for Text Classification</title>
<p>The introduction of attention mechanisms enables neural networks to focus selectively on salient features, yielding more faithful language modeling. Consequently, attention has been widely adopted across neural architectures. After its first use in NLP by Bhadauria et al. for machine translation, increasing numbers of studies incorporated attention to strengthen feature extraction in diverse models [<xref ref-type="bibr" rid="ref-31">31</xref>]. Sahu et al. proposed a sentence-modeling approach that combines a three-level attention mechanism with CNNs for text classification, achieving strong results [<xref ref-type="bibr" rid="ref-32">32</xref>]. Sherin et al. added bidirectional attention to GRUs for sentiment analysis, enhancing GRU-based classification [<xref ref-type="bibr" rid="ref-33">33</xref>]. Kamyab et al. introduced a bidirectional CNN&#x2013;RNN architecture fused with attention, where joint learning across deep models improved accuracy [<xref ref-type="bibr" rid="ref-34">34</xref>]. Beyond network hybrids and attention, other improvements have also been explored. Mars et al. introduced BERT, a bidirectional Transformer-based language model trained on large corpora that produces context-dependent word representations [<xref ref-type="bibr" rid="ref-35">35</xref>]. Du et al. integrated logical rules into deep neural networks to mitigate opacity and leverage prior knowledge [<xref ref-type="bibr" rid="ref-36">36</xref>]. Liu et al. proposed ACNN (Attention Convolutional Neural Network), which first uses CNNs for feature extraction and then feeds the features into a self-attention encoder and a context encoder [<xref ref-type="bibr" rid="ref-37">37</xref>]; their outputs are merged and passed to a fully connected layer and a softmax classifier. Chen et al. adapted neural machine translation&#x2013;style models to text classification by replacing recurrent units in a Seq2Seq framework with Transformer blocks and introducing multi-head attention, thereby better handling long sequences and capturing semantic relations [<xref ref-type="bibr" rid="ref-38">38</xref>]. Wang et al. proposed a model termed B_f that builds on FastText and employs bagging for ensembling [<xref ref-type="bibr" rid="ref-39">39</xref>]. While attention and Transformer-based advances have markedly improved text modeling, many methods remain over-parameterized, compute-intensive, and sensitive to domain drift and short-text sparsity; attention layers can overfit to noisy cues, hierarchical schemes add fusion complexity, and common pooling steps still lose fine-grained, position-aware signals. Moreover, gains from large pretrained encoders can taper off on brevity-dominated, slang-heavy corpora without tailored mechanisms for salient-token selection and feature preservation.</p>
<p>In the broader landscape of text feature augmentation, recent work has moved beyond raw TF-IDF (Term Frequency-Inverse Document Frequency) toward schemes that explicitly enrich or restructure document representations to better align with category semantics. Attieh et al. introduce a category-driven feature engineering (CFE) framework&#x2014;grounded in a TF-ICF variant&#x2014;that learns term&#x2013;category weight matrices and augments documents with synthetic, category-informative features, yielding higher accuracy on five benchmarks with markedly lower compute than deep models [<xref ref-type="bibr" rid="ref-40">40</xref>]. Complementing these supervised augmentations, Rathi et al. survey the mathematical underpinnings and taxonomy of term-weighting, contrasting supervised and statistical families and highlighting how Vector Space Model variants serve as the scaffolding on which many augmentation strategies are built [<xref ref-type="bibr" rid="ref-41">41</xref>]. At the level of weighting functions themselves, Li et al. propose TF-ERF, replacing the logarithmic global component of TF-RF with an exponential relevance frequency to better balance local/global contributions, improving robustness and classification quality on standard corpora [<xref ref-type="bibr" rid="ref-42">42</xref>]. Together, these efforts illustrate a continuum of augmentation&#x2014;from principled reweighting to category-aware projection and feature synthesis&#x2014;that improves downstream classifiers while controlling model complexity and training cost.</p>
<p>NLP plays a pivotal role in textual information retrieval. However, most existing classifiers are designed for balanced datasets, making model performance highly sensitive to the size and quality of training data. For Chinese, the relative scarcity and imbalance of corpora substantially degrade precision on positive classes, falling short of practical requirements [<xref ref-type="bibr" rid="ref-43">43</xref>]. Data augmentation is therefore widely used to alleviate these issues by transforming existing samples into novel ones&#x2014;commonly via synonym replacement or back-translation. Wei et al. introduced EDA (Easy Data Augmentation), which applies four simple edits&#x2014;random deletion, replacement, swapping, and insertion&#x2014;to improve text-classification performance and reduce overfitting [<xref ref-type="bibr" rid="ref-44">44</xref>], owing to its randomness, however, EDA can discard information, ignore context, and alter original meanings [<xref ref-type="bibr" rid="ref-45">45</xref>]. With advances in neural machine translation, back-translation has become a popular augmentation technique: text is translated from the source language to an intermediate language and then back to the source to create augmented samples [<xref ref-type="bibr" rid="ref-46">46</xref>]. Xie et al. proposed Unsupervised Data Augmentation (UDA), which augments manually labeled data with distantly supervised signals to construct a larger, more diverse training set and thereby improve generalization and performance across relation-extraction tasks [<xref ref-type="bibr" rid="ref-47">47</xref>]. Zheng et al. further introduced an unsupervised augmentation method that imposes a consistency loss so the model produces stable predictions across differently augmented views [<xref ref-type="bibr" rid="ref-48">48</xref>], experiments on multiple semi-supervised tasks demonstrated significant gains and broad applicability across domains and datasets. Liang et al. proposed a multi-channel neural text classification model that fuses ALBERT embeddings with CNN (local features), GCN and BiLSTM (global spatial&#x2013;temporal features), then applies softmax, achieving superior accuracy and recall to single-channel and other hybrid baselines on THUCNews [<xref ref-type="bibr" rid="ref-49">49</xref>]. Gao et al. introduced ALBERT-TextCNN-Attention (PATA), which dynamically fuses intermediate ALBERT layers via two channels, adds TextCNN for local sentiment cues and an attention mechanism for richer features, yielding a compact model (&#x2248;19.72% of BERT&#x2019;s parameters) that attains 90.63% accuracy on waimai-10k and outperforms recent ALBERT-based methods [<xref ref-type="bibr" rid="ref-50">50</xref>].</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Models and Methods</title>
<sec id="s3_1">
<label>3.1</label>
<title>Preliminaries</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>ALBERT Pretrained Language Model</title>
<p>We adopt the Chinese ALBERT pretrained language model, which encodes text features using a bidirectional Transformer encoder in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>ALBERT model architecture</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-1.tif"/>
</fig>
<p><inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the contextual feature vector of the <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>i</mml:mi></mml:math></inline-formula>-th token after encoding by the Transformer; <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>i</mml:mi></mml:math></inline-formula>-th token in the input sequence; and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:mtext>rm</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> represents the stack of bidirectional Transformer layers inside ALBERT. Unlike static word embeddings, contextual embeddings (as in BERT/ALBERT) fuse syntactic, lexical, and semantic cues and generate token representations that vary with context, thereby improving text understanding and markedly enhancing representation quality. However, BERT&#x2019;s large network and parameter count (the base configuration has <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mo>&#x223C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>110</mml:mn><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow></mml:math></inline-formula> parameters) entail substantial computation and training cost [<xref ref-type="bibr" rid="ref-51">51</xref>].</p>
<p>ALBERT addresses BERT&#x2019;s parameter and training-time overhead through architectural and objective modifications: factorized embedding parameterization decouples the word-embedding dimension <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>E</mml:mi></mml:math></inline-formula> from the hidden size <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>H</mml:mi></mml:math></inline-formula>. The original embedding matrix of size <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>V</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi></mml:math></inline-formula> (vocabulary size <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>V</mml:mi></mml:math></inline-formula>) is factorized into <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mo stretchy="false">(</mml:mo><mml:mi>V</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>E</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>E</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, drastically reducing parameters when <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>H</mml:mi></mml:math></inline-formula> is large; cross-layer parameter sharing ties weights across Transformer layers, avoiding parameter growth with depth; and replacing Next Sentence Prediction with Sentence Order Prediction improves inter-sentence relation modeling. Relative to BERT, ALBERT achieves faster training with far fewer parameters while preserving strong dynamic encoding performance. Accordingly, we employ ALBERT at the embedding layer to obtain contextualized token representations.</p>
<p>As a BERT variant, ALBERT inherits BERT&#x2019;s strengths while mitigating its drawbacks. In vanilla BERT, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mi>H</mml:mi></mml:math></inline-formula>; increasing <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>H</mml:mi></mml:math></inline-formula> forces <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>E</mml:mi></mml:math></inline-formula> to grow, risking parameter explosion. ALBERT &#x201C;unbinds&#x201D; <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>E</mml:mi></mml:math></inline-formula> from <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>H</mml:mi></mml:math></inline-formula> via the above factorization and shares parameters across layers, effectively learning one set of layer weights that is reused by all layers.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Dual Attention Mechanism</title>
<p>Attention was first introduced in vision and later adapted to text to extract word-level salient information, mimicking human selective focus while reading [<xref ref-type="bibr" rid="ref-52">52</xref>]. Readers naturally concentrate on informative spans to grasp a topic quickly. Formally, the attention module can be viewed as a single-layer neural network with a softmax output: its input is the preprocessed short-text word embeddings, and its output estimates each token&#x2019;s contribution to every class label.</p>
<p>Let an input sequence be <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>W</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>n</mml:mi></mml:math></inline-formula> is the number of tokens and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the embedding of the <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>i</mml:mi></mml:math></inline-formula>-th token. For each token <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, we compute a score vector
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>T</mml:mi></mml:math></inline-formula> is a weight matrix and <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>b</mml:mi></mml:math></inline-formula> is a bias term. Passing <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> through a nonlinearity (e.g., sigmoid) and then a softmax yields the class-posterior probabilities
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>K</mml:mi></mml:math></inline-formula> is the number of classes and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> is the <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>k</mml:mi></mml:math></inline-formula>-th component of <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. During training, we use pairs (<inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), where <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the class label of token <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in context. Model parameters are optimized by minimizing the negative loglikelihood:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>&#x2113;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mspace width="thinmathspace" /><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>with <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>m</mml:mi></mml:math></inline-formula> the number of training tokens. After training, each token <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is mapped to a probability vector over classes; we treat this vector as the token-level attention <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Stacking <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> forms the word-attention matrix <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in <xref ref-type="table" rid="table-1">Table 1</xref>, where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> encodes the confidence of token <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for each class.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Word attention matrix</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Attention vector</th>
<th>1</th>
<th>2</th>
<th><inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi mathvariant="bold-italic">k</mml:mi></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mo>&#x2026;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>T</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>SVD-Enhanced Pooling</title>
<p>To mitigate the loss of semantic and positional information caused by max pooling in TextRCNN, we replace the conventional pooling layer with singular value decomposition (SVD). This modification performs dimensionality reduction via low-rank matrix approximation, retaining principal components while preserving global semantic relations. Unlike max pooling&#x2014;which keeps only local extrema&#x2014;SVD captures the global structure of the feature matrix through spectral factorization, suppressing redundancy and reducing information loss [<xref ref-type="bibr" rid="ref-53">53</xref>]. As a result, the model attains stronger representations of deep semantic cues in text.</p>
<p>Singular value decomposition is a matrix factorization technique that expresses a matrix as the product of three matrices. For a nonzero real matrix <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> (assume <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>m</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mi>n</mml:mi></mml:math></inline-formula>), SVD writes <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mi>U</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>U</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is an orthogonal matrix whose columns are the left singular vectors, <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>V</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is an orthogonal matrix whose columns are the right singular vectors, and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is diagonal with nonnegative singular values in descending order:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>diag&#xA0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2265;</mml:mo><mml:mn>0.</mml:mn></mml:math></disp-formula></p>
<p>A standard procedure isEigen-decompose <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mi>A</mml:mi></mml:math></inline-formula> (an <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:math></inline-formula> symmetric matrix) <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> with eigenvalues <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and orthonormal eigenvectors <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, which form <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>Eigen-decompose <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>A</mml:mi><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> (an <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi></mml:math></inline-formula> symmetric matrix):
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mo>(</mml:mo><mml:mi>A</mml:mi><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>yielding orthonormal eigenvectors <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, which form <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. Let <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mrow><mml:mtext>rank&#xA0;</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>A</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>r</mml:mi></mml:math></inline-formula>. Then <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mi>A</mml:mi></mml:math></inline-formula> has <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>r</mml:mi></mml:math></inline-formula> positive eigenvalues and <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>r</mml:mi></mml:math></inline-formula> zeros:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></disp-formula>and the singular values are
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:msqrt><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>r</mml:mi></mml:math></disp-formula></p>
<p>With <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>diag&#xA0;</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, the diagonal matrix takes the block form
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>S</mml:mi></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Low-rank approximation. The singular-value spectrum typically decays rapidly; in practice, the top <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>k</mml:mi></mml:math></inline-formula> singular values (often a small fraction of the spectrum) capture the vast majority of the energy. Using the largest <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>k</mml:mi></mml:math></inline-formula> singular values and their corresponding singular vectors yields an efficient approximation:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2248;</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x226A;</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Thus, for <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, one may represent <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>A</mml:mi></mml:math></inline-formula> compactly by <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> with <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. After dimensionality reduction via SVD, the matrices <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (size <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula>) or <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> (size <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:math></inline-formula>) serve as reduced representations that preserve the most informative structure-effectively mapping an <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:math></inline-formula> feature matrix to <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows the dimension reduction and improvement method of the SVD algorithm. The improved TextRCNN model based on the SVD algorithm is shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Singular value decomposition algorithm for dimensionality reduction</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>TextRCNN-SVD network structure diagram</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-3.tif"/>
</fig>
</sec>
<sec id="s3_1_4">
<label>3.1.4</label>
<title>Cross-Entropy Loss Function</title>
<p>In this study, we adopt the cross-entropy loss function (CrossEntropyLoss) as the objective function for model training [<xref ref-type="bibr" rid="ref-54">54</xref>]. PyTorch provides a convenient implementation, nn.CrossEntropyLoss(), to compute this loss. Cross-entropy is a widely used loss function, particularly suited for classification tasks. It measures the discrepancy between two probability distributions and is extensively applied in deep learning for model optimization.</p>
<p>The principle of cross-entropy is grounded in two fundamental concepts from information theory: information content and entropy. Information content quantifies the uncertainty associated with an event, representing how much information is gained when the event occurs. This is commonly defined using a logarithmic function. We train with the standard cross-entropy loss&#x2014;implemented via PyTorch&#x2019;s nn.CrossEntropyLoss (which combines LogSoftmax and NLLLoss)&#x2014;to align the predicted class distribution with the ground-truth labels.</p>
</sec>
<sec id="s3_1_5">
<label>3.1.5</label>
<title>TextRCNN Model</title>
<p>To overcome the limitations of both RNNs and CNNs, the TextRCNN architecture was proposed. It leverages a bidirectional recurrent structure to capture both preceding and succeeding context, enabling the model to preserve word order over a broader range compared to CNNs. Additionally, it applies a max-pooling layer to extract the most important semantic features and automatically determine which parts of the text are most relevant for classification [<xref ref-type="bibr" rid="ref-55">55</xref>]. The architecture of the TextRCNN classification model is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. The model consists of three components: (1) a recurrent structure, (2) a max-pooling layer, and (3) an output layer. First, each input word is mapped to its corresponding word embedding through the input layer, resulting in a word embedding matrix for the sentence. The size of each sentence matrix is <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi>n</mml:mi></mml:math></inline-formula> is the number of words in the sentence and <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi>k</mml:mi></mml:math></inline-formula> is the dimensionality of the word embeddings.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>TextRCNN model architecture</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-4.tif"/>
</fig>
<p>The word embeddings are then fed into a bidirectional recurrent network to obtain both forward and backward contextual representations for each word:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>here, <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>i</mml:mi></mml:math></inline-formula>-th word in the input sequence. <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represent the left (preceding) and right (succeeding) contextual representations of <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, respectively. <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the word embedding of the previous word. The matrices <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>s</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> are weight parameters that propagate the left context and incorporate the semantics of the previous word into the current context. The function <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>f</mml:mi></mml:math></inline-formula> is a non-linear activation function.</p>
<p>By concatenating the contextual vectors with the original word embedding, the complete representation of word <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is defined as:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>&#x003A;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Next, a non-linear transformation is applied to compute the hidden semantic representation of each word:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>To identify the most salient semantic information across the entire sentence, a max-pooling operation is applied:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>Finally, <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>3</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is passed through a fully connected layer, followed by a Softmax function for classification.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Proposal</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Improvement of the ALBERT Pre-Trained Model</title>
<p>In this work we use albert_tiny_zh, pretrained on a large-scale Chinese corpus of roughly two billion samples [<xref ref-type="bibr" rid="ref-56">56</xref>]. Compared with BERT base, it retains comparable accuracy while using only about <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mn>4</mml:mn><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow></mml:math></inline-formula> parameters (approximately <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>28</mml:mn></mml:math></inline-formula> of BERT base), and yields roughly an order-of-magnitude speedup in training and inference.</p>
<p>Given a preprocessed input sequence
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mtext>text</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>ALBERT produces contextualized embeddings
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ALBERT&#xA0;</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>text</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>d</mml:mi></mml:math></inline-formula> is the embedding dimension, and <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the embedding of the <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>i</mml:mi></mml:math></inline-formula>-th token. These design choices substantially reduce parameters and training time while strengthening semantic modeling, yielding competitive classification performance.</p>
<p>Using the base version of the ALBERT model as an example, this study proposes an adaptive method to generate a semantic representation vector that best captures the meaning of each input text. Specifically, instead of relying solely on the [CLS] vector from the final Transformer layer-as done in the original ALBERT architecture-we extract the [CLS] vectors from all 12 Transformer encoder layers. We then compute dynamic weights to quantify the contribution of each layer&#x2019;s [CLS] vector to the final representation. These weights are applied via an attention mechanism, producing a fused [CLS] vector that integrates semantic information from all Transformer layers. This enhancement significantly improves the expressiveness and accuracy of the resulting text representation. Since each Transformer layer in ALBERT&#x2019;s pre-trained model consists only of an encoder block, we denote the [CLS] output from the <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mi>i</mml:mi></mml:math></inline-formula>-th encoder layer as <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msup><mml:mi>encoder_CLS</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. This yields a sequence of classification vectors: <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mrow><mml:mtext>encoder</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>encoder</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:msup><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mrow><mml:mtext>&#xA0;encoder</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:msup><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mtext>&#xA0;encoder</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:msup><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></inline-formula></p>
<p>To enable the model to dynamically assign weights to each of these vectors based on the input text, we design a recurrent module based on fully connected layers. For each classification vector in the encoder_CLS group, a corresponding output is produced via a linear transformation, capturing the information strength that layer contributes to the current classification task. After all layers are processed, their outputs are concatenated and passed through a softmax normalization to produce a probability distribution over the 12 layers, representing the learned importance (dynamic weight) of each layer&#x2019;s [CLS] vector. This adaptive fusion mechanism allows the model to focus on the most informative layers per input instance, yielding a richer and more context-sensitive representation than static, single-layer approaches.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Improvement of the Aggregation Mechanism of the Dual Attention Mechanism</title>
<p>We adopt a dual attention design, injecting attention at both the input layer and the bidirectional recurrent layer: Before final classification, we estimate each token&#x2019;s contribution to each class to enable token filtering. The goal is to allocate higher attention to semantically meaningful parts of speech-primarily nouns and verbs-while assigning little or no attention to function words such as prepositions, particles, or colloquialisms. This increases the weight of tokens that carry precise semantics in the downstream classifier.</p>
<p>We introduce a class-conditional hierarchical aggregation mechanism following the acquisition of contextual representations <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which are derived from input tokens filtered by input-level attention and subsequently encoded by a bidirectional recurrent layer. Specifically, token-level attention weights <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> are used to aggregate all tokens within a document, producing a class-specific document representation <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for each class <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>k</mml:mi></mml:math></inline-formula>. These representations <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> are then passed through a shared classification head to yield document-level class distributions. In this architecture, the input-level attention suppresses functional and noisy words before encoding, while the recurrent-layer attention emphasizes salient information in context. Together, they determine the magnitude of <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, enabling class-specific attention pooling to yield fixed-length document representations invariant to input length.</p>
<p>To enhance robustness, we apply temperature normalization and thresholding (or Top-<inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mi>r</mml:mi></mml:math></inline-formula> selection) to the attention scores prior to aggregation. Additionally, SVDenhanced pooling can be applied to <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> for low-rank reconstruction that preserves global semantic structure, which is then fused with the attention aggregated <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The entire process is trained end-to-end with supervision only from document-level labels. Attention parameters are learned through backpropagation, allowing the model to automatically infer which tokens, under what context, are most important for which class, thereby consistently elevating token-level saliency into document-level discriminative power.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>SVD Enhanced Pooling Improvement</title>
<p>In the TextRCNN module, the contextual token encodings from ALBERT <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are convolved with <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mi>n</mml:mi></mml:math></inline-formula> filters <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mi>w</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, producing <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mi>n</mml:mi></mml:math></inline-formula> column vectors <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mi>c</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. Stacking these yields the feature-map matrix
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>We then apply SVD to <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mi>C</mml:mi></mml:math></inline-formula> and retain the top <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mi>k</mml:mi></mml:math></inline-formula> singular values and their singular vectors to obtain a rank-<inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mi>k</mml:mi></mml:math></inline-formula> approximation
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Using <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> as a compact summary captures the dominant structure of the feature maps and reduces the dimensionality from <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula>. A subsequent flattening operation gives
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Flatten</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>For <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mi>l</mml:mi></mml:math></inline-formula> distinct kernel widths, we perform this SVD-based reduction and flattening for each feature-map matrix and then concatenate the resulting vectors to form the local textual feature representation
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:mtext>local</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2295;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2295;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2295;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>For a feature matrix <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> with sequence length <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mi>L</mml:mi></mml:math></inline-formula> and hidden dimension <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mi>d</mml:mi></mml:math></inline-formula>, the computational cost of a full SVD is approximately <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mi>L</mml:mi><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>d</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. To reduce this cost, we adopt truncated or randomized SVD, retaining only the top <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mi>r</mml:mi><mml:mo>&#x226A;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mi>L</mml:mi><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> singular components, which reduces the complexity to <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mi>d</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>-comparable to that of common attention mechanisms in the context of short texts and moderate-sized embeddings. SVD is differentiable almost everywhere on <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow></mml:math></inline-formula> when the singular values are distinct, enabling stable backpropagation. To address the non-smoothness at points of repeated singular values, we apply truncated SVD and impose thresholding or gradient stopping on near-zero singular values, restricting gradient computation to <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and its associated subspaces. This significantly improves gradient stability. During training, we enable full_matrices &#x003D; False and enforce singular value sorting. Input features are preprocessed with layer normalization and scale clipping. When necessary, a small Tikhonov regularization term is added to <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:msup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mrow><mml:mtext>X</mml:mtext></mml:mrow></mml:math></inline-formula> to improve the condition number. The retained rank <inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mi>r</mml:mi></mml:math></inline-formula> is subject to early stopping or upper-bound constraints, and gradient clipping is applied to prevent domination by large singular values. Under mixed precision training, we perform appropriate casting before and after SVD to ensure decomposition and backpropagation are maintained in FP32 precision.</p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>Improvement of the cross Entropy Loss Function</title>
<p>From the definitions of entropy, KL divergence, and cross-entropy, the following relationships can be derived:
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mspace width="thinmathspace" /><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mspace width="thinmathspace" /><mml:mrow><mml:mo>[</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Substituting into the definition of KL divergence yields: <disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>here, <inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the fixed distribution of the training data, so <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is constant. Therefore, minimizing the cross-entropy <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is equivalent to minimizing the KL divergence <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, with the goal of making the predicted distribution <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> as close as possible to the target distribution <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>.</p>
<p>Given the prevalence of long-tail categories in news corpora, we assign a class weight <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to each ground-truth label <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:mi>y</mml:mi></mml:math></inline-formula> to mitigate the dominance of mainstream sections in the loss function. Additionally, we apply label smoothing to the target distribution, replacing one-hot labels with <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, to reduce overfitting caused by annotation noise and semantically similar classes. To emphasize hard examples and easily confusable sections, we incorporate a focal modulation factor <inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> into the cross-entropy loss-where <inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the predicted probability of the true class and <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is a small value (e.g., 1&#x2013;2) thus adaptively amplifying the contribution of difficult samples without altering the gradient form. News labels often exhibit a hierarchical structure of &#x201C;channelsubcategory&#x201D;. To leverage this, we construct a cost matrix <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mi>C</mml:mi></mml:math></inline-formula> that assigns lower penalties when predictions fall within the same parent channel as the true class, and higher penalties for cross-channel misclassifications.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments and Results</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Data and Environment</title>
<p>The experimental dataset is a text classification dataset selected from THUCNews (URL: <ext-link ext-link-type="uri" xlink:href="https://github.com/jinmuxige0816/THUCNews/tree/main">https://github.com/jinmuxige0816/THUCNews/tree/main</ext-link>, accessed on 01 December 2025). It contains a total of 200,000 news headlines, each with a text length between 10 and 30 characters. The dataset covers ten categories: finance, real estate, stock, education, technology, society, current affairs, sports, games, and entertainment. Each category consists of 20,000 text samples, labeled with digits 0&#x2013;9. In the experiments, 180,000 samples were randomly selected as the training set, 10,000 as the test set, and another 10,000 as the validation set. The experimental environment configurations are shown in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Experimental environment configurations</title>
</caption>
<table>
<colgroup>
<col align="center" width="45mm"/>
<col align="center" width="55mm"/> </colgroup>
<thead>
<tr>
<th>Environment</th>
<th>Configuration</th>
</tr>
</thead>
<tbody>
<tr>
<td>Operating system</td>
<td>Windows 11</td>
</tr>
<tr>
<td>CPU</td>
<td>Intel Core i9-14900k</td>
</tr>
<tr>
<td>GPU</td>
<td>NVIDIA GeForce RTX 3090 Ti</td>
</tr>
<tr>
<td>Python</td>
<td>3.8</td>
</tr>
<tr>
<td>PyTorch</td>
<td>1.12.0&#x002B;cu113</td>
</tr>
<tr>
<td>Encoding format</td>
<td>UTF-8</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Hyperparameter Comparison Experiments</title>
<p>In this section, a series of comparative experiments are conducted on three hyperparameters&#x2014;namely the value of <italic>k</italic> in the SVD algorithm, the convolution kernel size, and the number of convolution kernels&#x2014;in order to examine their impact on the classification performance of the proposed model and to determine their optimal settings.</p>
<p>For the SVD algorithm, dimensionality reduction is achieved by retaining the top <inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mi>k</mml:mi></mml:math></inline-formula> singular values and representing the main features of the feature map matrix using the product of the <inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mi>k</mml:mi></mml:math></inline-formula>-order left singular matrix and the singular value matrix, <inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03A3;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The choice of <inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mi>k</mml:mi></mml:math></inline-formula> is crucial, as it determines both the reduced feature dimension and the quality of dimensionality reduction. In this study, the maximum dimension of the feature vectors requiring dimensionality reduction via SVD is 90. Therefore, with other parameters fixed, <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mi>k</mml:mi></mml:math></inline-formula> is varied from 1 to 9. The evaluation indicators used were the time consumption comparison before and after the SVD improvement and the macro-F1 score (macro-F1), in order to analyze the impact of <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mi>k</mml:mi></mml:math></inline-formula> on the model performance. The experimental results are shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Model performance and time under different <italic>k</italic> values in the SVD algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-5.tif"/>
</fig>
<p>The feature dimension retention parameter <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mi>k</mml:mi></mml:math></inline-formula> in the SVD algorithm has a significant impact on model performance. When <inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mi>k</mml:mi></mml:math></inline-formula> is too small, insufficient feature dimensions are preserved, leading to information loss and limiting the model&#x2019;s feature extraction ability. Conversely, when <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mi>k</mml:mi></mml:math></inline-formula> is too large, redundant dimensions increase model complexity, which may cause overfitting. When <inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>, the model achieves the highest macro-F1 score. Therefore, in this study, the SVD algorithm retains five dimensions after feature vector reduction. The runtime curves indicate that computational cost increases almost monotonically with <inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mi>k</mml:mi></mml:math></inline-formula>: for SVD, the time grows from 2.88 to 8.10 h; for MaxPool, from 0.96 to 4.75 h. The performance gap between the two is approximately threefold in the low-<inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mi>k</mml:mi></mml:math></inline-formula> regime, gradually narrowing as <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mi>k</mml:mi></mml:math></inline-formula> increases but remaining substantial, thereby quantifying the trade-off of performance for time. Balancing performance and efficiency, <inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> yields the optimal macro-level discriminative power with an acceptable computational cost under the current dataset and implementation. This result also supports a practical guideline: when the network and input scale are fixed, selecting a moderate <inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mi>k</mml:mi></mml:math></inline-formula> value that captures the major energy components without excessively increasing the rank allows SVD to effectively exploit its global structural expressiveness while avoiding unnecessary computational overhead and the risk of overfitting.</p>
<p>Model parameters are internal configuration variables, and different parameter settings can significantly affect model performance. Compared with traditional machine learning, deep learning involves larger-scale data and requires a broader range of parameter choices. The hyperparameter settings used in this study are summarized in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Hyperparameter settings of the model</title>
</caption>
<table>
<colgroup>
<col align="center" width="58mm"/>
<col align="center" width="42mm"/> </colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Text length</td>
<td>32</td>
</tr>
<tr>
<td>Epochs</td>
<td>3</td>
</tr>
<tr>
<td>Batch size</td>
<td>128</td>
</tr>
<tr>
<td>Hidden size</td>
<td>768</td>
</tr>
<tr>
<td>Convolution kernel size</td>
<td>2, 3, 4</td>
</tr>
<tr>
<td>Number of convolution kernels</td>
<td>256</td>
</tr>
<tr>
<td>Learning rate</td>
<td><inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Dropout</td>
<td>0.1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To ensure efficiency, we first conducted a small-scale hyperparameter search on a fixed validation set and determined the final configuration based on a trade-off between macro-averaged F1 (macro-F1) and training time. Specifically, the learning rate was fine-tuned over <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> using the AdamW optimizer with linear warm-up followed by a constant schedule. Results indicated that <inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> offered the best convergence speed and stability, and was therefore adopted. The number of training epochs was compared across <inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> with early stopping enabled. Performance gains on the validation set saturated after 3 epochs, with signs of slight overfitting, leading to the choice of 3 epochs. The batch size was constrained by GPU memory and effective batch size considerations; a size of 128 provided a good balance between throughput and generalization. The hidden dimension was set to 768, consistent with the backbone size of the pretrained encoder, with no further search. The input sequence length was set to 32 based on a prior analysis of corpus length distribution-this choice maintained a controllable truncation ratio, and increasing it to 64 or 128 yielded marginal macro-F1 improvements at the cost of significantly higher memory and latency. For convolutional layers, kernel sizes <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> were selected to cover diverse local patterns, showing greater robustness on the validation set compared to <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mo stretchy="false">(</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mn>5</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The number of convolutional kernels was set to 256, achieving a good trade-off between representational capacity and parameter efficiency. Dropout was fixed at 0.1, a common setting for fine-tuning pretrained models.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Effect Experiment of Algorithm Module</title>
<p>To further illustrate the impact of different architectural enhancements, <xref ref-type="fig" rid="fig-6">Fig. 6</xref> compares the performance of baseline models with their variants augmented by attention mechanisms, singular value decomposition (SVD), and the combination of both. The y-axis represents the performance metrics of the evaluated text classification models. Specifically, it shows the Accuracy (%) and F1-score (%), both expressed as percentages.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Evaluation results of classical text classification models. All models share identical data splits, hyperparameter configurations, and early-stopping criteria to ensure fair comparison. The reported metrics&#x2014;classification accuracy and macro-averaged F1&#x2014;are averaged over five random seeds, with standard deviations below 0.5% omitted for clarity. (1) Baseline Models: trained without attention or singular value decomposition (SVD) modules. (2) Baseline &#x002B; Attention: models augmented with token-level attention to highlight salient contextual information. (3) Baseline &#x002B; SVD: models enhanced with truncated SVD pooling for global structural representation. (4) Baseline &#x002B; SVD &#x002B; Attention: combined attention&#x2013;SVD variants integrating both local focus and global semantic compression</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-6.tif"/>
</fig>
<p>When evaluating the constructed dataset with classical text classification models, we select four deep learning approaches without BERT or attention&#x2014;TextCNN, TextRNN, DPCNN, and TextRCNN&#x2014;and compare performance using accuracy and F1 score. The accuracies of the four models on this dataset are 80.82%, 83.92%, 85.08%, and 81.78%, respectively. Overall, DPCNN performs best in terms of accuracy, reaching 85.08%; its stacked deep convolutional blocks effectively enhance contextual modeling, conferring advantages in capturing long-range dependencies and extracting deep semantic features. In contrast, TextRNN achieves the highest F1 score (86.94%), indicating a better balance of precision and recall for this task. Owing to its recurrent architecture, TextRNN models sequential dependencies particularly well, leading to superior class balance in predictions. Despite the computational efficiency and simplicity of TextCNN&#x2019;s local convolutional feature extraction, its accuracy (80.82%) and F1 score (81.81%) are lower than those of the other models, reflecting limitations in modeling long texts on this dataset. TextRCNN attains intermediate-to-lower performance, with an accuracy of 81.78% and an F1 score of 81.71%. Although its hybrid convolution&#x2013;recurrent design captures both local and global features, it does not deliver a marked advantage at larger experimental scales.</p>
<p>After introducing a dual attention mechanism on this dataset, TextRNN_Att stands out, with accuracy rising to 87.78% and F1 to 87.23%, improvements of 3.86 and 0.29 percentage points over the baseline, respectively. TextRCNN_Att also improves noticeably, with accuracy increasing from 81.78% to 84.70% and F1 from 81.71% to 85.80% (gains of 2.92 and 4.09 percentage points), yet its absolute performance still lags behind TextRNN_Att by 3.08 points in accuracy and 1.43 points in F1. For DPCNN, the dual attention mechanism yields a mixed effect&#x2014;accuracy decreases from 85.08% to 83.17% (&#x2212;1.91 points) while F1 increases from 82.08% to 83.42% (&#x002B;1.34 points)&#x2014;suggesting that dual attention helps balance precision and recall but its global reweighting may partially interfere with DPCNN&#x2019;s stable extraction of hierarchical n-gram structures. TextCNN_Att attains an accuracy of 80.61% and an F1 of 80.17%, both slightly below the baseline (&#x2212;0.21 and &#x2212;1.64 points), indicating that in shallow architectures dominated by local convolutions, dual attention provides limited benefit and may introduce minor performance fluctuations due to increased variance in weight distributions.</p>
<p>Building on the original models, further integrating singular value decomposition (SVD) for feature processing yields substantial gains across all four model types. SVD_TextRCNN exhibits the largest improvement, with accuracy rising from 81.78% to 87.45% (&#x002B;5.67 points) and achieving the best F1 overall. The dual attention mechanism, through two levels of weighting at the token and sentence levels, both highlights salient information and suppresses irrelevant features, while globally refining the semantic distribution of sentence representations&#x2014;thereby granting RNN/RCNN-type models stronger advantages in recall and F1. SVD, by performing low-rank matrix factorization for dimensionality reduction and denoising, preserves principal semantic directions and removes high-dimensional redundancy, simultaneously reducing model complexity and improving interclass linear separability; this effect is particularly pronounced for high-dimensional long-sequence representations.</p>
<p>Taken together merely introducing the dual attention mechanism improves F1 for some models but does not fundamentally alter TextRCNN&#x2019;s relative disadvantage; however, augmenting the representation with SVD markedly enhances separability and generalization for all models, with TextRCNN attaining the highest F1 (88.34%). These results indicate that combining a dual attention mechanism with SVD can further unlock performance potential in the subsequent fusion model and better balance precision and recall in text classification tasks.</p>
<p>To further enhance these SVD-processed models, we introduce an attention mechanism; in particular, adding attention to SVD-TextRCNN yields the method proposed in this paper. After introducing a dual attention mechanism into the four classical text classification models already processed by singular value decomposition (SVD), all models achieve varying degrees of improvement. Among them, SVD_TextRCNN_Att performs the best, reaching an accuracy of 89.76% and an F1 score of 89.49%, which substantiates the significant role of dual attention in strengthening long-sequence dependency modeling and optimizing the fusion of global and local features. SVD_TextRNN_Att attains an accuracy of 89.35% and an F1 of 89.40; compared with SVD_TextRNN (88.82% accuracy, 87.68% F1), this reflects gains of 0.53 percentage points in accuracy and 1.72 points in F1. These results indicate that the dual attention mechanism can further enhance semantic weighting in recurrent architectures that already possess strong sequential modeling capabilities, yielding a more balanced trade-off between precision and recall.</p>
<p>At the feature-processing stage, SVD uses low-rank approximation to compress high-dimensional sparse features, remove noise, and preserve the principal semantic directions, thereby making the input features more compact and discriminative and providing a clearer semantic distribution for downstream classification. Within the model, the dual attention mechanism applies two levels of weighting&#x2014;word-level and sentence-level&#x2014;which both highlights the tokens most critical for classification and refines the sentence representation at the global contextual level, organically combining salient local information with global semantic structure. The fusion of these components equips the model with efficient, robust feature representations while enabling dynamic reweighting during inference, markedly improving its ability to capture complex semantic patterns.</p>
<p>This fusion strategy is particularly effective for SVD_TextRCNN: leveraging RCNN&#x2019;s strong contextual capturing capacity, SVD&#x2019;s feature compression advantages, and dual attention&#x2019;s global&#x2013;local weighting, it attains the highest accuracy and F1, validating the effectiveness and superiority of the proposed method in text classification by balancing precision and recall and enhancing model generalization.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experiments before and after the Improvement</title>
<p>To systematically verify the effectiveness and trainability of the proposed ASSC-TextRCNN for text classification, we conduct controlled experiments using the classical TextRCNN as the baseline. The two models are trained under identical data splits and preprocessing pipelines, and they share the same optimizer, learning-rate schedule, batch size, maximum number of iterations, and regularization strategy. Except for introducing the ASSC module, all other network components and training hyperparameters remain unchanged to ensure fairness. We evaluate classification accuracy (Accuracy) and cross-entropy loss (Loss) on both the training and validation sets, recording them in sync and visualizing them against training steps, so as to examine optimization convergence, generalization behavior, and potential overfitting [<xref ref-type="bibr" rid="ref-57">57</xref>]. We compile statistics for all indicators throughout training before and after the improvement; the comparative trajectories are shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Change of iteration index before and after improvement</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-7.tif"/>
</fig>
<p>Under the same data and training configuration, the improved ASSC-TextRCNN exhibits clear advantages over the baseline TextRCNN in convergence speed, final performance, and training stability. From the training/validation loss and accuracy curves, ASSC-TextRCNN achieves a rapid loss drop and steep accuracy rise in the early iterations (&#x007E;1&#x2013;3k), with training accuracy quickly surpassing 90% and entering a plateau, whereas the baseline requires more iterations to approach a lower plateau&#x2014;evidence of better optimization attainability and sample efficiency. In terms of final performance, the validation accuracy of ASSC-TextRCNN stabilizes around &#x007E;90%, an absolute gain of about 10 percentage points over the &#x007E;80% baseline; meanwhile, its validation loss remains around &#x007E;0.32&#x2013;0.35, markedly lower than the baseline&#x2019;s &#x007E;0.55&#x2013;0.60, indicating reduced generalization error. The training and validation trajectories of ASSC-TextRCNN are overall smoother with smaller oscillations, suggesting a more stable optimization process and lower sensitivity to gradient noise and hyperparameter perturbations. By strengthening contextual representations and inter-class separability, the ASSC mechanism accelerates the formation of discriminative features and smooths decision boundaries, thereby improving both convergence efficiency and generalization. Across accuracy, convergence, and robustness, ASSC-TextRCNN consistently outperforms the baseline, demonstrating clear empirical advantages and practical value.</p>
<p>To comprehensively assess the effectiveness of ASSC-TextRCNN on a 10-class news classification task, we again adopt TextRCNN as the baseline and conduct comparable controlled experiments under identical data splits, preprocessing, and training hyperparameters. Evaluation metrics include per-class Precision, Recall, and F1, which are visualized in three heatmaps for the baseline, the improved model, and their difference, using both color intensity and cell values to encode performance magnitude and improvement.</p>
<p>The improved model achieves consistent and substantial gains on all three macro-averaged metrics in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, Precision rises from 0.817 to 0.916, Recall from 0.818 to 0.915, and F1 from 0.817 to 0.915, indicating simultaneous enhancement of discriminability and coverage and a significant reduction in overall error. At the class level, improvements are broad and particularly pronounced for previously weaker categories: for example, stocks see F1 increase from 0.658 to 0.853, with Recall improving from 0.639 to 0.875 and Precision from 0.679 to 0.832; finance improves in F1 from 0.774 to 0.908, with Precision and Recall gains of 0.130 and 0.136, respectively. Even for categories that already performed well, the improved model delivers steady gains&#x2014;for instance, education F1 rises from 0.903 to 0.953, and sports from 0.938 to 0.977&#x2014;showing that errors can still be further reduced in high-baseline regimes.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Heatmap of per-class metrics</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-8.tif"/>
</fig>
<p>The difference heatmap exhibits a predominantly positive and continuous coloration, reflecting concurrent improvements in Precision and Recall for most categories. These gains do not stem from threshold trade-offs but from systematic enhancements in representational capacity and class separability. Coupled with the larger gains observed in semantically overlapping categories such as stocks, finance, and science, this suggests that the ASSC mechanism more effectively captures long-range context and topical cues, reduces cross-category confusion, and yields smoother, more robust decision boundaries. In summary, under a rigorously fair setting, ASSC-TextRCNN significantly outperforms the baseline at both macro and micro levels, demonstrating broad effectiveness and strong generalization, and providing solid empirical support for scaling to larger and cross-domain datasets.</p>
<p>To examine the discriminative capacity and error structure of ASSC-TextRCNN on the 10-class news task, we conduct strictly comparable controlled experiments with the same data split, preprocessing, and training hyperparameters as TextRCNN, and visualize the test-set confusion matrices of both models side-by-side to characterize inter-class misclassification patterns and their trends (a LogNorm color map is used to enhance readability across different count magnitudes) in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Comparison of confusion matrix of classification task before and after improvement</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72064-fig-9.tif"/>
</fig>
<p>The diagonal counts for ASSC-TextRCNN increase markedly across all categories, and overall accuracy improves from 81.78% to 91.47%, an absolute gain of &#x002B;9.69 percentage points. This gain represents a &#x201C;net improvement&#x201D; driven by a substantial reduction in systematic confusions rather than a trade-off between categories. For key confusable pairs, bidirectional mistakes between finance and stocks shrink significantly: &#x201C;finance-stocks&#x201D; drops from 163 to 75, and &#x201C;stocks-finance&#x201D; from 123 to 30, with the recalls of the two classes rising from 0.735 and 0.639 to 0.871 and 0.875, respectively. Confusions due to semantic proximity between science and games also diminish notably&#x2014;for instance, &#x201C;games-science&#x201D; falls from 83 to 30, and &#x201C;science-stocks/politics/games&#x201D; decrease from 71/50/58 to 33/20/27 (a total reduction of 99)&#x2014;raising the science recall from 0.705 to 0.866. For high-baseline categories, the diagonal strengthening in education and sports remains evident, while residual confusions such as &#x201C;education-entertainment&#x201D; and &#x201C;entertainment-sports/games&#x201D; drop from 34 to 10 and from 27/23 to 12/5, respectively. By integrating context more effectively and shaping decision boundaries, ASSC-TextRCNN systematically reduces cross-class errors caused by topical overlap and lexical co-occurrence, improving recall for weaker classes while further lowering errors for stronger ones. The remaining few symmetric or near-symmetric confusions suggest that future work could incorporate finer-grained domain features and cross-paragraph semantic constraints to eliminate edge cases.</p>
<p>To comprehensively evaluate the overall advantages of the proposed ASSC-TextRCNN, which fuses SVD and attention, on the 10-class news classification task, we conduct comparable controlled experiments under the same data splits, preprocessing pipeline, and training settings as the competing models, and perform a horizontal comparison using accuracy and F1 as metrics across classical sequence models (BiGRU, BiLSTM, FastText), a self-attention model (Transformer), and pretrained language models (BERT, MacBERT, ELECTRA). The results are presented in <xref ref-type="table" rid="table-4">Table 4</xref>. To ensure a fair comparison, all models were trained using an equivalent number of optimization steps and a unified early stopping criterion, rather than a fixed number of epochs. Specifically, we employed the AdamW optimizer with <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mn>10</mml:mn><mml:mrow><mml:mtext>\%&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula> warm-up and cosine annealing for learning rate scheduling. Early stopping was monitored based on macro-F1 on the validation set, with a minimum improvement threshold of 0.001 and identical evaluation frequency across models. Additional checkpoints were performed at 3, 5, and 10 epochs. Each experiment was repeated with five different random seeds, and we report the mean and standard deviation of the results. For baselines that did not meet the early stopping criterion within 3 epochs, training was continued until either early stopping was triggered or the shared maximum step count was reached. This setup ensures that all methods are evaluated at their respective optimal performance points.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparison of deep learning text classification models</title>
</caption>
<table>
<colgroup>
<col align="center" width="35mm"/>
<col align="center" width="31mm"/>
<col align="center" width="35mm"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Accuracy (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>BiGRU [<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
<td>87.08</td>
<td>87.07</td>
</tr>
<tr>
<td>BiLSTM [<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>89.19</td>
<td>89.20</td>
</tr>
<tr>
<td>MacBERT [<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
<td>88.34</td>
<td>88.34</td>
</tr>
<tr>
<td>ELECTRA [<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
<td>87.17</td>
<td>87.14</td>
</tr>
<tr>
<td>FastText [<xref ref-type="bibr" rid="ref-62">62</xref>]</td>
<td>89.25</td>
<td>89.26</td>
</tr>
<tr>
<td>Transformer [<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td>90.15</td>
<td>90.12</td>
</tr>
<tr>
<td>BERT [<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td>89.50</td>
<td>88.50</td>
</tr>
<tr>
<td>ModernBERT [<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
<td>90.68</td>
<td>87.26</td>
</tr>
<tr>
<td>DeBERTa-v3 [<xref ref-type="bibr" rid="ref-66">66</xref>]</td>
<td>89.82</td>
<td>91.41</td>
</tr>
<tr>
<td>RoPE [<xref ref-type="bibr" rid="ref-67">67</xref>]</td>
<td>88.37</td>
<td>89.34</td>
</tr>
<tr>
<td>RoBERTa [<xref ref-type="bibr" rid="ref-68">68</xref>]</td>
<td>89.42</td>
<td>90.24</td>
</tr>
<tr>
<td>ASSC-TextRCNN</td>
<td>91.47</td>
<td>91.49</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>ASSC-TextRCNN achieves the best performance on both metrics, with accuracy and F1 reaching 91.47% and 91.49%, respectively&#x2014;an absolute gain of &#x002B;1.32/&#x002B;1.37 percentage points over the strongest current comparator, Transformer (90.15%/90.12%), corresponding to an approximate 13.4% reduction in relative error rate. Compared with FastText (89.25%/89.26%) and BERT (89.50%/88.50%), the accuracy gains are &#x002B;2.22 and &#x002B;1.97 points, and the F1 gains are &#x002B;2.23 and &#x002B;2.99 points, respectively, representing a more pronounced advantage over earlier RNN-based models and related pretrained variants. The concurrent improvements in accuracy and F1 indicate that the gains do not stem from thresholding or incidental class-distribution effects, but from systematic enhancements in representation quality and decision-boundary separability. In latent space, SVD performs low-rank compression that de-redundifies and denoises long texts and topic-mixed scenarios, highlighting sentence-level stable semantic factors; on this foundation, the attention mechanism further focuses on discriminative segments and mitigates confusions caused by cross-class lexical co-occurrence. Consequently, consistent gains are realized atop strong baselines spanning diverse architectures and training paradigms. Together with the evidence from the confusion matrices and per-class metrics, ASSC-TextRCNN substantially boosts recall for weaker classes while further reducing residual errors for stronger classes, demonstrating broad-spectrum efficacy and robust generalization.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion and Conclusions</title>
<sec id="s5_1">
<label>5.1</label>
<title>Research Significance</title>
<p>Addressing the challenge that digital-culture texts&#x2014;characterized by being sparse, short, noisy, and highly time-sensitive&#x2014;are difficult to model effectively with traditional methods, this study proposes and systematically investigates a classification framework based on ASSC-TextRCNN. On top of TextRCNN&#x2019;s context-aware convolution&#x2013;recurrent architecture, we introduce ALBERT as a pretrained semantic encoder to enhance deep semantic alignment; we design a short-text&#x2013;oriented dual-attention mechanism that adaptively focuses on latent key information along the word&#x2013;context and channel&#x2013;semantics dimensions; we replace traditional max pooling with singular value decomposition (SVD) to mitigate feature loss from extreme pooling and preserve discriminative subspaces; and we adopt cross-entropy loss to robustly capture class distributions. These architectural refinements strike a better balance among representation depth, information fidelity, and discriminability.</p>
<p>Experiments demonstrate that ASSC-TextRCNN substantially outperforms multiple mainstream deep-learning baselines on digital-culture text classification: relative to the baseline, accuracy improves by 11.85%, F1 by 11.97%, and the relative error rate decreases by 53.18%; the confusion matrix exhibits overall contraction, and class boundaries become clearer. These results indicate that deep coupling of pretrained semantics with contextual structure effectively alleviates the information scarcity of short texts; dual attention can reliably distill key cues under sparsity and noise, enhancing sensitivity to fine-grained class differences; SVD-based subspace compression and reconstruction preserve global discriminative information better than max pooling, improving robustness; and end-to-end optimization with cross-entropy enables stable decision surfaces even under class imbalance and fuzzy boundaries.</p>
<p>Robust understanding of short and nonstandard expressions improves automatic archiving, topic clustering, and metadata annotation for cultural materials. More accurate, fine-grained classification markedly enhances topic routing and audience targeting efficiency, reducing exposure waste from misclassification&#x2013;misalignment and activating the long tail of high-quality cultural content for precise reach. The &#x201C;structure-preserving &#x002B; key-selection&#x201D; synergy formed by dual attention and SVD effectively suppresses noise interference and semantic drift, supporting public-opinion early warning, content compliance review, and intellectual-property protection for more fine-grained governance. Stable classification signals provide quantifiable evidence for hotspot discovery, issue-evolution tracking, and curation optimization, forming a data-driven, closed-loop content-operations pipeline. ASSC-TextRCNN embodies a synergistic paradigm&#x2014;from deep semantic pretraining to attention-based selection, structured compression, and robust optimization&#x2014;validating the effectiveness of structure-preserving dimensionality reduction (SVD) &#x002B; attention-based discrimination and furnishing core infrastructure for organized representation and at-scale distribution of digital-culture content in complex settings. With stronger semantic capacity and lower misclassification costs, the model advances a shift from experience-driven to data&#x2013;algorithm co-driven cultural dissemination, accelerating value realization and the sustainable amplification of social impact.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Research Outlook</title>
<p>Given the complex ecology of digital-culture texts&#x2014;sparse, short, noisy, and time-critical&#x2014;future work on ASSC-TextRCNN should both deepen the technical pathway and broaden cultural applications. On the technical front, we will first tackle the core issue of inadequate information in short texts by exploring a synergy of external knowledge, context modeling, and attention-based selection: atop the current ALBERT backbone, incorporate retrieval-augmented and knowledge-graph&#x2013;constrained semantic completion, injecting structured relations among cultural entities into encoding and decoding; aggregate multi-granularity context across paragraphs and conversations to mitigate semantic sparsity in individual short texts. Building on dual attention, we will develop interpretable hierarchical attention, enabling causal visualization and controllable guidance of the word&#x2013;context and channel&#x2013;semantics foci, thus providing a traceable evidence chain for cultural concepts and narrative themes. In line with the structure-preserving idea behind replacing max pooling with SVD, we will investigate dynamic low-rank decomposition and randomized approximations to reduce latency and memory while maintaining the integrity of the discriminative space, and coordinate with parameter-efficient finetuning to ensure deployability and scalability under platform-level, high-concurrency workloads.</p>
<p>Digital-culture content exhibits substantial diversity in topic, style, and genre, alongside pronounced distributional drift; models must retain stable memory while updating flexibly over time: employ time-aware train/validation splits and imbalance-aware long-tail reweighting to boost recognition of emerging subcultures and niche themes; adapt across languages and dialects to build a multilingual, multi-domain shared semantic subspace, using contrastive learning to align styles across platforms and genres and thereby reduce out-of-domain performance collapse. Weak and distant supervision can rapidly expand label coverage, while active learning&#x2014;via uncertainty and cost-sensitive sampling&#x2014;guides human annotation to form an efficient human-in-the-loop pipeline, significantly lowering the cost of introducing new classes and topics.</p>
<p>For interpretability, jointly visualize the contributions of dual attention and the low-rank subspace to clarify which words/subspaces are pivotal to final decisions. Uncertainty estimation and probability calibration will help the system respond cautiously in high-risk or semantically ambiguous scenarios, reducing governance risks from misclassification. Regarding adversarial and noise robustness, systematically evaluate the impact of colloquialisms, orthographic variants, detection-evading metaphors and homophones, and machine-translation noise on decision boundaries, and reinforce with data augmentation and robust optimization to bolster resilience in real-world dissemination environments. For large-scale deployment, integrate fairness and bias auditing end-to-end, with targeted evaluations for minority-language communities, regional dialects, and youth/elderly users, preventing the technological amplification of existing discourse imbalances.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The APC was funded by China National Innovation and Entrepreneurship Project Fund Innovation Training Program (202410451009).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Zixuan Guo: conceptualization, methodology, writing; Houbin Wang: conceptualization, methodology, computation, analysis, visualization, writing; Yuanfang Chen: validation, writing, review, editing; Sameer Kumar: conceptualization, review, supervisor, fund acquisition. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The authors confirm that the data supporting the findings of this study are available within the article.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations</title>
<def-list>
<def-item>
<term>SVD</term>
<def>
<p>Singular Value Decomposition</p>
</def>
</def-item>
<def-item>
<term>ASR</term>
<def>
<p>Automatic speech recognition</p>
</def>
</def-item>
<def-item>
<term>SVM</term>
<def>
<p>Support vector machine</p>
</def>
</def-item>
<def-item>
<term>CNN</term>
<def>
<p>Convolutional neural network</p>
</def>
</def-item>
<def-item>
<term>EDA</term>
<def>
<p>Easy Data Augmentation</p>
</def>
</def-item>
<def-item>
<term>UDA</term>
<def>
<p>Unsupervised Data Augmentation</p>
</def>
</def-item>
<def-item>
<term>NLP</term>
<def>
<p>Natural language processing</p>
</def>
</def-item>
<def-item>
<term>ASSC</term>
<def>
<p>ALBERT, SVD, Self-Attention and Cross-Entropy</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Research on innovative development strategies of cultural industry under digital transformation</article-title>. In: <conf-name>Proceedings of 2024 6th International Conference on Economic Management and Cultural Industry (ICEMCI 2024); 2024 Nov 1&#x2013;3</conf-name>; <publisher-loc>Chengdu, China. Dordrecht, the Netherlands</publisher-loc>: <publisher-name>Atlantis Press International BV</publisher-name>; <year>2025</year>. p. <fpage>451</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.2991/978-94-6463-642-0_48</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hong</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Digital-media-based interaction and dissemination of traditional culture integrating using social media data analytics</article-title>. <source>Comput Intell Neurosci</source>. <year>2022</year>;<volume>2022</volume>:<fpage>5846451</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2022/5846451</pub-id>; <pub-id pub-id-type="pmid">35685140</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Correia</surname> <given-names>RA</given-names></string-name>, <string-name><surname>Ladle</surname> <given-names>R</given-names></string-name>, <string-name><surname>Jari&#x0107;</surname> <given-names>I</given-names></string-name>, <string-name><surname>Malhado</surname> <given-names>ACM</given-names></string-name>, <string-name><surname>Mittermeier</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Roll</surname> <given-names>U</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Digital data sources and methods for conservation culturomics</article-title>. <source>Conserv Biol</source>. <year>2021</year>;<volume>35</volume>(<issue>2</issue>):<fpage>398</fpage>&#x2013;<lpage>411</lpage>. doi:<pub-id pub-id-type="doi">10.1111/cobi.13706</pub-id>; <pub-id pub-id-type="pmid">33749027</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>George</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Baskar</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Leveraging big data and sentiment analysis for actionable insights: a review of data mining approaches for social media</article-title>. <source>Partn Univers Int Innov J</source>. <year>2024</year>;<volume>2</volume>(<issue>4</issue>):<fpage>39</fpage>&#x2013;<lpage>59</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>E</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Metaphors as semantic anchors: a label-constrained contrastive learning approach for Chinese text classification</article-title>. <source>IEEE Trans Comput Soc Syst</source>. <year>2025</year>;<volume>12</volume>(<issue>6</issue>):<fpage>5080</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tcss.2025.3585577</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xing</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Cultural bias mitigation in vision-language models for digital heritage documentation: a comparative analysis of debiasing techniques</article-title>. <source>Artif Intell Mach Learn Rev</source>. <year>2024</year>;<volume>5</volume>(<issue>3</issue>):<fpage>28</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.69987/aimlr.2024.50303</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gani Ganie</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dadvandipour</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Exploring the impact of informal language on sentiment analysis models for social media text using convolutional neural networks</article-title>. <source>Multidiszcip Tudom&#x00E1;nyok</source>. <year>2023</year>;<volume>13</volume>(<issue>1</issue>):<fpage>244</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.35925/j.multi.2023.1.17</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Baviskar</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ahirrao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Potdar</surname> <given-names>V</given-names></string-name>, <string-name><surname>Kotecha</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Efficient automated processing of the unstructured documents using artificial intelligence: a systematic literature review and future directions</article-title>. <source>IEEE Access</source>. <year>2021</year>;<volume>9</volume>:<fpage>72894</fpage>&#x2013;<lpage>936</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2021.3072900</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ranjgar</surname> <given-names>B</given-names></string-name>, <string-name><surname>Sadeghi-Niaraki</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shakeri</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rahimi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>SM</given-names></string-name></person-group>. <article-title>Cultural heritage information retrieval: past, present, and future trends</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>42992</fpage>&#x2013;<lpage>3026</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2024.3374769</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Minaee</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kalchbrenner</surname> <given-names>N</given-names></string-name>, <string-name><surname>Cambria</surname> <given-names>E</given-names></string-name>, <string-name><surname>Nikzad</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chenaghlu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep learning&#x2014;based text classification: a comprehensive review</article-title>. <source>ACM Comput Surv</source>. <year>2022</year>;<volume>54</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3439726</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chatzimparmpas</surname> <given-names>A</given-names></string-name>, <string-name><surname>Martins</surname> <given-names>RM</given-names></string-name>, <string-name><surname>Kerren</surname> <given-names>A</given-names></string-name></person-group>. <article-title>VisRuler: visual analytics for extracting decision rules from bagged and boosted decision trees</article-title>. <source>Inf Vis</source>. <year>2023</year>;<volume>22</volume>(<issue>2</issue>):<fpage>115</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1177/14738716221142005</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hollmann</surname> <given-names>N</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hutter</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Large language models for automated data science: introducing caafe for context-aware automated feature engineering</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2023</year>;<volume>36</volume>:<fpage>44753</fpage>&#x2013;<lpage>75</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Chinnaiyan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Balasubaramanian</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jeyabalu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Warrier</surname> <given-names>GS</given-names></string-name></person-group>. <chapter-title>AI applications-computer vision and natural language processing</chapter-title>. In: <source>Model optimization methods for efficient and edge AI: federated learning architectures, frameworks and applications</source>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons, Inc.</publisher-name>; <year>2025</year>. p. <fpage>25</fpage>&#x2013;<lpage>41</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Interpretable deep learning: interpretation, interpretability, trustworthiness, and beyond</article-title>. <source>Knowl Inf Syst</source>. <year>2022</year>;<volume>64</volume>(<issue>12</issue>):<fpage>3197</fpage>&#x2013;<lpage>234</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10115-022-01756-8</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>R</given-names></string-name>, <string-name><surname>He</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Huo</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Asian and low-resource language information processing</article-title>. <source>ACM Trans</source>. <year>2024</year>;<volume>23</volume>(<issue>4</issue>):<fpage>3613605</fpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Naithani</surname> <given-names>K</given-names></string-name>, <string-name><surname>Raiwani</surname> <given-names>YP</given-names></string-name></person-group>. <article-title>Realization of natural language processing and machine learning approaches for text-based sentiment analysis</article-title>. <source>Expert Syst</source>. <year>2023</year>;<volume>40</volume>(<issue>5</issue>):<fpage>e13114</fpage>. doi:<pub-id pub-id-type="doi">10.1111/exsy.13114</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sarker</surname> <given-names>A</given-names></string-name>, <string-name><surname>O&#x2019;Connor</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ginn</surname> <given-names>R</given-names></string-name>, <string-name><surname>Scotch</surname> <given-names>M</given-names></string-name>, <string-name><surname>Smith</surname> <given-names>K</given-names></string-name>, <string-name><surname>Malone</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Social media mining for toxicovigilance: automatic monitoring of prescription medication abuse from twitter</article-title>. <source>Drug Saf</source>. <year>2016</year>;<volume>39</volume>(<issue>3</issue>):<fpage>231</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s40264-015-0379-4</pub-id>; <pub-id pub-id-type="pmid">26748505</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Building multimodal knowledge bases with multimodal computational sequences and generative adversarial networks</article-title>. <source>IEEE Trans Multimed</source>. <year>2024</year>;<volume>26</volume>:<fpage>2027</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TMM.2023.3291503</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alizadeh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Seilsepour</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A novel self-supervised sentiment classification approach using semantic labeling based on contextual embeddings</article-title>. <source>Multimed Tools Appl</source>. <year>2025</year>;<volume>84</volume>(<issue>12</issue>):<fpage>10195</fpage>&#x2013;<lpage>220</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-024-19086-y</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bashiri</surname> <given-names>H</given-names></string-name>, <string-name><surname>Naderi</surname> <given-names>H</given-names></string-name></person-group>. <article-title>SyntaPulse: an unsupervised framework for sentiment annotation and semantic topic extraction</article-title>. <source>Pattern Recognit</source>. <year>2025</year>;<volume>164</volume>:<fpage>111593</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2025.111593</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>L</given-names></string-name></person-group>. <article-title>An adaptive LDA optimal topic number selection method in news topic identification</article-title>. <source>IEEE Access</source>. <year>2023</year>;<volume>11</volume>:<fpage>92273</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2023.3308520</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>T</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>F</given-names></string-name></person-group>. <article-title>LayoutLMv3: pre-training for document AI with unified text and image masking</article-title>. In: <conf-name>Proceedings of the 30th ACM International Conference on Multimedia; 2022 Oct 10&#x2013;14</conf-name>; <publisher-loc>Lisboa, Portugal. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>; <year>2022</year>. p. <fpage>4083</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3503161.3548112</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on text classification: from traditional to deep learning</article-title>. <source>ACM Trans Intell Syst Technol</source>. <year>2022</year>;<volume>13</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3495162</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Breit</surname> <given-names>A</given-names></string-name>, <string-name><surname>Waltersdorfer</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ekaputra</surname> <given-names>FJ</given-names></string-name>, <string-name><surname>Sabou</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ekelhart</surname> <given-names>A</given-names></string-name>, <string-name><surname>Iana</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Combining machine learning and semantic web: a systematic mapping study</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>14s</issue>):<fpage>1</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3586163</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bu</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A comprehensive survey on deep clustering: taxonomy, challenges, and future directions</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>57</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3689036</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Song</surname> <given-names>B</given-names></string-name></person-group>. <article-title>A novel approach for multiclass sentiment analysis on Chinese social media with ERNIE-MCBMA</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>18675</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-03875-y</pub-id>; <pub-id pub-id-type="pmid">40437070</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Soni</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chouhan</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Rathore</surname> <given-names>SS</given-names></string-name></person-group>. <article-title>TextConvoNet: a convolutional neural network based architecture for text classification</article-title>. <source>Appl Intell</source>. <year>2023</year>;<volume>53</volume>(<issue>11</issue>):<fpage>14249</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10489-022-04221-9</pub-id>; <pub-id pub-id-type="pmid">36310755</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Akpatsa</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>H</given-names></string-name></person-group>. <chapter-title>A survey and future perspectives of hybrid deep learning models for text classification</chapter-title>. In: <source>Artificial intelligence and security</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2021</year>. p. <fpage>358</fpage>&#x2013;<lpage>69</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-78609-0_31</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Multilingual pretrained based multi-feature fusion model for English text classification</article-title>. <source>Comput Sci Inf Syst</source>. <year>2025</year>;<volume>22</volume>(<issue>1</issue>):<fpage>133</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.2298/csis240630004z</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Long</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Deep learning for reddit text classification: TextCNN and TextRNN approaches</article-title>. In: <conf-name>Proceedings of the 2024 4th Interdisciplinary Conference on Electrics and Computer (INTCEC); 2024 Jun 11&#x2013;13</conf-name>; <publisher-loc>Chicago, IL, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/intcec61833.2024.10603240</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bhadauria</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Shukla</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>P</given-names></string-name>, <string-name><surname>Dwivedi</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Attention-based neural machine translation for multilingual communication</article-title>. In: <conf-name>Proceedings of the Fifth International Conference on Trends in Computational and Cognitive Engineering; 2023 Nov 24&#x2013;25</conf-name>; <publisher-loc>K&#x0101;npur, India</publisher-loc>. p. <fpage>129</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-97-1923-5_10</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sahu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Vechtomova</surname> <given-names>O</given-names></string-name>, <string-name><surname>Bahdanau</surname> <given-names>D</given-names></string-name>, <string-name><surname>Laradji</surname> <given-names>I</given-names></string-name></person-group>. <article-title>PromptMix: a class boundary augmentation method for large language model distillation</article-title>. In: <conf-name>Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; 2023 Dec 6&#x2013;10</conf-name>; <publisher-loc>Singapore. Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2023</year>. p. <fpage>5316</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2023.emnlp-main.323</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sherin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jasmine Selvakumari Jeya</surname> <given-names>I</given-names></string-name>, <string-name><surname>Deepa</surname> <given-names>SN</given-names></string-name></person-group>. <article-title>Enhanced Aquila optimizer combined ensemble Bi-LSTM-GRU with fuzzy emotion extractor for tweet sentiment analysis and classification</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>141932</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3464091</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kamyab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Rasool</surname> <given-names>A</given-names></string-name>, <string-name><surname>Adjeisah</surname> <given-names>M</given-names></string-name></person-group>. <article-title>ACR-SA: attention-based deep model through two-channel CNN and Bi-RNN for sentiment analysis</article-title>. <source>PeerJ Comput Sci</source>. <year>2022</year>;<volume>8</volume>:<fpage>e877</fpage>. doi:<pub-id pub-id-type="doi">10.7717/peerj-cs.877</pub-id>; <pub-id pub-id-type="pmid">35494855</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mars</surname> <given-names>M</given-names></string-name></person-group>. <article-title>From word embeddings to pre-trained language models: a state-of-the-art walkthrough</article-title>. <source>Appl Sci</source>. <year>2022</year>;<volume>12</volume>(<issue>17</issue>):<fpage>8805</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app12178805</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Learning credible DNNs via incorporating prior knowledge and model local explanation</article-title>. <source>Knowl Inf Syst</source>. <year>2021</year>;<volume>63</volume>(<issue>2</issue>):<fpage>305</fpage>&#x2013;<lpage>32</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10115-020-01517-5</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Channel attention TextCNN with feature word extraction for Chinese sentiment analysis</article-title>. <source>ACM Trans Asian Low-Resour Lang Inf Process</source>. <year>2023</year>;<volume>22</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3571716</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Paradigm construction and improvement strategy of English long text translation based on self-attention mechanism</article-title>. <source>J Combin Math Combin Comput</source>. <year>2025</year>;<volume>127</volume>:<fpage>8945</fpage>&#x2013;<lpage>60</lpage>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>L</given-names></string-name></person-group>. <article-title>New bagging based ensemble learning algorithm distinguishing short and long texts for document classification</article-title>. <source>ACM Trans Asian Low-Resour Lang Inf Process</source>. <year>2025</year>;<volume>24</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3718740</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Attieh</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tekli</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Supervised term-category feature weighting for improved text classification</article-title>. <source>Knowl Based Syst</source>. <year>2023</year>;<volume>261</volume>:<fpage>110215</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2022.110215</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rathi</surname> <given-names>RN</given-names></string-name>, <string-name><surname>Mustafi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>The importance of Term Weighting in semantic understanding of text: a review of techniques</article-title>. <source>Multimed Tools Appl</source>. <year>2023</year>;<volume>82</volume>(<issue>7</issue>):<fpage>9761</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-022-12538-3</pub-id>; <pub-id pub-id-type="pmid">35437420</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>An improved term weighting method based on relevance frequency for text classification</article-title>. <source>Soft Comput</source>. <year>2023</year>;<volume>27</volume>(<issue>7</issue>):<fpage>3563</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00500-022-07597-5</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Addressing class-imbalance challenges in cross-lingual aspect-based sentiment analysis: dynamic weighted loss and anti-decoupling</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>257</volume>:<fpage>125059</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.125059</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A novel textual data augmentation method for identifying comparative text from user-generated content</article-title>. <source>Electron Commer Res Appl</source>. <year>2022</year>;<volume>53</volume>:<fpage>101143</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.elerap.2022.101143</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Qutaber: task-based exploratory data analysis with enriched context awareness</article-title>. <source>J Vis</source>. <year>2024</year>;<volume>27</volume>(<issue>3</issue>):<fpage>503</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s12650-024-00975-1</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Body</surname> <given-names>T</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Using back-and-forth translation to create artificial augmented textual data for sentiment analysis models</article-title>. <source>Expert Syst Appl</source>. <year>2021</year>;<volume>178</volume>:<fpage>115033</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2021.115033</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Data-driven unsupervised anomaly detection of manufacturing processes with multi-scale prototype augmentation and multi-sensor data</article-title>. <source>J Manuf Syst</source>. <year>2024</year>;<volume>77</volume>:<fpage>26</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jmsy.2024.08.027</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>M</given-names></string-name>, <string-name><surname>You</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Ressl: relational self-supervised learning with weak augmentation</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2021</year>;<volume>34</volume>:<fpage>2543</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Research on multi channel text classification method based on neural network</article-title>. In: <conf-name>Proceedings of the 2024 14th International Conference on Information Technology in Medicine and Education (ITME); 2024 Sep 13&#x2013;15</conf-name>; <publisher-loc>Guiyang, China</publisher-loc>. p. <fpage>850</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/itme63426.2024.00171</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Research on high precision Chinese text sentiment Classification based on ALBERT Optimization</article-title>. In: <conf-name>Proceedings of the 2023 15th International Conference on Advanced Computational Intelligence (ICACI); 2023 May 6&#x2013;9</conf-name>; <publisher-loc>Seoul, Republic of Korea</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icaci58115.2023.10146191</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>JX</given-names></string-name>, <string-name><surname>Tu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Laskar</surname> <given-names>MTR</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Utilizing BERT for information retrieval: survey, applications, resources, and challenges</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>56</volume>(<issue>7</issue>):<fpage>1</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3648471</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ming</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Visuals to text: a comprehensive review on automatic image captioning</article-title>. <source>IEEE/CAA J Autom Sin</source>. <year>2022</year>;<volume>9</volume>(<issue>8</issue>):<fpage>1339</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JAS.2022.105734</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jia</surname> <given-names>W</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Feature dimensionality reduction: a review</article-title>. <source>Complex Intell Syst</source>. <year>2022</year>;<volume>8</volume>(<issue>3</issue>):<fpage>2663</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s40747-021-00637-x</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mao</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mohri</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Cross-entropy loss functions: theoretical analysis and applications</article-title>. In: <conf-name>Proceedings of the International Conference on Machine Learning; 2023 Jul 23&#x2013;29</conf-name>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>23803</fpage>&#x2013;<lpage>28</lpage>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Magee</surname> <given-names>CL</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep learning for technical document classification</article-title>. <source>IEEE Trans Eng Manage</source>. <year>2024</year>;<volume>71</volume>:<fpage>1163</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tem.2022.3152216</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Comparison of sarcasm detection models based on lightweight BERT model&#x2014;differences between ALBERT-Chinese-tiny and TinyBERT in small sample scenarios</article-title>. <source>Sci Technol Eng Chem Environ Prot</source>. <year>2025</year>;<volume>1</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.61173/6jpm4n27</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on scheduling techniques in computing and network convergence</article-title>. <source>IEEE Commun Surv Tutorials</source>. <year>2024</year>;<volume>26</volume>(<issue>1</issue>):<fpage>160</fpage>&#x2013;<lpage>95</lpage>. doi:<pub-id pub-id-type="doi">10.1109/comst.2023.3329027</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Attention-based BiGRU-CNN for Chinese question classification</article-title>. <source>J Ambient Intell Humaniz Comput</source>. <year>2019</year>;<fpage>1</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s12652-019-01344-9</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Joshi</surname> <given-names>VM</given-names></string-name>, <string-name><surname>Ghongade</surname> <given-names>RB</given-names></string-name>, <string-name><surname>Joshi</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Kulkarni</surname> <given-names>RV</given-names></string-name></person-group>. <article-title>Deep BiLSTM neural network model for emotion detection using cross-dataset approach</article-title>. <source>Biomed Signal Process Control</source>. <year>2022</year>;<volume>73</volume>:<fpage>103407</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2021.103407</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shao</surname> <given-names>D</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Chinese named entity recognition based on MacBERT and Joint Learning</article-title>. In: <conf-name>Proceedings of the 2023 IEEE 9th International Conference on Cloud Computing and Intelligent Systems (CCIS); 2023 Aug 12&#x2013;13</conf-name>; <publisher-loc>Dali, China</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ccis59572.2023.10262869</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mala</surname> <given-names>JB</given-names></string-name>, <string-name><surname>Anisha Angel</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>Alex Raj</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Rajan</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Efficacy of ELECTRA-based language model in sentiment analysis</article-title>. In: <conf-name>Proceedings of the 2023 International Conference on Intelligent Systems for Communication, IoT and Security (ICISCoIS); 2023 Feb 9&#x2013;11</conf-name>; <publisher-loc>Coimbatore, India</publisher-loc>. p. <fpage>682</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICISCoIS56541.2023.10100342</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Text classification model based on fastText</article-title>. In: <conf-name>Proceedings of the 2020 IEEE International Conference on Artificial Intelligence and Information Systems (ICAIIS); 2020 Mar 20&#x2013;22</conf-name>; <publisher-loc>Dalian, China</publisher-loc>. p. <fpage>154</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icaiis49377.2020.9194939</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rahali</surname> <given-names>A</given-names></string-name>, <string-name><surname>Akhloufi</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>End-to-end transformer-based models in textual-based NLP</article-title>. <source>AI</source>. <year>2023</year>;<volume>4</volume>(<issue>1</issue>):<fpage>54</fpage>&#x2013;<lpage>110</lpage>. doi:<pub-id pub-id-type="doi">10.3390/ai4010004</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Garrido-Merchan</surname> <given-names>EC</given-names></string-name>, <string-name><surname>Gozalo-Brizuela</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gonzalez-Carvajal</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Comparing BERT against traditional machine learning models in text classification</article-title>. <source>J Comput Cogn Eng</source>. <year>2023</year>;<volume>2</volume>(<issue>4</issue>):<fpage>352</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.47852/bonviewjcce3202838</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Molokwu</surname> <given-names>BC</given-names></string-name>, <string-name><surname>Rahimi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Molokwu</surname> <given-names>RC</given-names></string-name></person-group>. <article-title>Fine-grained sentiment mining, at document level on big data, using a state-of-the-art representation-based transformer: ModernBERT [Internet]</article-title>. <year>2025 [cited 2025 Dec 1]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5418377">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5418377</ext-link>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Team Bohan Li at PAN: DeBERTa-v3 with R-drop regularization for human-AI collaborative text classification</article-title>. In: <conf-name>Proceedings of the Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2024); 2024 Sep 9&#x2013;12</conf-name>; <publisher-loc>Grenoble, France</publisher-loc>. p. <fpage>2658</fpage>&#x2013;<lpage>64</lpage>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bo</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>RoFormer: enhanced transformer with rotary position embedding</article-title>. <source>Neurocomputing</source>. <year>2024</year>;<volume>568</volume>:<fpage>127063</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2023.127063</pub-id>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Adoma</surname> <given-names>AF</given-names></string-name>, <string-name><surname>Henry</surname> <given-names>NM</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Comparative analyses of bert, roberta, distilbert, and xlnet for text-based emotion recognition</article-title>. In: <conf-name>Proceedings of the 2020 17th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP); 2020 Dec 18&#x2013;20</conf-name>; <publisher-loc>Chengdu, China</publisher-loc>. p. <fpage>117</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccwamtip51612.2020.9317379</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>