<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">75119</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2025.075119</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Neuro-Symbolic Graph Learning for Causal Inference and Continual Learning in Mental-Health Risk Assessment</article-title>
<alt-title alt-title-type="left-running-head">Neuro-Symbolic Graph Learning for Causal Inference and Continual Learning in Mental-Health Risk Assessment</alt-title>
<alt-title alt-title-type="right-running-head">Neuro-Symbolic Graph Learning for Causal Inference and Continual Learning in Mental-Health Risk Assessment</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Jena</surname><given-names>Monalisa</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Khan</surname><given-names>Noman</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>noman@yonsei.ac.kr</email></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Lee</surname><given-names>Mi Young</given-names></name><xref ref-type="aff" rid="aff-3">3</xref><email>miylee@cau.ac.kr</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Rho</surname><given-names>Seungmin</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science, Fakir Mohan University</institution>, <addr-line>Balasore, 756019, Odisha</addr-line>, <country>India</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Architecture and Architectural Engineering, Yonsei University</institution>, <addr-line>Seoul, 03722</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Industrial Security, Chung-Ang University</institution>, <addr-line>Seoul, 06974</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Noman Khan. Email: <email>noman@yonsei.ac.kr</email>; Mi Young Lee. Email: <email>miylee@cau.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>29</day><month>1</month><year>2026</year>
</pub-date>
<volume>146</volume>
<issue>1</issue>
<elocation-id>44</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>12</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_75119.pdf"></self-uri>
<abstract>
<p>Mental-health risk detection seeks early signs of distress from social media posts and clinical transcripts to enable timely intervention before crises. When such risks go undetected, consequences can escalate to self-harm, long-term disability, reduced productivity, and significant societal and economic burden. Despite recent advances, detecting risk from online text remains challenging due to heterogeneous language, evolving semantics, and the sequential emergence of new datasets. Effective solutions must encode clinically meaningful cues, reason about causal relations, and adapt to new domains without forgetting prior knowledge. To address these challenges, this paper presents a Continual Neuro-Symbolic Graph Learning (CNSGL) framework that unifies symbolic reasoning, causal inference, and continual learning within a single architecture. Each post is represented as a symbolic graph linking clinically relevant tags to textual content, enriched with causal edges derived from directional Point-wise Mutual Information (PMI). A two-layer Graph Convolutional Network (GCN) encodes these graphs, and a Transformer-based attention pooler aggregates node embeddings while providing interpretable tag-level importances. Continual adaptation across datasets is achieved through the Multi-Head Freeze (MH-Freeze) strategy, which freezes a shared encoder and incrementally trains lightweight task-specific heads (small classifiers attached to the shared embedding). Experimental evaluations across six diverse mental-health datasets ranging from Reddit discourse to clinical interviews, demonstrate that MH-Freeze consistently outperforms existing continual-learning baselines in both discriminative accuracy and calibration reliability. Across six datasets, MH-Freeze achieves up to 0.925 accuracy and 0.923 F1-Score, with AUPRC <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo>&#x2265;</mml:mo><mml:mn>0.934</mml:mn></mml:math></inline-formula> and AUROC <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mo>&#x2265;</mml:mo><mml:mn>0.942</mml:mn></mml:math></inline-formula>, consistently surpassing all continual-learning baselines. The results confirm the framework&#x2019;s ability to preserve prior knowledge, adapt to domain shifts, and maintain causal interpretability, establishing CNSGL as a promising step toward robust, explainable, and lifelong mental-health risk assessment.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Catastrophic forgetting</kwd>
<kwd>causal inference</kwd>
<kwd>continual learning</kwd>
<kwd>deep learning</kwd>
<kwd>graph convolutional network</kwd>
<kwd>mental health monitoring</kwd>
<kwd>transformer</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Research Foundation of Korea</funding-source>
<award-id>RS-2025-00518960</award-id>
<award-id>RS-2025-00563192</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Mental-health disorders such as depression, anxiety, and suicidal ideation are rising globally, posing serious risks to individuals and society [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. With the growth of online platforms and clinical records, vast amounts of unstructured text now capture personal experiences and distress signals [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>]. Automatically analyzing this text for early detection of psychological risk has become an urgent research problem. However, it remains highly challenging due to the ambiguity of natural language, and the subtlety of psychological cues [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. Timely detection enables early intervention and targeted allocation of limited mental-health resources, and real-world deployment demands models that are not only accurate but also interpretable and well-calibrated so that clinicians and moderators can act on predictions with confidence [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>In many real-world scenarios, posts and interviews carry implicit cues (e.g., hopelessness, insomnia, self-harm) that are difficult to capture with surface features alone [<xref ref-type="bibr" rid="ref-8">8</xref>]. Purely neural approaches can learn powerful representations but often lack interpretability and causal grounding; purely symbolic methods offer transparency but struggle with linguistic variability and generalization [<xref ref-type="bibr" rid="ref-9">9</xref>]. Moreover, data arrive over time from different communities and collection protocols, creating domain shift and exposing models to <italic>catastrophic forgetting</italic> when retrained sequentially [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>]. These factors emphasize an integrated approach that can (i) structure free text into clinically meaningful graphs, (ii) model directional relations among risk factors, and (iii) learn continually across datasets without erasing earlier competencies. In addition, class imbalance, where rare but critical signals such as suicidal ideation are underrepresented, biases predictions and reduces reliability. Addressing these issues requires models that are both <italic>interpretable</italic> and <italic>adaptable</italic>, while retaining stability across evolving datasets.</p>
<p>Continual learning (CL), also referred to as <italic>lifelong</italic> or <italic>incremental learning</italic>, aims to enable models to acquire new knowledge over time without forgetting previously learned information [<xref ref-type="bibr" rid="ref-12">12</xref>]. Unlike traditional retraining approaches that require access to all past data, CL supports sequential learning across tasks or domains by reusing shared representations and adapting to new inputs efficiently. This paradigm is particularly valuable in mental-health applications, where new linguistic trends, populations, and annotation protocols continuously emerge. A robust continual learning mechanism ensures that models remain up to date while preserving earlier competencies, enabling sustainable and realistic deployment in evolving digital health environments.</p>
<p>To address these challenges, a <italic>Continual Neuro-Symbolic Graph Learning (CNSGL)</italic> framework is proposed in this work for causal inference and continual learning in mental-health risk assessment. CNSGL represents each post as a symbolic graph in which a post node connects to tag nodes derived from a risk lexicon. Beyond simple co-occurrence, graphs are enriched with directional edges using a variant of point-wise mutual information to reflect likely precursors and consequents among risk factors. A two-layer Graph Convolutional Network (GCN) propagates information over this structure, and a lightweight Transformer attention pooler, anchored by a learnable <monospace>[CLS]</monospace> token, aggregates node embeddings while producing tag-level importances for interpretability.</p>
<p>To enable continual adaptation, the proposed framework employs a Multi-Head Freeze (MH-Freeze) strategy that freezes the shared encoder after the first dataset and incrementally attaches lightweight task-specific heads for subsequent datasets. Here, &#x201C;task-specific head&#x201D; refers to a small linear-sigmoid classifier attached to the shared embedding for each dataset. This form is adopted to enable lightweight adaptation on a fixed embedding space: the frozen GCN-Transformer encoder produces a stable representation, and the head maps it directly to a calibrated probability via binary cross-entropy (BCE) loss. This keeps updates simple and efficient, reduces the risk of cross-task gradient interference, and preserves calibration. In contrast, deeper or non-linear heads introduce extra trainable layers that can overfit to a single dataset and reintroduce interference with previously learned tasks. Each dataset is treated as a separate task (T1-T6), allowing systematic evaluation of domain transfer and retention. Our evaluation further incorporates both discrimination and calibration metrics (AUROC, AUPRC, Brier score, and Expected Calibration Error) to quantify predictive reliability under domain shift and sequential learning conditions.</p>
<sec id="s1_1">
<label>1.1</label>
<title>Motivation</title>
<p>In mental health risk assessment and monitoring, systems are deployed in dynamic and evolving scenarios. In many real world scenarios, diverse signals (clinical notes, social media text, speech, and wearable biosignals) are used to detect risk states such as depression, anxiety, self-harm intent, and acute stress. However, most pipelines are trained on static datasets with fixed labels and vocabularies. When new expressions, populations, or risk patterns appear, performance is often degraded. In practice, full retraining on new data is often required, while incremental updates risk catastrophic forgetting that overwrites previously learned knowledge [<xref ref-type="bibr" rid="ref-13">13</xref>]. Therefore, continual learning is increasingly regarded as important for mental health tasks. Rather than retraining from scratch, new risk categories or domains can be incorporated as they arise while preserving recognition of earlier ones [<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Research Gaps</title>
<p>Despite rapid progress, Several unresolved challenges continue to hinder dependable mental-health risk detection from text:
<list list-type="bullet">
<list-item>
<p>Missing causal structure: Most models treat symptoms as flat labels; they do not encode directed tag<inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>tag influences or use them during message passing.</p></list-item>
<list-item>
<p>Limited interpretability: Explanations are often post-hoc for text tokens, not concept-level (symbolic tags) nor pathway-level (causal paths).</p></list-item>
<list-item>
<p>Catastrophic forgetting: Models struggle to retain prior knowledge as new datasets arrive, while simple, deployable continual-learning solutions are still lacking.</p></list-item>
<list-item>
<p>Opaque pooling: Mean/max pooling blurs which symbolic tags matter per post; attention over concept nodes is rarely leveraged.</p></list-item>
</list></p>
<p>In light of the above, we propose a continual neuro-symbolic framework that builds per-post symbolic tag graphs with directed links, encodes them via a two-layer GCN, and uses a lightweight attention head to form calibrated, interpretable post representations. For sequential datasets, we adopt a frozen-encoder, multi-head protocol to prevent forgetting while keeping adaptation lightweight. The key contributions of this work are summarized below:
<list list-type="bullet">
<list-item>
<p>A neuro-symbolic graph learning framework is proposed that combines symbolic reasoning, causal inference, graph-based neural encoding, and continual learning for mental-health risk assessment.</p></list-item>
<list-item>
<p>Symbolic graphs are constructed from text using risk-related tags, ensuring interpretability by grounding predictions in clinically meaningful indicators.</p></list-item>
<list-item>
<p>A causal-aware enrichment mechanism introduces directed tag&#x2013;tag edges, capturing potential causal influences among symptoms rather than simple co-occurrence.</p></list-item>
<list-item>
<p>A graph convolutional encoder is employed to propagate symbolic and causal features, followed by a lightweight Transformer-based attention head weights the post&#x2013;tag embeddings and a classifier that outputs binary risk predictions through probability estimation and thresholding.</p></list-item>
<list-item>
<p>A continual learning strategy (multi-head frozen-encoder) is implemented to preserve knowledge across datasets while enabling adaptation to new domains, mitigating catastrophic forgetting.</p></list-item>
<list-item>
<p>Extensive experiments are conducted on multiple datasets, and results are compared against strong continual learning baselines, demonstrating improved robustness, interpretability, and adaptability.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> presents an extensive survey of related research and methods relevant to this work, <xref ref-type="sec" rid="s3">Section 3</xref> presents the proposed CNSGL framework, including symbolic graph construction, causal enrichment, the GCN&#x2013;Transformer encoder, and the MH-Freeze continual-learning strategy. <xref ref-type="sec" rid="s4">Section 4</xref> describes the experimental details, and evaluation metrics. <xref ref-type="sec" rid="s5">Section 5</xref> presents the ablation study analyzing the contribution of each component. <xref ref-type="sec" rid="s6">Section 6</xref> concludes the paper and highlights future directions.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Detecting mental-health risk from text is a challenging and active area with direct applications to screening, risk stratification, and clinical decision support. This section reviews related research across the categories listed in the subsections, aligning each category of methods pertinent to this work.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Mental-Health Risk Detection from Text</title>
<p>This subsection reviews key efforts on detecting mental-health signals from social-media text, ranging from traditional machine learning (ML) to deep learning (DL). Hemmatirad et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] showed that lexicon and handcrafted features paired with support vector machine or logistic regression classifiers can distinguish high-risk users using linguistic and emotional cues. With the advent of contextual embeddings, hybrid models such as BERT&#x002B;BiLSTM have been proposed by Zhou and Mohd [<xref ref-type="bibr" rid="ref-16">16</xref>] to better handle informal language, emojis, and sequential patterns in depression-related posts.</p>
<p>Prior surveys consolidate the literature and shed light on persistent challenges. Garg [<xref ref-type="bibr" rid="ref-17">17</xref>] reviewed 92 studies, introduced an updatable suicide-detection repository, and emphasized the need for real-time, responsible AI. Skaik and Inkpen [<xref ref-type="bibr" rid="ref-18">18</xref>] surveyed NLP/ML approaches for public mental-health surveillance, summarizing data collection strategies, modeling tools, and remaining gaps. Other studies explore emotion-aware and efficiency-focused systems. Benrouba and Boudour [<xref ref-type="bibr" rid="ref-19">19</xref>] proposed an emotion-aware content-filtering framework that classifies posts into basic emotions and compares them with an &#x201C;ideal&#x201D; lexicon to flag potentially harmful content. Ding et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] compared ML models (logistic regression, random forest, LightGBM) with DL models (ALBERT, GRU) for binary and multi-class mental-health classification, finding that ML methods offer better interpretability and efficiency on medium-sized datasets, whereas DL models better capture complex linguistic patterns.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Causal Inference</title>
<p>Causal reasoning in language means modeling directional influence (<inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>B</mml:mi></mml:math></inline-formula>) so that changing <italic>A</italic> would change the likelihood of <italic>B</italic>, beyond simple correlation. This has been studied using temporal precedence and directional association measures, causal discovery on event graphs, and counterfactual analyses. However, in social-media risk assessment, such causal structure is rarely embedded within the encoder itself. Choudhury and Kiciman [<xref ref-type="bibr" rid="ref-21">21</xref>] examined the causal impact of online social-support language in Reddit mental-health communities on future suicidal-ideation risk. Using human assessments within a stratified propensity-score framework to form comparable cohorts, they estimated treatment effects of support types and found that esteem and network support significantly reduce subsequent risk, with implications for tools that enhance support provision.</p>
<p>Zhang et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a causal framework based on a counterfactual neural temporal point process (TPP) to estimate the individual treatment effect (ITE) of misinformation on user beliefs and actions at scale, using a neural TPP with Gaussian mixtures for efficient inference. Experiments on synthetic data and a real COVID-19 vaccine dataset showed identifiable causal effects of misinformation, including negative shifts in users&#x2019; vaccine-related sentiments. Cheng et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] surveyed Event Causality Identification (ECI) and proposed a systematic taxonomy split into sentence-level ECI (SECI) and document-level ECI (DECI) tasks, reviewing approaches from feature/ML methods to deep semantic encoding, event-graph reasoning, and prompt/causal-knowledge pretraining, with notes on multilingual, cross-lingual, and zero-shot large language model settings.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Neural Encoders: Graph &#x002B; Attention</title>
<p>GCN encoders capture relational structure for text via message passing on graphs, benefiting settings with explicit concept relations [<xref ref-type="bibr" rid="ref-24">24</xref>]. Hamilton et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed GraphSAGE, an inductive framework that extended GCNs to unsupervised learning and introduced trainable aggregation functions beyond simple convolutions. The method generated embeddings for unseen nodes by sampling and aggregating neighborhood features, leveraging node attributes for generalization. Yao et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] introduced Text GCN, which built a corpus-level graph from word co-occurrence and document-word relations to jointly learn word and document embeddings. Without relying on external embeddings, Text GCN outperformed state-of-the-art methods on multiple benchmarks and showed strong robustness with limited training data.</p>
<p>Transformers, driven by self-attention, excel at weighting inputs and can be used as interpretable pooling over concept embeddings [<xref ref-type="bibr" rid="ref-27">27</xref>]. Vaswani et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed the Transformer, a sequence transduction architecture based solely on attention mechanisms, removing recurrence and convolutions. The model achieved state-of-the-art results on Workshop on Machine Translation 2014 English to German and English to French translation tasks, while being more parallelizable and significantly faster to train than prior approaches.</p>
<p>Devlin et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] introduced Bidirectional Encoder Representations from Transformers (BERT), a bidirectional Transformer-based model pre-trained on unlabeled text by jointly conditioning on left and right context. With simple fine-tuning, BERT achieved state-of-the-art results on eleven NLP tasks. Yang et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] proposed a hierarchical attention network (HAN) for document classification, which reflected the hierarchical structure of documents and applied attention at both word and sentence levels. The model outperformed prior methods on six large-scale benchmarks and provided interpretable document representations by highlighting informative words and sentences.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Continual Learning for Mental Health</title>
<p>Continual learning addresses time-varying, patient-specific data in mental-health scenarios by incrementally updating models from electronic health records, speech/text, and wearable streams while preserving prior knowledge. Although the CL for mental health literature remains limited, we highlight a few representative systems that show feasibility under practical constraints. Gamel and Talaat [<xref ref-type="bibr" rid="ref-31">31</xref>] proposed SleepSmart, an Internet of Things (IoT) enabled continual learning framework for intelligent sleep enhancement. The system employed wearable biosensors to capture physiological signals during sleep, which were processed via an IoT platform to deliver personalized recommendations. By leveraging continual learning, SleepSmart improved recommendation accuracy over time, and a pilot study demonstrated its effectiveness in enhancing sleep quality and reducing disturbances.</p>
<p>Lee and Lee [<xref ref-type="bibr" rid="ref-32">32</xref>] explored the role of continual learning in medicine, where models adapt to new patient data without forgetting prior knowledge. They emphasized challenges such as catastrophic forgetting and regulatory constraints, but argued that continual learning offers advantages over non-adaptive Food and Drug Administration approved systems by incrementally improving diagnostic and decision-support performance. Li and Jha [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed DOCTOR, a continual-learning framework for multi-disease detection on wearable medical sensors at the edge. The system used a multi-headed deep neural network with replay-based CL, via exemplar data preservation or synthetic data generation to mitigate catastrophic forgetting while sequentially adding tasks with new classes and distributions. In experiments, a single model maintained high accuracy, yielding up to 43% higher test accuracy, 25% higher F1 score, and 0.41 higher backward transfer over naive fine-tuning.</p>
<p>A structured comparison is presented in <xref ref-type="table" rid="table-1">Table 1</xref> to more clearly contextualize CNSGL within existing work. Prior methods typically incorporate only one or two of the components, symbolic representations, causal edge modeling, graph-based encoders, Transformer attention mechanisms, or continual-learning strategies, rather than unifying all of them within a single framework. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, approaches that employ symbolic reasoning rarely integrate GNN encoders or explicit causal edge construction; causal GNN models generally do not use symbolic tag vocabularies or Transformer-based pooling; and continual-learning systems commonly operate without symbolic graphs or causal modeling. In contrast, CNSGL combines directional PMI-derived causal edges, a symbolic tag graph, a two-layer GCN encoder, a Transformer attention pooler, and a multi-head freeze continual-learning strategy within one architecture tailored for mental-health risk detection. This integrated design forms the central novelty of the approach and demonstrates how CNSGL extends beyond existing component-wise methods.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison of CNSGL with representative neuro-symbolic, causal-modeling, and continual-learning architectures</title>
</caption>
 
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Author</th>
<th>Model</th>
<th>Symbolic</th>
<th>Causal Edge</th>
<th>Graph Encoder</th>
<th>Transformer</th>
<th>Continual</th>
<th>Application</th>
</tr>
<tr>
<th></th>
<th></th>
<th>Representation</th>
<th>Modeling</th>
<th>(GCN/GNN)</th>
<th>Attention</th>
<th>Learning</th>
<th>Domain</th>
</tr>
</thead>
<tbody>
<tr>
<td>Nie et al. (2022) [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>Incremental GCN</td>
<td>No (utterances &#x0026; speakers as nodes; no symbolic tags/lexicons)</td>
<td>No</td>
<td>GC/N</td>
<td>Yes, multi-head attention for utterance correlation</td>
<td>Yes, Incremental fine-tuning with new utterances</td>
<td>Conversation emotion detection</td>
</tr>
<tr>
<td>Kaur et al. (2022) [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Transformer-based causal categorization</td>
<td>No</td>
<td>Causal labels, but no causal edges</td>
<td>No</td>
<td>Yes</td>
<td>No</td>
<td>Mental-health causal categorization on social-media posts</td>
</tr>
<tr>
<td>Kodati &#x0026; Tene (2023) [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>Context-based bidirectional gated recurrent unit with multi-head attention and a convolutional neural network</td>
<td>Partial, POS tags &#x002B; lexicon features (not symbolic graphs)</td>
<td>No</td>
<td>No, CNN used</td>
<td>Yes, Multi-Head Attention &#x002B; BERT MLM/Self-attention</td>
<td>No</td>
<td>Suicidal-emotion detection on social-media text</td>
</tr>
<tr>
<td>Kumar (2023) [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>Neuro-Symbolic AI framework</td>
<td>Yes, structured knowledge graphs, symbolic reasoning, cognitive theories</td>
<td>General causal reasoning mentioned, but no graph construction method</td>
<td>No</td>
<td>No</td>
<td>No</td>
<td>personalized mental health therapy, computational psychiatry</td>
</tr>
<tr>
<td>Tang et al. (2023) [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>Causality-Driven GCN Framework</td>
<td>No</td>
<td>Yes, causal interventions &#x002B; invariant prediction principle &#x002B; causality scoring</td>
<td>GCN</td>
<td>No</td>
<td>No</td>
<td>Automated classification of postural abnormalities in Parkinson&#x2019;s disease</td>
</tr>
<tr>
<td>Bhuyan et al. (2024) [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>Conceptual Neuro- Symbolic AI framework</td>
<td>Yes, symbolic reasoning &#x0026; discrete logic</td>
<td>No</td>
<td>GNN</td>
<td>No</td>
<td>Yes</td>
<td>General AI/Neuro-Symbolic reasoning</td>
</tr>
<tr>
<td>Dalkic (2025) [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>Context-Aware EEG Emotion Recognition System</td>
<td>No (raw EEG &#x002B; context embeddings; no symbolic tags or lexicons)</td>
<td>No</td>
<td>No</td>
<td>Yes, Temporal Transformer encoder</td>
<td>Yes, EWC-based continual learning</td>
<td>EEG-based emotion recognition/affective computing</td>
</tr>
<tr>
<td>Patan&#x00E8; et al. (2025) [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>Prompt-based continual learning framework</td>
<td>No (mobile sensing features; no symbolic tags/lexicons)</td>
<td>No</td>
<td>No</td>
<td>Yes, Transformer backbone with task prompts</td>
<td>Yes, Replay buffer &#x002B; prompt-based adaptation</td>
<td>personalized mental well-being monitoring</td>
</tr>
<tr>
<td>Febrinanto et al. (2025) [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>Causal Graphs for Brains</td>
<td>No</td>
<td>Yes, causal discovery &#x002B; transfer entropy &#x002B; curvature-based rewiring</td>
<td>Yes. GNN models refined causal graphs</td>
<td>No</td>
<td>No</td>
<td>Brain disease classification (neuroscience)</td>
</tr>
<tr>
<td>Gosala et al. (2025) [<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
<td>GCN-LSTM; 12-layer GCN</td>
<td>No (EEG electrodes as nodes, not symbolic tags)</td>
<td>No, edges from cohesion/ phase-locking, not causal)</td>
<td>Yes, GCN &#x002B; hybrid GCN-LSTM</td>
<td>No</td>
<td>No</td>
<td>Schizophrenia classification from EEG (clinical neuroimaging)</td>
</tr>
<tr>
<td>Our Work</td>
<td>Continual Neuro- Symbolic Graph Learning (CNSGL)</td>
<td>Yes, symbolic mental-health tags with clinical relevance</td>
<td>Yes, directional PMI edges encoding causal tendencies between tags</td>
<td>Yes, 2-layer GCN encoder</td>
<td>Yes, Transformer-based attention pooling with CLS-to-tag weights for explainability</td>
<td>Yes, Multi-Head Freeze (encoder frozen after Task-1) to prevent forgetting</td>
<td>Mental-health risk detection from social media posts</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Work</title>
<p>A <italic>Continual Neuro-Symbolic Graph Learning (CNSGL) framework</italic> is proposed for causal inference and continual learning in mental-health risk assessment. In this framework, symbolic reasoning, causal graph construction, graph neural encoding, and continual learning are combined within a single architecture. The details of the proposed work are presented in the following subsections.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Symbolic Graph Construction</title>
<p>Mental-health text from online platforms or clinical records is largely unstructured and often contains implicit cues about psychological conditions that are difficult to analyze directly. To impose structure and enhance interpretability, each post is represented as a <italic>symbolic graph</italic> that captures both semantic content and clinically meaningful indicators. A compact set of ten symbolic tags, sleep, anxiety, depression, stress, anger, lonely, health, fear, coping, and suicidal, was constructed based on well-established constructs in computational mental-health research and their frequent annotation in benchmark datasets. The vocabulary was further validated through manual inspection and an expert-informed review to ensure clinical relevance. A small, consistent set was intentionally maintained to minimize noise and support stable directional PMI estimation during causal graph construction. Let <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow></mml:math></inline-formula> denote the dataset of posts, and let <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow></mml:math></inline-formula> denote a post in this collection. A vocabulary <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow></mml:math></inline-formula> of risk-related terms, including <italic>hopelessness</italic>, <italic>insomnia</italic>, and <italic>self-harm</italic>, is predefined. Using this lexicon, a set of symbolic tags is assigned to the post,
<disp-formula id="ueqn-1"><mml:math id="mml-ueqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where each <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> denotes a tag identified in <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. These tags provide explicit signals of potential risk factors, linking the unstructured narrative of a post to interpretable constructs grounded in psychology.</p>
<p>The symbolic graph is defined as <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> consisting of:
<list list-type="bullet">
<list-item>
<p><bold>Nodes:</bold> <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>&#x222A;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>v</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> represents the post and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> denote the symbolic tags.</p></list-item>
<list-item>
<p><bold>Edges:</bold> <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>E</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> establish bidirectional links between the post and its tags.</p></list-item>
<list-item>
<p><bold>Features:</bold> Node attributes encode symbolic and semantic information. Tag node features use Term Frequency-Inverse Document Frequency (TF-IDF) scores [<xref ref-type="bibr" rid="ref-44">44</xref>] of the tag <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> in post <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. For a tag node <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>,
<disp-formula id="ueqn-24"><mml:math id="mml-ueqn-24" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>TFIDF</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
while the post node <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>v</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> is enriched with a 64-dimensional embedding obtained through truncated singular value decomposition (SVD) of the TF-IDF representation,</p></list-item>
</list>
<disp-formula id="ueqn-3"><mml:math id="mml-ueqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>SVD</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>TFIDF</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>At this stage, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> uses undirected post-tag links (implemented as bidirectional pairs) to encode association rather than causality. Directional tag&#x2013;tag edges are introduced in the subsequent Causal-Aware Graph Enrichment step, where order-sensitive statistics are is used to determine edge orientation.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Causal-Aware Graph Enrichment</title>
<p>Edges based solely on co-occurrence capture statistical associations but cannot distinguish whether one factor precedes or influences another. For example, <italic>sleep deprivation</italic> may frequently appear with <italic>stress</italic>, yet in many cases it precedes and contributes to <italic>suicidal ideation</italic>. To incorporate such directional relationships, graphs are enriched with causal edges in addition to co-occurrence links. To quantify how strongly tags co-occur in the input space, point-wise mutual information (PMI) is used here. Consider the tags <italic>sleep</italic> and <italic>anxiety</italic>. These tags may frequently appear together in posts, which would yield a symmetric, undirected edge in a standard co-occurrence graph. Directional PMI instead focuses on ordered pairs and estimates whether one tag is more likely to appear before the other. If ordered counts show that mentions of <italic>sleep</italic> problems systematically precede <italic>anxiety</italic> indicators more often than the reverse, then <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>PMI</mml:mi><mml:mrow><mml:mtext>dir</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mtext>sleep</mml:mtext><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mtext>anxiety</mml:mtext><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> exceeds <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>PMI</mml:mi><mml:mrow><mml:mtext>dir</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mtext>anxiety</mml:mtext><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mtext>sleep</mml:mtext><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and the graph includes the directed edge <italic>sleep</italic> <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> <italic>anxiety</italic>. In this way, directional PMI encodes asymmetric, precedence-aware relationships that cannot be represented by undirected co-occurrence edges alone. PMI between two given tags <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula> is calculated in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> [<xref ref-type="bibr" rid="ref-45">45</xref>]:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mtext>PMI</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mspace width="thinmathspace" /></mml:math></disp-formula></p>
<p>Probabilities were estimated by counting how often each tag and each tag pair appeared within the same post and then normalizing by the total. A small smoothing constant was applied so that rare tags didn&#x2019;t get zero probability. Because PMI treats a pair the same in either order, it captures association only and does not encode direction. Let <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x227A;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> be the number of posts in which <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> occurs before <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula>, and let <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mtext>pairs</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> be the total number of ordered tag pairs considered. <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> shows the mathematical definition of the directional PMI:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msup><mml:mrow><mml:mtext>PMI</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>dir</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mo>&#x227A;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mspace width="thinmathspace" /><mml:mo>,</mml:mo><mml:mspace width="2em" /></mml:math></disp-formula>where, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x227A;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2248;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mfrac><mml:mrow><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mo>&#x227A;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mtext>pairs</mml:mtext></mml:mrow></mml:msub></mml:mfrac></mml:math></inline-formula> is the probability of ordered event. A directional PMI threshold <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is applied to filter out weak or noisy associations. In practice, <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> was selected through validation by sweeping values in <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0.0</mml:mn><mml:mo>,</mml:mo><mml:mn>0.05</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> and choosing the smallest threshold that removed spurious edges while preserving clinically meaningful relations. The framework was most stable for <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.1</mml:mn></mml:math></inline-formula>, which we adopt for all experiments. Sensitivity analysis showed that the graph structure remained consistent for thresholds within a small neighborhood (<inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mn>0.05</mml:mn></mml:math></inline-formula>&#x2013;<inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mn>0.15</mml:mn></mml:math></inline-formula>), indicating that results are not overly sensitive to the exact choice of <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula>.</p>
<p>A directed edge, <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula> is introduced whenever <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msup><mml:mrow><mml:mi mathvariant="normal">P</mml:mi><mml:mi mathvariant="normal">M</mml:mi><mml:mi mathvariant="normal">I</mml:mi></mml:mrow><mml:mrow><mml:mtext>dir</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula>. This indicates that the occurrence of <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> increases the likelihood of <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>t</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula>, suggesting a causal tendency rather than a mere correlation. After adding the causality factor, the resulting enriched graph becomes:
<disp-formula id="ueqn-6"><mml:math id="mml-ueqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>co</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x222A;</mml:mo><mml:msubsup><mml:mi>E</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>causal</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msubsup><mml:mi>E</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mtext>co</mml:mtext></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes post&#x2013;tag associations, and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msubsup><mml:mi>E</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mtext>causal</mml:mtext></mml:mrow></mml:msubsup></mml:math></inline-formula> represents directed tag&#x2013;tag relations. For example: If mentions of <italic>sleep problems</italic> are frequently followed by references to <italic>anxiety</italic>, and the directional PMI exceeds the threshold <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula>, an edge <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mtext>sleep</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mtext>anxiety</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is added. In this way, the representation allows the model to reason not only about what terms appear together, but also about which factors may act as potential precursors of others. During the execution of the GCN, the directed causal edges are converted into an undirected form (with self-loops) so that message passing remains symmetric. Throughout this work, the directed edges are treated as precedence-aware statistical associations rather than definitive cause&#x2013;effect links; accordingly, the term &#x201C;causal&#x201D; is used in an operational sense to describe directional risk-tendency patterns observed in mental-health text.</p>
<p>The proposed CNSGL framework is depicted as an end-to-end architecture in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, emphasizing the left-to-right progression from symbolic/causal structuring to representation learning and, then to sequential adaptation. The diagram marks where the shared encoder is frozen and where dataset-specific heads are attached, making clear how prior knowledge is preserved while new tasks are added. The Transformer Encoder used for graph-level pooling in the encoding stage is described separately in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed CNSGL framework consisting of three key components: (<bold>1</bold>) Preprocessing, where a post-tag graph is built and enriched with directional causal edges (red), and causal tag nodes (green). (<bold>2</bold>) Encoding, consisting of two-layer GCN followed by a light-weight Transformer <monospace>[CLS]</monospace> pooler, produces a post vector and tag importances.(<bold>3</bold>) Continual learning (MH-Freeze): the encoder and <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>Head</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> were trained on <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>T</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, after which the encoder was frozen; for subsequent datasets, a small linear-sigmoid head was attached and trained, with a dataset specific threshold calibrated. At inference, a post was encoded once, the appropriate head was selected by dataset, and the resulting probability was thresholded to yield <italic>risky</italic>/<italic>non-risky</italic> posts</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_75119-fig-1.tif"/>
</fig><fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Architecture of the encoder-only Transformer used as an attention pooler. The CLS token and tag embeddings are fed into a single Transformer block, where multi-head attention computes CLS <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> tag attention scores while masking tag-to-tag interactions. Residual connections, layer normalization, and a feed-forward network refine the CLS representation, which becomes the final pooled embedding for classification</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_75119-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>GCN Encoder</title>
<p>The symbolic graphs enriched with causal relations are processed by a two-layer GCN. The GCN propagates information across connected nodes so that each representation reflects both its own features and those of neighboring nodes. In this way, a post node aggregates signals from its tags, while tag nodes incorporate both symbolic and causal context. <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> shows how the node representations are updated at layer <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>l</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-46">46</xref>].
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>&#x03C3;</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>D</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="false" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mspace width="thinmathspace" /><mml:msup><mml:mrow><mml:mover><mml:mi>D</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="false" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:msup><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mrow><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mo>+</mml:mo><mml:mi>I</mml:mi></mml:math></inline-formula> is the adjacency with self-loops, <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mrow><mml:mover><mml:mi>D</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is its degree matrix, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> are trainable weights, and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is an element-wise ReLU activation function. The input is <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, the node-feature matrix of graph <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. Since causal edges are directed, their weights are first assembled in a directed matrix <italic>W</italic> and then symmetrized to form <italic>A</italic> (e.g., <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="false" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo stretchy="false">(</mml:mo><mml:mi>W</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>) prior to normalization.</p>
<p>Although the final GCN uses a symmetrized adjacency matrix for stable message passing, the causal interpretability of the framework is retained because symmetrization occurs only after causal tendencies have been encoded in the edge-selection stage. Directional PMI determines which tag pairs are connected and the strength of those connections, thereby shaping the underlying causal structure even if the GCN operates on an undirected form of the graph. The interpretability comes from this directed edge construction and from the subsequent analysis of causal paths and attention weights, whereas symmetrization serves primarily as a computational requirement of the canonical GCN rather than a removal of causal information. After two layers, the node embeddings <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> are passed to a learned classification token [CLS] as a global query over the tag nodes to produce a graph-level representation <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> with tag-level importances. This pooled vector <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> encodes symbolic, semantic, and causal structure and is used for classification.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Transformer-Based Attention</title>
<p>While mean pooling provides a simple mechanism for aggregating node embeddings into a graph-level representation, it treats all nodes equally and fails to highlight which risk factors are more influential in a particular post. To address this limitation, the node embeddings produced by the GCN are passed through a Transformer encoder to perform attention-based pooling [<xref ref-type="bibr" rid="ref-47">47</xref>]. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates the Transformer encoder block employed as an attention pooler over GCN-derived node embeddings. The [CLS] token attends to tag embeddings to produce a pooled representation while simultaneously providing interpretable tag-level importance scores through attention weights.</p>
<p>Let <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> be the node embeddings (post &#x002B; tags) for post <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>i</mml:mi></mml:math></inline-formula> after the 2-layer GCN. The input sequence formed is presented in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mo>;</mml:mo><mml:mspace width="thinmathspace" /><mml:msup><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula>where <monospace>[CLS]</monospace> is a learnable pooling token. The encoder computes linear projections
<disp-formula id="ueqn-9"><mml:math id="mml-ueqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mi>Q</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mi>K</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mi>V</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>with trainable <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>W</mml:mi><mml:mi>Q</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>K</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>V</mml:mi></mml:msub></mml:math></inline-formula> and per-head key dimension <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>. We apply a <italic>pooler mask</italic> so that only the <monospace>[CLS]</monospace> row of <italic>Q</italic> issues queries (tag<inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mo stretchy="false">&#x2194;</mml:mo></mml:math></inline-formula>tag attention is masked). Let <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> denote the <monospace>[CLS]</monospace> query. The attention weights from <monospace>[CLS]</monospace> to all tokens are calculated in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>softmax</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msup><mml:mi>K</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:msqrt><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula></p>
<p>These are then used to form a weighted sum of the values for <monospace>[CLS]</monospace> (computed per head, concatenated, and projected). Each encoder block applies Add&#x0026;LayerNorm around multi-head attention and a Feed-Forward Network (FFN). The final <monospace>[CLS]</monospace> vector is taken as the pooled embedding
<disp-formula id="ueqn-11"><mml:math id="mml-ueqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The weights <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> (averaged over heads) serve as <italic>tag importances</italic>, providing a transparent summary of which tags influenced the decision. The pooled vector <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is passed to a linear sigmoid head (<inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>) to obtain the risk label, explained in the subsequent section in detail. The attention pooler improves aggregation of symbolic and causal information while preserving interpretability via tag-level weights.</p>
<p>The pooling module is implemented as a single Transformer-style encoder block with multi-head self-attention and a position-wise feed-forward network (FFN). In practice, the CLS attention Pooler uses an embedding dimension of <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula>, four attention heads, and an FFN hidden size of 256 with dropout 0.1. A learned <monospace>[CLS]</monospace> token is prepended to the node embeddings, and its final output vector is used as the graph-level representation, while the CLS-to-tag attention weights provide interpretable tag importances. The same configuration is used across datasets to maintain consistency in the continual-learning setup.</p>
<p>To illustrate how the Transformer attention pooler provides qualitative interpretability, two example posts are shown below. In each case, the model highlights the most influential tags and their causal relations when producing a risk prediction.</p>
<p><monospace>Example 1 (High-risk post):</monospace> <italic>&#x201C;I have not slept properly for days, and the constant anxiety is making everything feel overwhelming. Lately I keep thinking that things would be easier if I just disappeared.&#x201D;</italic> The attention pooler assigns high importance to the tags <italic>sleep</italic>, <italic>anxiety</italic>, and <italic>suicidal</italic>, with a strong causal pathway <italic>sleep</italic><inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> <italic>anxiety</italic><inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula><italic>suicidal</italic>. These attended tags correspond to clinically salient risk indicators, leading the classifier to assign a high-risk label.</p>
<p><monospace>Example 2 (Low-risk post):</monospace> <italic>&#x201C;Feeling a bit stressed about exams next week, but talking to friends has helped and I&#x2019;m trying to stay positive.&#x201D;</italic> The model focuses primarily on <italic>stress</italic>, with low attention weights on other tags and no causal escalation toward <italic>depression</italic> or <italic>suicidal</italic>. The attention pattern reflects a non-escalatory emotional state, leading to a low-risk prediction.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Continual Learning Strategy</title>
<p>In real-world applications, data arrive in stages <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mi>K</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> with evolving language, populations, and even label definitions. Sequential training on <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> risks <italic>catastrophic forgetting</italic> of knowledge learned on earlier datasets. Continual learning is incorporated to address the sequential arrival of mental-health datasets and the risk of catastrophic forgetting. This ensured that knowledge acquired from earlier domains was preserved while adaptation to new sources was achieved, thereby enhancing robustness and practical applicability of the framework. A simple, effective continual learning technique, MH-Freeze is used, which preserves a shared encoder while adding a small task-specific head per dataset. Let <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> be a post graph and let <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>f</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub></mml:math></inline-formula> denote the encoder mapping <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> to a pooled embedding. The pooled embedding is computed in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>f</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>f</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub></mml:math></inline-formula> comprises causal-aware graph construction, two GCN layers, and the lightweight Transformer encoder. Thus <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> integrates symbolic structure, causal links, and attention-based tag weighting.</p>
<p>For dataset <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>G</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> with a <italic>single</italic> binary risk label <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, the task head <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> is a linear&#x2013;sigmoid classifier unit that produces a probability
<disp-formula id="ueqn-13"><mml:math id="mml-ueqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>&#x03C3;</mml:mi><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mtext>&#x00A0;</mml:mtext><mml:mo>&#x2208;</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>with parameters <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>b</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> and sigmoid <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. A threshold <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> (calibrated on a small validation split) yields the binary decision. At inference, predictions are labeled as positive if <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>p</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>.</p>
<p><italic>MH-Freeze</italic></p>
<p>At first, the shared encoder <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the initial linear&#x2013;sigmoid head <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> are jointly trained on <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> by minimizing the average BCE loss, which is the negative log-likelihood of a Bernoulli target and directly trains calibrated probabilities <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> from the sigmoid head. <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref> presents formulation of BCE for a single label <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> and predicted probability <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is [<xref ref-type="bibr" rid="ref-48">48</xref>]:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>BCE</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">[</mml:mo></mml:mrow></mml:mstyle><mml:mspace width="thinmathspace" /><mml:mi>y</mml:mi><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">]</mml:mo></mml:mrow></mml:mstyle><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The shared encoder <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the initial linear&#x2013;sigmoid head <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> are then jointly trained on <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> by minimizing the average BCE, as shown in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>:
<disp-formula id="ueqn-15"><mml:math id="mml-ueqn-15" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:munder><mml:mtext>&#x00A0;</mml:mtext><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mtext>BCE</mml:mtext></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mi>w</mml:mi><mml:mn>1</mml:mn><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:msub><mml:mi>f</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>After convergence, the encoder parameters <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> are frozen, so that <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> provides a fixed representation for all subsequent heads. Freezing is applied after Task 1 because the first dataset provides the broadest and most diverse distribution of symbolic tags, allowing the encoder to learn generalizable representations before domain-specific heads are introduced. We also examined variants where the encoder is frozen after Task 2 or Task 3. These alternatives showed higher forgetting on earlier datasets and reduced overall stability, as the encoder continued adapting toward the later-task distributions and drifted away from the symbolic-causal structure learned initially. Freezing after Task 1 therefore offered the best balance between preserving prior knowledge and supporting effective multi-head adaptation.</p>
<p>For each subsequent dataset <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> (<inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi>k</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2265;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>2</mml:mn></mml:math></inline-formula>), a lightweight head is instantiated and only its parameters are optimized while keeping the encoder fixed, as mentioned in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>:
<disp-formula id="ueqn-17"><mml:math id="mml-ueqn-17" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mtext>&#x00A0;</mml:mtext><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mtext>BCE</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:msub><mml:mi>f</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0.</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>As <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> is fixed, adaptation reduces to fitting task-specific <italic>linear separators</italic> in the common embedding space <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>z</mml:mi></mml:math></inline-formula>, which avoids cross-task interference and sharply limits forgetting. Intuitively, the shared encoder captures domain-general structure (symbolic and causal relations plus attention-based tag weighting), while each head accounts for dataset-specific prevalence, wording, or scope. Given dataset, the corresponding head <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> is selected and its probability <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> is thresholded to yield the label:
<disp-formula id="ueqn-19"><mml:math id="mml-ueqn-19" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>risky if</mml:mtext></mml:mrow><mml:msup><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2265;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mrow><mml:mtext>non-risky otherwise.</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results</title>
<p>The experiments were conducted on a high-performance workstation equipped with an AMD Ryzen Threadripper 2950X (16 cores, 3.50 GHz) and 32 GB RAM. The experiments were implemented in Python 3.11 using key libraries such as PyTorch 2.2, PyTorch Geometric 2.5, NumPy, Scikit&#x2013;learn, and Matplotlib. All codes were executed in a Jupyter Notebook environment configured on Windows 11, ensuring a consistent and reproducible experimental setup.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset Description</title>
<p>To evaluate the effectiveness and generalizability of the proposed Continual Neuro-Symbolic Graph Learning framework, six diverse datasets were used spanning Reddit-based mental health discourse and clinician-guided interviews. A brief summary is provided in <xref ref-type="table" rid="table-2">Table 2</xref>. These corpora differ in annotation protocols, linguistic style, and risk indicators, enabling a comprehensive assessment of both the symbolic reasoning components and the graph-based learning modules.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Summary of datasets used as continual-learning tasks, showing source, text type, and size</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Source</th>
<th>Type</th>
<th>Size</th>
</tr>
</thead>
<tbody>
<tr>
<td>DASH-2020</td>
<td>Zenodo</td>
<td>Reddit posts</td>
<td>3151</td>
</tr>
<tr>
<td>Dereaddit</td>
<td>Kaggle</td>
<td>Reddit posts</td>
<td>3553</td>
</tr>
<tr>
<td>Kaggle MH</td>
<td>Kaggle</td>
<td>Reddit posts</td>
<td>5957</td>
</tr>
<tr>
<td>SWMH</td>
<td>Zenodo</td>
<td>Reddit posts (split)</td>
<td>54,412</td>
</tr>
<tr>
<td>Go_emotions</td>
<td>Kaggle</td>
<td>Reddit posts</td>
<td>58,011</td>
</tr>
<tr>
<td>E-DAIC</td>
<td>USC-ICT</td>
<td>Clinical dialogues</td>
<td>418 transcripts</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><list list-type="bullet">
<list-item>
<p>Data Analytics for Smart Health (DASH-2020) [<xref ref-type="bibr" rid="ref-49">49</xref>]: It consists of reddit posts annotated for substance use, addiction, and recovery. For our binary setup, we merged all recovery-related categories into a single non-addicted class, while posts explicitly labeled as addicted are retained as the positive class.</p></list-item>
<list-item>
<p>Go_emotions [<xref ref-type="bibr" rid="ref-50">50</xref>]: It is a reddit-based dataset annotated with 27 fine-grained emotion categories plus neutral. It contains about 58,000 unique comments collected from diverse subreddits. For binary mental-health risk classification in our work, all emotion categories associated with distress (e.g., sadness, anger, fear, anxiety) were grouped as risky, while the rest were treated as non-risky.</p></list-item>
<list-item>
<p>Kaggle Mental Health [<xref ref-type="bibr" rid="ref-51">51</xref>]: This dataset sourced from Kaggle repository, contains Reddit posts labeled across five mental health conditions. For binary classification, all risk-associated categories were merged into a single risky class, while the remaining category was treated as non-risky.</p></list-item>
<list-item>
<p>Dreaddit [<xref ref-type="bibr" rid="ref-52">52</xref>]: This dataset also sourced from Kaggle, consists of reddit corpus for stress detection across five community categories. The authors collected around 190 K posts and crowd-sourced stress labels for around 3.5 K text segments. The public release provides official splits (<inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mo>&#x2248;</mml:mo></mml:math></inline-formula>2838 train/715 test) with roughly balanced stress vs. non-stress. In our setup, we used the provided binary label (1 &#x003D; stressful/risky), and the official train/test.</p></list-item>
<list-item>
<p>Reddit SuicideWatch and Mental Health Collection (SWMH) [<xref ref-type="bibr" rid="ref-53">53</xref>]: This is a Reddit-derived dataset released via Zenodo, combining posts from the SuicideWatch subreddit and other mental health communities. Posts from SuicideWatch are categorized as the risky class, while those from broader mental health forums are assigned to the non-risky class.</p></list-item>
<list-item>
<p>Extended DAIC (E-DAIC) [<xref ref-type="bibr" rid="ref-54">54</xref>]: E-DAIC is an extended version of the original Distress Analysis Interview Corpus with Wizard-of-Oz (DAIC-WOZ) corpus [<xref ref-type="bibr" rid="ref-55">55</xref>]. The dataset, sourced from University of Southern California-Institute for Creative Technologies (USC-ICT), includes semi-structured interviews conducted by a virtual agent named Ellie, controlled either by a human wizard or an autonomous AI system. It contains transcribed clinical interviews annotated using PHQ-8 scores.</p></list-item>
</list></p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Continual Learning Baselines</title>
<p>The proposed MH-Freeze framework is compared with several representative continual learning techniques, each reflecting a distinct strategy for mitigating catastrophic forgetting in sequential task scenarios.
<list list-type="bullet">
<list-item>
<p>Elastic Weight Consolidation (EWC) [<xref ref-type="bibr" rid="ref-56">56</xref>]: EWC addresses catastrophic forgetting in sequential learning by estimating the importance of each parameter for previously learned tasks (via a Fisher-based approximation) and selectively slowing changes to those important weights when learning a new task. This preserves prior expertise while allowing plasticity on less critical parameters.</p></list-item>
<list-item>
<p>Gradient Episodic Memory (GEM) [<xref ref-type="bibr" rid="ref-57">57</xref>]: GEM uses an episodic memory of past tasks and projects the current gradient to satisfy inequality constraints that do not increase loss on stored past-task examples. This enforces update compatibility with earlier tasks and can yield positive backward transfer when gradients align. In our experiments, we adopt the efficient A-GEM variant with the same memory protocol as ER and apply projection at every step before the optimizer update.</p></list-item>
<list-item>
<p>Learning without Forgetting (LWF) [<xref ref-type="bibr" rid="ref-58">58</xref>]: LWF adapts a network to new tasks using only new-task data while preserving prior capabilities via knowledge distillation: the current model is trained to match the frozen previous model&#x2019;s outputs on the new data, alongside the new-task loss. This avoids storing old datasets, competes with multitask training that has access to old data, and often outperforms plain feature extraction or finetuning when old and new tasks are similar.</p></list-item>
<list-item>
<p>Experience Replay (ER) [<xref ref-type="bibr" rid="ref-59">59</xref>]: It mitigates forgetting by maintaining a small episodic memory of past-task examples and interleaving them with current-task batches during training. This simple rehearsal stabilizes prior decision boundaries while preserving plasticity on new data, yielding a strong, low-complexity baseline. We have kept a fixed-size, class-balanced buffer. Each minibatch mixes current-task samples with buffer samples at a fixed ratio. Buffer size and ratio are tuned on validation.</p></list-item>
<list-item>
<p>Finetuning: The finetuning (Sequential Learning) across tasks without any anti-forgetting mechanism serves as a lower-bound baseline [<xref ref-type="bibr" rid="ref-60">60</xref>]. In this work, a single shared head and encoder are updated sequentially across tasks under the same optimizer/schedule and validation-based thresholding; no replay or regularization terms are added.</p></list-item>
<list-item>
<p>Synaptic Intelligence (SI) [<xref ref-type="bibr" rid="ref-61">61</xref>]: This CL technique assigns an &#x2018;importance&#x2019; score to each weight based on how much it contributed during training on a task. At the end of a task, those importance scores are retained as a summary of what mattered most. When the next task arrives, SI adds a lightweight penalty that discourages large changes to previously important weights while leaving the others free to adapt. We applied SI to the encoder (and pooler), snapshot parameters at each task boundary, and tune the overall regularization strength and a small stabilizer constant on the validation split.</p></list-item>
</list></p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation Metrics</title>
<p>The effectiveness of the proposed framework and the baseline continual learning techniques is evaluated using multiple performance metrics. These metrics capture not only classification accuracy but also robustness under class imbalance and calibration of probabilistic outputs.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Accuracy</title>
<p>Accuracy measures the proportion of correctly classified instances, as computed in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>. It reflects how effectively each continual-learning method distinguishes risky posts from non-risky ones across sequential mental-health datasets:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mtext>Accuracy</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <italic>TP</italic>, <italic>TN</italic>, <italic>FP</italic>, and <italic>FN</italic> denote true positive, true negative, false positive, and false negative instances, respectively.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>F1-Score</title>
<p>It is the harmonic mean of precision and recall, rewarding models that balance both low false positives and low false negatives, as shown in <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>. Precision is the proportion of predicted positives that are correct, recall is the proportion of actual positives that are correctly identified. F1 score reflects how well each continual-learning method maintains balanced risky vs. non-risky decisions across sequential datasets.:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mtext>F1-Score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Area under ROC Curve (AUROC)</title>
<p>The AUROC evaluates the trade-off between true positive rate (TPR) and false positive rate (FPR) across varying thresholds. It is defined as the probability that a randomly chosen positive is ranked higher than a randomly chosen negative.</p>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Area under Precision-Recall Curve (AUPRC)</title>
<p>The AUPRC integrates the precision- recall curve, which is more informative under class imbalance. It summarizes the trade-off between precision and recall across thresholds.</p>
</sec>
<sec id="s4_3_5">
<label>4.3.5</label>
<title>Brier Score</title>
<p>The Brier score evaluates the accuracy of probabilistic predictions by measuring the mean squared error between predicted probabilities <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and true labels <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. It can be computed using <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mtext>Brier</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s4_3_6">
<label>4.3.6</label>
<title>Expected Calibration Error (ECE)</title>
<p>ECE measures the alignment between predicted probabilities and observed accuracy. Predictions are partitioned into <italic>M</italic> bins according to confidence, and the weighted average gap between accuracy and confidence is reported. It can be computed using <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>. In our work, ECE is computed per task (dataset) to assess whether MH-Freeze and the baselines produce well-calibrated risk probabilities after threshold calibration in the continual-learning sequence.
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mtext>ECE</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mi>N</mml:mi></mml:mfrac><mml:mspace width="thinmathspace" /><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">|</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mtext>acc</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>conf</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">|</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula>where <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:math></inline-formula> is the set of samples in bin <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mi>m</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mtext>acc</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> the empirical accuracy, and <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mtext>conf</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> the mean confidence.</p>
</sec>
<sec id="s4_3_7">
<label>4.3.7</label>
<title>Matthews Correlation Coefficient (MCC)</title>
<p>MCC quantifies how well the classifier balances positive/negative decisions across sequential tasks and shifting, imbalanced class distributions, penalizing asymmetric error patterns that F1 or accuracy may hide. A higher value of MCC indicates better predictive performance. The MCC is computed in <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref> [<xref ref-type="bibr" rid="ref-62">62</xref>].
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mtext>MCC</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:msqrt><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:msqrt></mml:mfrac></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Results and Discussions</title>
<p>The proposed MH-Freeze framework demonstrates strong continual-learning behavior across heterogeneous datasets as seen in <xref ref-type="table" rid="table-3">Table 3</xref>. MH-Freeze performs continual learning by freezing a shared GCN-Transformer encoder and training lightweight, task-specific heads as new tasks arrive. MH-Freeze is compared against six continual-learning baselines across six tasks, where each task corresponds to a different dataset in a fixed sequential order. The six tasks correspond to distinct datasets: T1 &#x003D; DASH, T2 &#x003D; Dreaddit, T3 &#x003D; SWMH, T4 &#x003D; Go_emotions, T5 &#x003D; DAIC-WOZ, and T6 &#x003D; Kaggle-MH. For brevity and consistency, these datasets are referenced as T1-T6 throughout the remainder of the paper. In all experiments, tasks are encountered in the fixed order T1-T6; balanced mini-batches are used to counter dataset-level imbalance, and the MH-Freeze architecture prevents dominance by any single dataset due to differing label distributions. The discriminative capability remains uniformly high, with AUPRC <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mo>&#x2265;</mml:mo></mml:math></inline-formula> 0.934 and AUROC <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mo>&#x2265;</mml:mo></mml:math></inline-formula> 0.942 throughout the task sequence, indicating robust separability between risk and non-risk classes. Both Accuracy (0.898 to 0.925) and F1-Score (0.886 to 0.923) follow a steady upward trajectory from T1 (DASH) to T6 (Kaggle-MH), reflecting positive forward transfer without degradation of earlier competencies. The MCC also improves from 0.829 to 0.873, confirming balanced predictive behavior under label imbalance. In parallel, Brier and ECE scores decline from 0.069 to 0.060 and 0.023 to 0.014, respectively, demonstrating progressive improvement in probability calibration. These metrics affirm that MH-Freeze effectively preserves prior knowledge while adapting to new domains, achieving well-calibrated, generalizable predictions with minimal catastrophic forgetting. It maintains a clear advantage in both discrimination and calibration.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Continual learning performance of MH-Freeze across all datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Task</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>T1 (DASH)</bold></td>
<td>0.898</td>
<td>0.886</td>
<td>0.934</td>
<td>0.942</td>
<td>0.829</td>
<td>0.069</td>
<td>0.023</td>
</tr>
<tr>
<td><bold>T2 (Dreaddit)</bold></td>
<td>0.906</td>
<td>0.893</td>
<td>0.94</td>
<td>0.949</td>
<td>0.836</td>
<td>0.065</td>
<td>0.020</td>
</tr>
<tr>
<td><bold>T3 (SWMH)</bold></td>
<td>0.912</td>
<td>0.898</td>
<td>0.939</td>
<td>0.948</td>
<td>0.841</td>
<td>0.066</td>
<td>0.021</td>
</tr>
<tr>
<td><bold>T4 (Go_emotions)</bold></td>
<td>0.918</td>
<td>0.905</td>
<td>0.947</td>
<td>0.955</td>
<td>0.849</td>
<td>0.061</td>
<td>0.016</td>
</tr>
<tr>
<td><bold>T5 (DAIC-WOZ)</bold></td>
<td>0.923</td>
<td>0.911</td>
<td>0.945</td>
<td>0.954</td>
<td>0.852</td>
<td>0.062</td>
<td>0.018</td>
</tr>
<tr>
<td><bold>T6 (Kaggle-MH)</bold></td>
<td>0.925</td>
<td>0.923</td>
<td>0.947</td>
<td>0.965</td>
<td>0.873</td>
<td>0.060</td>
<td>0.014</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Among the CL baselines used for comparison, GEM performs better, attaining moderately high AUROC (0.94&#x2013;0.95) and balanced MCC values, though it gains plateau beyond mid-sequence tasks. LWF exhibits reasonable F1-scores but suffers from high calibration error and inconsistent reliability across datasets. Experience Replay and Synaptic Intelligence provide stable yet lower performance, with AUROC typically below 0.91 and limited robustness to domain shifts. EWC achieves comparable mid-range results but shows greater sensitivity to task transitions, while Finetuning performs worst overall, displaying rapid accuracy decay (0.73&#x2013;0.76) and high Brier/ECE values indicative of severe forgetting. In contrast, MH-Freeze sustains near-optimal metrics across all six tasks, confirming that its frozen encoder with task-specific heads yields superior retention, adaptation, and calibration in continual-learning environments. The detailed results for each continual-learning baseline are presented in <xref ref-type="table" rid="table-4">Tables 4</xref>&#x2013;<xref ref-type="table" rid="table-9">9</xref>, providing a comprehensive comparison across all tasks. Experience Replay and Synaptic Intelligence show early performance saturation, as seen in <xref ref-type="table" rid="table-6">Tables 6</xref> and <xref ref-type="table" rid="table-7">7</xref>, respectively. This likely reflects limited forward transfer and calibration instability under domain shift. A plausible cause is that ER&#x2019;s small replay buffer cannot adequately represent later datasets, while SI&#x2019;s weight-importance penalty restricts the flexibility needed to adapt. Consistently higher Brier and ECE on the final tasks (T4-T6) reinforce this interpretation, indicating less reliable probabilities and weaker calibration as the data distribution changes. While MH-Freeze exhibits a monotonic increase in MCC from T1 to T6, GEM stabilizes at a slightly lower range (<xref ref-type="table" rid="table-4">Table 4</xref>). Brier and ECE generally decline for MH-Freeze, indicating progressively better calibration, whereas replay and regularization-based baselines (ER, EWC, SI) show smaller or inconsistent reductions.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Continual learning performance of GEM across all datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Task</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>T1 (DASH)</bold></td>
<td>0.888</td>
<td>0.876</td>
<td>0.929</td>
<td>0.937</td>
<td>0.819</td>
<td>0.072</td>
<td>0.025</td>
</tr>
<tr>
<td><bold>T2 (Dreaddit)</bold></td>
<td>0.896</td>
<td>0.883</td>
<td>0.935</td>
<td>0.944</td>
<td>0.826</td>
<td>0.068</td>
<td>0.022</td>
</tr>
<tr>
<td><bold>T3 (SWMH)</bold></td>
<td>0.902</td>
<td>0.888</td>
<td>0.934</td>
<td>0.943</td>
<td>0.831</td>
<td>0.069</td>
<td>0.023</td>
</tr>
<tr>
<td><bold>T4 (Go_emotions)</bold></td>
<td>0.908</td>
<td>0.895</td>
<td>0.942</td>
<td>0.950</td>
<td>0.839</td>
<td>0.064</td>
<td>0.018</td>
</tr>
<tr>
<td><bold>T5 (DAIC-WOZ)</bold></td>
<td>0.913</td>
<td>0.902</td>
<td>0.940</td>
<td>0.949</td>
<td>0.842</td>
<td>0.065</td>
<td>0.02</td>
</tr>
<tr>
<td><bold>T6 (Kaggle-MH)</bold></td>
<td>0.913</td>
<td>0.903</td>
<td>0.942</td>
<td>0.955</td>
<td>0.863</td>
<td>0.063</td>
<td>0.016</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Continual learning performance of LWF across all datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Task</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>T1 (DASH)</bold></td>
<td>0.865</td>
<td>0.854</td>
<td>0.906</td>
<td>0.915</td>
<td>0.791</td>
<td>0.181</td>
<td>0.172</td>
</tr>
<tr>
<td><bold>T2 (Dreaddit)</bold></td>
<td>0.889</td>
<td>0.877</td>
<td>0.926</td>
<td>0.934</td>
<td>0.821</td>
<td>0.17</td>
<td>0.129</td>
</tr>
<tr>
<td><bold>T3 (SWMH)</bold></td>
<td>0.875</td>
<td>0.864</td>
<td>0.921</td>
<td>0.928</td>
<td>0.809</td>
<td>0.273</td>
<td>0.23</td>
</tr>
<tr>
<td><bold>T4 (Go_emotions)</bold></td>
<td>0.882</td>
<td>0.87</td>
<td>0.928</td>
<td>0.935</td>
<td>0.816</td>
<td>0.191</td>
<td>0.127</td>
</tr>
<tr>
<td><bold>T5 (DAIC-WOZ)</bold></td>
<td>0.895</td>
<td>0.883</td>
<td>0.932</td>
<td>0.941</td>
<td>0.829</td>
<td>0.172</td>
<td>0.123</td>
</tr>
<tr>
<td><bold>T6 (Kaggle-MH</bold>)</td>
<td>0.902</td>
<td>0.888</td>
<td>0.931</td>
<td>0.94</td>
<td>0.834</td>
<td>0.112</td>
<td>0.091</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Continual learning performance of Experience Replay across all datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Task</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>T1 (DASH)</bold></td>
<td>0.845</td>
<td>0.835</td>
<td>0.887</td>
<td>0.895</td>
<td>0.77</td>
<td>0.199</td>
<td>0.139</td>
</tr>
<tr>
<td><bold>T2 (Dreaddit)</bold></td>
<td>0.858</td>
<td>0.845</td>
<td>0.893</td>
<td>0.905</td>
<td>0.782</td>
<td>0.195</td>
<td>0.136</td>
</tr>
<tr>
<td><bold>T3 (SWMH)</bold></td>
<td>0.864</td>
<td>0.848</td>
<td>0.892</td>
<td>0.908</td>
<td>0.784</td>
<td>0.184</td>
<td>0.123</td>
</tr>
<tr>
<td><bold>T4 (Go_emotions)</bold></td>
<td>0.872</td>
<td>0.857</td>
<td>0.902</td>
<td>0.913</td>
<td>0.795</td>
<td>0.182</td>
<td>0.119</td>
</tr>
<tr>
<td><bold>T5 (DAIC-WOZ)</bold></td>
<td>0.875</td>
<td>0.862</td>
<td>0.902</td>
<td>0.915</td>
<td>0.792</td>
<td>0.171</td>
<td>0.117</td>
</tr>
<tr>
<td><bold>T6 (Kaggle-MH)</bold></td>
<td>0.867</td>
<td>0.852</td>
<td>0.895</td>
<td>0.912</td>
<td>0.783</td>
<td>0.182</td>
<td>0.121</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Continual learning performance of Synaptic Intelligence across all datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Task</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>T1 (DASH)</bold></td>
<td>0.866</td>
<td>0.845</td>
<td>0.891</td>
<td>0.905</td>
<td>0.785</td>
<td>0.085</td>
<td>0.038</td>
</tr>
<tr>
<td><bold>T2 (Dreaddit)</bold></td>
<td>0.845</td>
<td>0.833</td>
<td>0.878</td>
<td>0.892</td>
<td>0.776</td>
<td>0.088</td>
<td>0.042</td>
</tr>
<tr>
<td><bold>T3 (SWMH)</bold></td>
<td>0.851</td>
<td>0.843</td>
<td>0.877</td>
<td>0.911</td>
<td>0.795</td>
<td>0.085</td>
<td>0.035</td>
</tr>
<tr>
<td><bold>T4 (Go_emotions)</bold></td>
<td>0.848</td>
<td>0.832</td>
<td>0.851</td>
<td>0.895</td>
<td>0.772</td>
<td>0.087</td>
<td>0.039</td>
</tr>
<tr>
<td><bold>T5 (DAIC-WOZ)</bold></td>
<td>0.858</td>
<td>0.842</td>
<td>0.887</td>
<td>0.903</td>
<td>0.781</td>
<td>0.083</td>
<td>0.037</td>
</tr>
<tr>
<td><bold>T6 (Kaggle-MH)</bold></td>
<td>0.851</td>
<td>0.839</td>
<td>0.877</td>
<td>0.901</td>
<td>0.762</td>
<td>0.081</td>
<td>0.036</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Continual learning performance of EWC across all datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Task</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>T1 (DASH)</bold></td>
<td>0.856</td>
<td>0.841</td>
<td>0.863</td>
<td>0.902</td>
<td>0.780</td>
<td>0.173</td>
<td>0.126</td>
</tr>
<tr>
<td><bold>T2 (Dreaddit)</bold></td>
<td>0.854</td>
<td>0.839</td>
<td>0.886</td>
<td>0.907</td>
<td>0.785</td>
<td>0.084</td>
<td>0.138</td>
</tr>
<tr>
<td><bold>T3 (SWMH)</bold></td>
<td>0.859</td>
<td>0.852</td>
<td>0.885</td>
<td>0.918</td>
<td>0.803</td>
<td>0.081</td>
<td>0.132</td>
</tr>
<tr>
<td><bold>T4 (Go_emotions)</bold></td>
<td>0.859</td>
<td>0.848</td>
<td>0.895</td>
<td>0.921</td>
<td>0.771</td>
<td>0.167</td>
<td>0.133</td>
</tr>
<tr>
<td><bold>T5 (DAIC-WOZ)</bold></td>
<td>0.866</td>
<td>0.851</td>
<td>0.895</td>
<td>0.910</td>
<td>0.792</td>
<td>0.179</td>
<td>0.124</td>
</tr>
<tr>
<td><bold>T6 (Kaggle-MH)</bold></td>
<td>0.874</td>
<td>0.854</td>
<td>0.898</td>
<td>0.927</td>
<td>0.793</td>
<td>0.151</td>
<td>0.135</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Continual learning performance of Finetuning across all datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Task</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>T1 (DASH)</bold></td>
<td>0.725</td>
<td>0.775</td>
<td>0.765</td>
<td>0.775</td>
<td>0.645</td>
<td>0.311</td>
<td>0.265</td>
</tr>
<tr>
<td><bold>T2 (Dreaddit)</bold></td>
<td>0.763</td>
<td>0.741</td>
<td>0.797</td>
<td>0.805</td>
<td>0.681</td>
<td>0.267</td>
<td>0.251</td>
</tr>
<tr>
<td><bold>T3 (SWMH)</bold></td>
<td>0.755</td>
<td>0.739</td>
<td>0.785</td>
<td>0.803</td>
<td>0.677</td>
<td>0.263</td>
<td>0.257</td>
</tr>
<tr>
<td><bold>T4 (Go_emotions)</bold></td>
<td>0.735</td>
<td>0.735</td>
<td>0.775</td>
<td>0.799</td>
<td>0.665</td>
<td>0.325</td>
<td>0.311</td>
</tr>
<tr>
<td><bold>T5 (DAIC-WOZ)</bold></td>
<td>0.748</td>
<td>0.745</td>
<td>0.781</td>
<td>0.785</td>
<td>0.655</td>
<td>0.322</td>
<td>0.255</td>
</tr>
<tr>
<td><bold>T6 (Kaggle-MH)</bold></td>
<td>0.731</td>
<td>0.722</td>
<td>0.745</td>
<td>0.798</td>
<td>0.672</td>
<td>0.326</td>
<td>0.267</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Methods that incorporate explicit memory mechanisms or parameter regularization, such as GEM and EWC, demonstrate better retention than Finetuning or LWF across all six tasks, confirming that constraining weight drift mitigates forgetting. However, these approaches still exhibit limited calibration stability, as indicated by elevated Brier and ECE values across late tasks. Synaptic Intelligence achieves moderate balance between accuracy and calibration, but its adaptation saturates beyond mid-sequence datasets, revealing difficulty in scaling to domain shifts. In contrast, MH-Freeze consistently maintains high discriminative accuracy while achieving the lowest calibration errors.</p>
<p>From T1 (DASH) to T6 (Kaggle-MH), most baselines show mild fluctuations in F1-Score and AUROC due to changing dataset characteristics and label imbalance. However, MH-Freeze exhibits smooth performance progression, achieving improvements in accuracy and F1-Score compared with the strongest baseline (GEM). This trend demonstrates strong forward transfer and minimal backward interference. Moreover, the consistently low Brier (<inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mo>&#x2248;</mml:mo></mml:math></inline-formula>0.06) and ECE (<inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mo>&#x2248;</mml:mo></mml:math></inline-formula>0.014) values emphasize its reliability in producing well-calibrated probabilities, critical in sensitive applications such as mental-health risk prediction, where overconfident misclassifications can have severe consequences.</p>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows accuracy trends for all methods across the six datasets in sequence. The accuracy curves show a clear and persistent margin for the proposed MH-Freeze approach on every task. From <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>T</mml:mi><mml:mn>1</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mi>T</mml:mi><mml:mn>6</mml:mn></mml:math></inline-formula>, the accuracy of MH-Freeze rises from 0.898 to 0.925 (an increase of 3.0%), which indicates that knowledge gained on earlier tasks is retained while useful information from later tasks is added. In contrast, Finetuning changes only slightly (0.725 to 0.731; 0.8% increase) and shows signs of forgetting as new tasks are introduced. Regularization methods yield smaller gains, EWC improves by 2.1% and SI decreases by 1.7%, suggesting limited ability to adapt to the domain shifts in this sequence. Replay (ER), distillation (LWF), and gradient-projection (GEM) make training more stable, but they still perform worse than the proposed method. The highest final accuracy of 0.925 and the steady rise across datasets suggest that freezing the shared GCN-Transformer encoder and using separate heads for each dataset reduces interference and supports reliable forward transfer.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Performance comparison of Continual Learning techniques in terms of Accuracy</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_75119-fig-3.tif"/>
</fig>
<p>Similarly, a comparison based on F1-Score is presented in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. This metric is informative under class imbalance and helps assess whether decisions remain balanced as new datasets are introduced. The F1-Score of MH-Freeze increases from 0.886 at <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mi>T</mml:mi><mml:mn>1</mml:mn></mml:math></inline-formula> to 0.923 at <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi>T</mml:mi><mml:mn>6</mml:mn></mml:math></inline-formula>, a 4.2% increase, indicating that earlier decision boundaries are preserved while new patterns are learned. The baselines follow a consistent ordering: Finetuning decreases from 0.775 to 0.722 (a 6.8% decrease), reflecting forgetting; EWC shows a small improvement of 1.5%; SI decreases slightly by 0.7%; ER and LWF achieve moderate increases of 2.0% and 4.0%, respectively; and GEM improves by 3.1% but remains below MH-Freeze. These results suggest that the multi-head freeze strategy maintains balanced predictions across datasets while enabling steady gains, which is the intended behavior in a continual-learning environment.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Performance comparison of Continual Learning techniques in terms of F1-Score</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_75119-fig-4.tif"/>
</fig>
<p>In addition, ROC curves are presented to provide an intuitive visualization of classification trade-offs across varying decision thresholds. Unlike single-value metrics, ROC curves reveal how each model balances the true-positive and false-positive rates, offering a deeper understanding of their discriminative behavior. The ROC curves shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> illustrates the comparative classification performance of all continual-learning baselines and the proposed MH-Freeze model across six sequential tasks (T1&#x2013;T6). Each subplot corresponds to a specific task, where the x-axis represents the False Positive Rate (FPR) and the y-axis represents the True Positive Rate (TPR). The ROC trajectories of all models are plotted within each panel, while only the area under the ROC curve (AUC) of the MH-Freeze model, representing the top-performing method, is explicitly annotated. The solid black curve corresponds to MH-Freeze and demonstrates its consistently superior separability across all tasks. In contrast, baseline models such as Finetuning, EWC, and SI are depicted with thinner colored curves to provide visual benchmarking and highlight relative performance differences. This visualization clearly demonstrates that MH-Freeze maintains stable and near-optimal discriminative capability across all incremental tasks, confirming its strong resistance to catastrophic forgetting and enhanced adaptability in continual-learning environments.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>ROC curves across six sequential tasks (T1&#x2013;T6) for continual-learning baselines and the proposed MH-Freeze model. Each task corresponds to a distinct dataset used in the continual-learning sequence</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_75119-fig-5.tif"/>
</fig>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Computational Efficiency and Model Complexity</title>
<p>The CNSGL framework is designed to remain computationally efficient while still using the same encoder as the continual-learning baselines. The shared encoder, comprising a two-layer GCN (hidden size 128) and a single Transformer attention block with four heads (<inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula>, FFN size 256), contains approximately 482 k learnable parameters, which are used in all models, including the proposed MH-Freeze. Each task also has a linear&#x2013;sigmoid classifier head with 129 parameters, giving a total of about 482 k&#x002B;129 parameters per model. The major difference is not in how many parameters exist, but in how many are updated during each new task. In the proposed framework, MH-Freeze freezes the encoder after first task (T1), and trains only the 129-parameter head, whereas baseline models continue to update all 482 k&#x002B;129 parameters for every new task. Training on T1, where the encoder and head are jointly optimized, takes about 19 min, while training on subsequent tasks (T2&#x2013;T6), where only the head is updated, completes in roughly 2 to 4 min per task. A single forward pass requires approximately 1.47 ms per post and about 1.37 million floating point operations per second (FLOPs).</p>
<p>All continual-learning baselines use the same encoder architecture for a fair comparison, so their inference FLOPs are identical to CNSGL. However, they differ substantially in how many parameters are updated during each new task and in the resulting training time. CNSGL (MH-Freeze) updates only 129 parameters per new task, yielding the lowest per-task training time, whereas baselines must update the full encoder (<inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mo>&#x2248;</mml:mo></mml:math></inline-formula> 482 k&#x002B;129 parameters), leading to longer training times despite identical inference complexity.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Ablation Study</title>
<p>An ablation study was performed to examine individual contribution of each component in the proposed CNSGL framework, where symbolic reasoning, causal enrichment, GCN-based message passing, Transformer attention pooling, and the continual-learning mechanism are removed one at a time. <xref ref-type="table" rid="table-10">Table 10</xref> summarizes the performance of these variants.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Ablation study evaluating the contribution of each component in the CNSGL framework</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Accuracy</th>
<th>F1-Score</th>
<th>AUPRC</th>
<th>AUROC</th>
<th>MCC</th>
<th>Brier</th>
<th>ECE</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>No Symbolic &#x0026; Causal (GCN&#x002B;Transformer&#x002B;CL)</bold></td>
<td>0.919</td>
<td>0.911</td>
<td>0.924</td>
<td>0.932</td>
<td>0.851</td>
<td>0.073</td>
<td>0.021</td>
</tr>
<tr>
<td><bold>No GCN (Symbolic-Causal&#x002B; Transformer&#x002B;CL)</bold></td>
<td>0.872</td>
<td>0.861</td>
<td>0.892</td>
<td>0.893</td>
<td>0.824</td>
<td>0.173</td>
<td>0.152</td>
</tr>
<tr>
<td><bold>No Transformer (Symbolic-Causal&#x002B; GCN&#x002B;CL)</bold></td>
<td>0.893</td>
<td>0.887</td>
<td>0.913</td>
<td>0.911</td>
<td>0.885</td>
<td>0.085</td>
<td>0.023</td>
</tr>
<tr>
<td><bold>No Continual Learning (Symbolic-Causal&#x002B;GCN&#x002B; Transformer)</bold></td>
<td>0.831</td>
<td>0.829</td>
<td>0.853</td>
<td>0.861</td>
<td>0.811</td>
<td>0.195</td>
<td>0.141</td>
</tr>
<tr>
<td><bold>Full CNSGL Symbolic-Causal&#x002B;GCN&#x002B; Transformer&#x002B;MH-Freeze)</bold></td>
<td>0.925</td>
<td>0923</td>
<td>0.947</td>
<td>0.965</td>
<td>0.873</td>
<td>0.060</td>
<td>0.014</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Removing symbolic tags and causal edges (&#x201C;No Symbolic &#x0026; Causal&#x201D;) yields a model that operates purely on GCN &#x002B; Transformer embeddings without structured risk concepts. While performance remains reasonably strong (F1 &#x003D; 0.911, AUPRC &#x003D; 0.924), a noticeable drop appears compared to the full system, particularly in calibration (Brier &#x003D; 0.073 vs. 0.060). This confirms that symbolic grounding provides clinically meaningful structure that enhances predictive reliability. When the GCN encoder is removed (&#x201C;No GCN&#x201D;), performance declines sharply across all metrics (F1 &#x003D; 0.861, AUPRC &#x003D; 0.892), and calibration degrades significantly (ECE &#x003D; 0.152). This indicates that graph-based message passing is essential for leveraging symbolic&#x2013;causal structure; replacing it with flat representations harms both accuracy and stability. Removing the Transformer attention pooler (&#x201C;No Transformer&#x201D;) further demonstrates the role of attention in extracting concept-level importance. Although the model still performs moderately well due to symbolic&#x2013;causal structure (F1 &#x003D; 0.887), it shows lower AUROC (0.911) and poorer calibration relative to the full framework.</p>
<p>The performance degrades drastically when continual learning is removed (&#x201C;No Continual Learning&#x201D;), where sequential fine-tuning leads to catastrophic forgetting (F1 &#x003D; 0.829, AUPRC &#x003D; 0.853, Brier &#x003D; 0.195). This highlights the necessity of the MH-Freeze strategy; without it, performance on earlier tasks collapses, and calibration becomes unstable. At last, the full CNSGL model, integrating symbolic tags, directional associations, GCN encoding, Transformer pooling, and MH-Freeze, achieves the strongest and most consistent performance across all metrics (F1 &#x003D; 0.923, AUROC &#x003D; 0.965, Brier &#x003D; 0.060, ECE &#x003D; 0.014). These results confirm that each component contributes meaningfully and that the full architecture offers the best balance of predictive accuracy, stability, and interpretability.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion and Future Work</title>
<p>The proposed Continual Neuro-Symbolic Graph Learning framework successfully integrates symbolic reasoning, causal inference, and continual learning to address the evolving nature of mental-health risk detection. By constructing symbolic graphs enriched with directional causal edges, the framework enables interpretable reasoning about risk factors and their interrelations. The hybrid encoder, comprising a two-layer GCN and a Transformer-based attention pooler, effectively captures both structural and contextual dependencies, producing discriminative yet interpretable graph-level embeddings. The MH-Freeze strategy, which freezes the shared encoder and attaches task-specific heads, ensures strong retention of prior knowledge while allowing efficient adaptation to new datasets. Experimental results across six tasks (datasets) validate its robustness, showing that MH-Freeze consistently achieves the highest AUROC and F1-Score values, alongside superior calibration metrics (Brier and ECE), compared to all other continual-learning baselines. These findings confirm that MH-Freeze mitigates catastrophic forgetting and sustains stable, generalizable decision boundaries across diverse domains. Ablation analysis further confirms that each component contributes meaningfully to overall performance and calibration, and that the MH-Freeze continual-learning scheme is particularly critical for preserving performance and stability as new tasks are introduced.</p>
<p>Despite its advantages, this work also has a few limitations. The symbolic tag vocabulary is kept intentionally small and manually curated to ensure clarity and cross-dataset consistency. This focused design works well for the current scenario, but future extensions could incorporate richer or domain-specific tags to capture more subtle risk cues in broader clinical or social media data. In addition, the directional PMI module models precedence-based associations rather than fully validated causal relations. Future work can address these points by learning richer tag sets in a data-driven way and by integrating stronger causal discovery or longitudinal validation to refine the directed edges.</p>
<p>Future extensions will broaden the scope of the framework in several ways. Integrating multimodal signals, such as speech, facial expressions, and physiological markers can enrich early detection by complementing text with non-textual cues. Enhancing causal graph enrichment with temporal and counterfactual reasoning can deepen interpretability and strengthen the reliability of causal claims. Adopting federated or other privacy-preserving continual learning schemes can enable secure training across distributed mental-health datasets without direct data sharing. These advances would move CNSGL toward a more explainable, adaptive, and ethically deployable system for real-world mental-health risk assessment.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2025-00518960) and in part by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS-2025-00563192).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Monalisa Jena and Noman Khan; methodology, Monalisa Jena; software, Monalisa Jena and Noman Khan; validation, Monalisa Jena and Noman Khan; formal analysis, Mi Young Lee; investigation, Seungmin Rho; data curation, Monalisa Jena; writing&#x2014;original draft preparation, Monalisa Jena; writing&#x2014;review and editing, Monalisa Jena and Noman Khan; visualization, Monalisa Jena and Noman Khan; supervision, Mi Young Lee and Seungmin Rho; project administration, Mi Young Lee; funding acquisition, Mi Young Lee and Seungmin Rho. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The SWMH (SuicideWatch and Mental Health), E-DAIC (DAIC-WOZ), and DASH-2020 datasets are available on request from the respective authors due to ethical and access restrictions (SWMH: <ext-link ext-link-type="uri" xlink:href="https://10.5281/zenodo.6476178">https://10.5281/zenodo.6476178</ext-link>, accessed on 16 May 2025; E-DAIC: <ext-link ext-link-type="uri" xlink:href="https://dcapswoz.ict.usc.edu/wwwedaic/">https://dcapswoz.ict.usc.edu/wwwedaic/</ext-link>, accessed on 16 May 2025; DASH-2020: <ext-link ext-link-type="uri" xlink:href="https://zenodo.org/record/4278895#.X7T6cgzY2w">https://zenodo.org/record/4278895#.X7T6cgzY2w</ext-link>, accessed on 18 May 2025), whereas the Dreaddit, Kaggle Mental_Health, and GoEmotions datasets are openly available via public repositories on Kaggle (Dreaddit: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/rishantenis/dreaddit-train-test">https://www.kaggle.com/datasets/rishantenis/dreaddit-train-test</ext-link>, accessed on 20 May 2025; Kaggle Mental Health: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/entenam/reddit-mental-health-dataset?resource=download-directory">https://www.kaggle.com/datasets/entenam/reddit-mental-health-dataset?resource=download-directory</ext-link>, accessed on 22 May 2025; GoEmotions: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/debarshichanda/goemotions">https://www.kaggle.com/datasets/debarshichanda/goemotions</ext-link>, accessed on 22 May 2025).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations</title>
<def-list>
<def-item>
<term>AUPRC</term>
<def>
<p>Area Under the Precision-Recall Curve</p>
</def>
</def-item>
<def-item>
<term>AUROC</term>
<def>
<p>Area Under the Receiver Operating Characteristic Curve</p>
</def>
</def-item>
<def-item>
<term>BERT</term>
<def>
<p>Bidirectional Encoder Representation from Transformers</p>
</def>
</def-item>
<def-item>
<term>CL</term>
<def>
<p>Continual Learning</p>
</def>
</def-item>
<def-item>
<term>CNSGL</term>
<def>
<p>Continual Neuro-Symbolic Graph Learning</p>
</def>
</def-item>
<def-item>
<term>DASH</term>
<def>
<p>Data Analytics for Smart Health</p>
</def>
</def-item>
<def-item>
<term>DL</term>
<def>
<p>Deep Learning</p>
</def>
</def-item>
<def-item>
<term>EWC</term>
<def>
<p>Elastic Weight Consolidation</p>
</def>
</def-item>
<def-item>
<term>ECI</term>
<def>
<p>Event Causality Identification</p>
</def>
</def-item>
<def-item>
<term>ECE</term>
<def>
<p>Expected Calibration Error</p>
</def>
</def-item>
<def-item>
<term>ER</term>
<def>
<p>Experience Replay</p>
</def>
</def-item>
<def-item>
<term>FP</term>
<def>
<p>False Positive</p>
</def>
</def-item>
<def-item>
<term>FPR</term>
<def>
<p>False Positive Rate</p>
</def>
</def-item>
<def-item>
<term>FN</term>
<def>
<p>False Negative</p>
</def>
</def-item>
<def-item>
<term>FFN</term>
<def>
<p>Feed-Forward Network</p>
</def>
</def-item>
<def-item>
<term>FLOP</term>
<def>
<p>Floating Point Operations Per second</p>
</def>
</def-item>
<def-item>
<term>GEM</term>
<def>
<p>Gradient Episodic Memory</p>
</def>
</def-item>
<def-item>
<term>GCN</term>
<def>
<p>Graph Convolutional Network</p>
</def>
</def-item>
<def-item>
<term>HAN</term>
<def>
<p>Hierarchical Attention Network</p>
</def>
</def-item>
<def-item>
<term>IoT</term>
<def>
<p>Internet of Things</p>
</def>
</def-item>
<def-item>
<term>LWF</term>
<def>
<p>Learning without Forgetting</p>
</def>
</def-item>
<def-item>
<term>MCC</term>
<def>
<p>Matthews Correlation Coefficient</p>
</def>
</def-item>
<def-item>
<term>MH-Freeze</term>
<def>
<p>Multi-Head Freeze</p>
</def>
</def-item>
<def-item>
<term>ML</term>
<def>
<p>Machine Learning</p>
</def>
</def-item>
<def-item>
<term>NLP</term>
<def>
<p>Natural Language processing</p>
</def>
</def-item>
<def-item>
<term>PMI</term>
<def>
<p>Point-wise Mutual Information</p>
</def>
</def-item>
<def-item>
<term>SI</term>
<def>
<p>Synaptic Intelligence</p>
</def>
</def-item>
<def-item>
<term>SVD</term>
<def>
<p>singular value decomposition</p>
</def>
</def-item>
<def-item>
<term>SWMH</term>
<def>
<p>SuicideWatch and Mental Health Collection</p>
</def>
</def-item>
<def-item>
<term>TF-IDF</term>
<def>
<p>Term Frequency Inverse Document Frequency</p>
</def>
</def-item>
<def-item>
<term>TPP</term>
<def>
<p>Temporal Point Process</p>
</def>
</def-item>
<def-item>
<term>TN</term>
<def>
<p>True Negative</p>
</def>
</def-item>
<def-item>
<term>TP</term>
<def>
<p>True Positive</p>
</def>
</def-item>
<def-item>
<term>TPR</term>
<def>
<p>True Positive Rate</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mihalcea</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wilson</surname> <given-names>SR</given-names></string-name></person-group>. <chapter-title>Text-based detection and understanding of changes in mental health</chapter-title>. In: <source>Social informatics (SocInfo 2018)</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2018</year>. p. <fpage>176</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-01159-8_17</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hossain</surname> <given-names>E</given-names></string-name>, <string-name><surname>Alazeb</surname> <given-names>A</given-names></string-name>, <string-name><surname>Almudawi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Alshehri</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gazi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Faruque</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Forecasting mental stress using machine learning algorithms</article-title>. <source>Comput Mater Contin</source>. <year>2022</year>;<volume>72</volume>(<issue>3</issue>):<fpage>4945</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2022.027058</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gabriel</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hendler</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Mental-LLM: leveraging large language models for mental health prediction via online text data</article-title>. <source>Proc ACM Interact Mob Wearable Ubiquitous Technol</source>. <year>2024</year>;<volume>8</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>32</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hossain</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Hossain</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Mridha</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Safran</surname> <given-names>M</given-names></string-name>, <string-name><surname>Alfarhood</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Multi-task opinion enhanced hybrid BERT model for mental health analysis</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>3332</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-86124-6</pub-id>; <pub-id pub-id-type="pmid">39870711</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Omarov</surname> <given-names>B</given-names></string-name>, <string-name><surname>Narynov</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhumanov</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Artificial intelligence-enabled chatbots in mental health: a systematic review</article-title>. <source>Comput Mater Contin</source>. <year>2023</year>;<volume>74</volume>(<issue>3</issue>):<fpage>5105</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2023.034655</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tejaswini</surname> <given-names>V</given-names></string-name>, <string-name><surname>Sathya Babu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Sahoo</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Depression detection from social media text analysis using natural language processing techniques and hybrid deep learning model</article-title>. <source>ACM Trans Asian Low-Resour Lang Inf Process</source>. <year>2024</year>;<volume>23</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3569580</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vajrobol</surname> <given-names>V</given-names></string-name>, <string-name><surname>Saxena</surname> <given-names>GJ</given-names></string-name>, <string-name><surname>Pundir</surname> <given-names>A</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gaurav</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bansal</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A comprehensive survey on federated learning applications in computational mental healthcare</article-title>. <source>Comput Model Eng Sci</source>. <year>2025</year>;<volume>142</volume>(<issue>1</issue>):<fpage>49</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2024.056500</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thekkekara</surname> <given-names>JP</given-names></string-name>, <string-name><surname>Yongchareon</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liesaputra</surname> <given-names>V</given-names></string-name></person-group>. <article-title>An attention-based CNN-BiLSTM model for depression detection on social media text</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>249</volume>:<fpage>123834</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.123834</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mazurets</surname> <given-names>O</given-names></string-name>, <string-name><surname>Tymofiiev</surname> <given-names>I</given-names></string-name>, <string-name><surname>Dydo</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Approach for using neural network BERT-GPT2 dual transformer architecture for detecting persons depressive state</article-title>. In: <conf-name>VI International Scientific and Practical Conference; 2024 Nov 15</conf-name>; <publisher-loc>Bologna, Italy</publisher-loc>. p. <fpage>147</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Helmy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nassar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ramdan</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Depression detection for Twitter users using sentiment analysis in English and Arabic tweets</article-title>. <source>Artif Intell Med</source>. <year>2024</year>;<volume>147</volume>:<fpage>102716</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.artmed.2023.102716</pub-id>; <pub-id pub-id-type="pmid">38184345</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kodati</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tene</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Advancing mental health detection in texts via multi-task learning with soft-parameter sharing transformers</article-title>. <source>Neural Comput Appl</source>. <year>2025</year>;<volume>37</volume>(<issue>5</issue>):<fpage>3077</fpage>&#x2013;<lpage>110</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-024-10753-7</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huai</surname> <given-names>T</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Recent advances of foundation language models-based continual learning: a survey</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>57</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3705725</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thuseethan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rajasegarar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yearwood</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep continual learning for emerging emotion recognition</article-title>. <source>IEEE Trans Multimedia</source>. <year>2021</year>;<volume>24</volume>:<fpage>4367</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2021.3116434</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Mascolo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Andr&#x00E9;</surname> <given-names>E</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep learning for mobile mental health: challenges and recent advances</article-title>. <source>IEEE Signal Process Mag</source>. <year>2021</year>;<volume>38</volume>(<issue>6</issue>):<fpage>96</fpage>&#x2013;<lpage>105</lpage>. doi:<pub-id pub-id-type="doi">10.1109/msp.2021.3099293</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hemmatirad</surname> <given-names>K</given-names></string-name>, <string-name><surname>Bagherzadeh</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fazl-Ersi</surname> <given-names>E</given-names></string-name>, <string-name><surname>Vahedian</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Detection of mental illness risk on social media through multi-level SVMs</article-title>. In: <conf-name>Proceedings of the 8th Iranian Joint Congress on Fuzzy and Intelligent Systems (CFIS); 2020 Sep 2&#x2013;4</conf-name>; <publisher-loc>Mashhad, Iran</publisher-loc>. p. <fpage>116</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mohd</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Mental health safety and depression detection in social media text data: a classification approach based on a deep learning model</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>:<fpage>63284</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3559170</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Garg</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Mental health analysis in social media posts: a survey</article-title>. <source>Arch Comput Methods Eng</source>. <year>2023</year>;<volume>30</volume>(<issue>3</issue>):<fpage>1819</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s11831-022-09863-z</pub-id>; <pub-id pub-id-type="pmid">36619138</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Skaik</surname> <given-names>R</given-names></string-name>, <string-name><surname>Inkpen</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Using social media for mental health surveillance: a review</article-title>. <source>ACM Comput Surv</source>. <year>2020</year>;<volume>53</volume>(<issue>6</issue>):<fpage>1</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3422824</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Benrouba</surname> <given-names>F</given-names></string-name>, <string-name><surname>Boudour</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Emotional sentiment analysis of social media content for mental health safety</article-title>. <source>Soc Netw Anal Min</source>. <year>2023</year>;<volume>13</volume>(<issue>1</issue>):<fpage>17</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s13278-022-01000-9</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Trade-offs between machine learning and deep learning for mental illness detection on social media</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>14497</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-99167-6</pub-id>; <pub-id pub-id-type="pmid">40281061</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>De Choudhury</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kiciman</surname> <given-names>E</given-names></string-name></person-group>. <article-title>The language of social support in social media and its effect on suicidal ideation risk</article-title>. In: <conf-name>Proceedings of the International AAAI Conference on Web and Social Media</conf-name>. <publisher-loc>Palo Alto, CA, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>; <year>2017</year>. p. <fpage>32</fpage>&#x2013;<lpage>41</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Counterfactual neural temporal point process for estimating causal influence of misinformation on social media</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2022</year>;<volume>35</volume>:<fpage>10643</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Si</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A survey of event causality identification: taxonomy, challenges, assessment, and prospects</article-title>. <source>ACM Comput Surv</source>. <year>2025</year>;<volume>58</volume>(<issue>3</issue>):<fpage>59</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3756009</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kipf</surname> <given-names>TN</given-names></string-name></person-group>. <article-title>Semi-supervised classification with graph convolutional networks</article-title>. <comment>arXiv:1609.02907</comment>. <year>2016</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hamilton</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ying</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Leskovec</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Inductive representation learning on large graphs</article-title>. In: <conf-name>NIPS&#x2019;17: Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc.</publisher-name>; <year>2017</year>. p. <fpage>1025</fpage>&#x2013;<lpage>35</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Graph convolutional networks for text classification</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence</conf-name>. <publisher-loc>Palo Alto, CA, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>; <year>2019</year>. Vol. 33. p. <fpage>7370</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v33i01.33017370</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Subakan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ravanelli</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cornell</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bronzi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Attention is all you need in speech separation</article-title>. In: <conf-name>Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>21</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Attention is all you need</article-title>. In: <conf-name>NIPS&#x2019;17: Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc.</publisher-name>; <year>2017</year>. p. <fpage>6000</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>M-W</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>K</given-names></string-name>, <string-name><surname>Toutanova</surname> <given-names>K</given-names></string-name></person-group>. <article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>. In: <conf-name>Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT)</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2019</year>. p. <fpage>4171</fpage>&#x2013;<lpage>86</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Dyer</surname> <given-names>C</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name>, <string-name><surname>Smola</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hovy</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Hierarchical attention networks for document classification</article-title>. In: <conf-name>Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT)</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2016</year>. p. <fpage>1480</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gamel</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Talaat</surname> <given-names>FM</given-names></string-name></person-group>. <article-title>SleepSmart: an IoT-enabled continual learning algorithm for intelligent sleep enhancement</article-title>. <source>Neural Comput Appl</source>. <year>2024</year>;<volume>36</volume>(<issue>8</issue>):<fpage>4293</fpage>&#x2013;<lpage>309</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-023-09310-5</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>CS</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>AY</given-names></string-name></person-group>. <article-title>Clinical applications of continual learning machine learning</article-title>. <source>Lancet Digit Health</source>. <year>2020</year>;<volume>2</volume>(<issue>6</issue>):<fpage>e279</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1016/s2589-7500(20)30102-3</pub-id>; <pub-id pub-id-type="pmid">33328120</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>C-H</given-names></string-name>, <string-name><surname>Jha</surname> <given-names>NK</given-names></string-name></person-group>. <article-title>DOCTOR: a multi-disease detection continual learning framework based on wearable medical sensors</article-title>. <source>ACM Trans Embed Comput Syst</source>. <year>2024</year>;<volume>23</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3679050</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nie</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>M</given-names></string-name>, <string-name><surname>Su</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>A</given-names></string-name></person-group>. <article-title>I-GCN: incremental graph convolution network for conversation emotion detection</article-title>. <source>IEEE Trans Multimedia</source>. <year>2021</year>;<volume>24</volume>:<fpage>4471</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2021.3118881</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kaur</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bhardwaj</surname> <given-names>R</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Garg</surname> <given-names>M</given-names></string-name>, <string-name><surname>Saxena</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Causal categorization of mental health posts using transformers</article-title>. In: <conf-name>Proceedings of the 14th Annual Meeting of the Forum for Information Retrieval Evaluation</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2022</year>. p. <fpage>43</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kodati</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tene</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Identifying suicidal emotions on social media through transformer-based deep learning</article-title>. <source>Appl Intell</source>. <year>2023</year>;<volume>53</volume>(<issue>10</issue>):<fpage>11885</fpage>&#x2013;<lpage>11917</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10489-022-04060-8</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Neuro Symbolic AI in personalized mental health therapy: bridging cognitive science and computational psychiatry</article-title>. <source>World J Adv Res Rev</source>. <year>2023</year>;<volume>19</volume>(<issue>2</issue>):<fpage>1663</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.30574/wjarr.2023.19.2.1516</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhuang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A causality-driven graph convolutional network for postural abnormality diagnosis in Parkinsonians</article-title>. <source>IEEE Trans Med Imaging</source>. <year>2023</year>;<volume>42</volume>(<issue>12</issue>):<fpage>3752</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmi.2023.3305378</pub-id>; <pub-id pub-id-type="pmid">37581959</pub-id></mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Bhuyan</surname> <given-names>BP</given-names></string-name>, <string-name><surname>Ramdane-Cherif</surname> <given-names>A</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>TP</given-names></string-name>, <string-name><surname>Tomar</surname> <given-names>R</given-names></string-name></person-group>. <chapter-title>Neuro-Symbolic AI: the integration of continuous learning and discrete reasoning</chapter-title>. In: <source>Neuro-symbolic artificial intelligence: bridging logic and learning</source>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>29</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-97-8171-3_3</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dalkic</surname> <given-names>H</given-names></string-name></person-group>. <article-title>CognEmoSense: a continual learning and context-aware EEG emotion recognition system using transformer-augmented brain-state modeling</article-title>. <source>J Brain Sci Ment Health</source>. <year>2025</year>;<volume>1</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Patan&#x00E8;</surname> <given-names>G</given-names></string-name>, <string-name><surname>Sorrenti</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bellitto</surname> <given-names>G</given-names></string-name>, <string-name><surname>Palazzo</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Continual learning strategies for personalized mental well-being monitoring from mobile sensing data</article-title>. In: <conf-name>Proceedings of the International Workshop on Personalized Incremental Learning in Medicine</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2025</year>. p. <fpage>9</fpage>&#x2013;<lpage>17</lpage>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Febrinanto</surname> <given-names>FG</given-names></string-name>, <string-name><surname>Simango</surname> <given-names>A</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tyagi</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Refined causal graph structure learning via curvature for brain disease classification</article-title>. <source>Artif Intell Rev</source>. <year>2025</year>;<volume>58</volume>(<issue>8</issue>):<fpage>222</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-025-11231-9</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gosala</surname> <given-names>B</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>AR</given-names></string-name>, <string-name><surname>Tiwari</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>M</given-names></string-name></person-group>. <article-title>GCN-LSTM: a hybrid graph convolutional network model for schizophrenia classification</article-title>. <source>Biomed Signal Process Control</source>. <year>2025</year>;<volume>105</volume>(<issue>1</issue>):<fpage>107657</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.bspc.2025.107657</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>L-C</given-names></string-name></person-group>. <article-title>An extended TF-IDF method for improving keyword extraction in traditional corpus-based research: an example of a climate change corpus</article-title>. <source>Data Knowl Eng</source>. <year>2024</year>;<volume>153</volume>(<issue>2</issue>):<fpage>102322</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.datak.2024.102322</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>KC-C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>X</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A principled decomposition of pointwise mutual information for intention template discovery</article-title>. In: <conf-name>Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM)</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2023</year>. p. <fpage>1746</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ghorbani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Baghshah</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Rabiee</surname> <given-names>HR</given-names></string-name></person-group>. <article-title>MGCN: semi-supervised classification in multi-layer graphs with graph convolutional networks</article-title>. In: <conf-name>Proceedings of the 2019 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>208</fpage>&#x2013;<lpage>11</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Point Transformer V2: grouped vector attention and partition-based pooling</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2022</year>;<volume>35</volume>:<fpage>33330</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Terven</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cordova-Esparza</surname> <given-names>D-M</given-names></string-name>, <string-name><surname>Romero-Gonz&#x00E1;lez</surname> <given-names>J-A</given-names></string-name>, <string-name><surname>Ram&#x00ED;rez-Pedraza</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ch&#x00E1;vez-Urbiola</surname> <given-names>EA</given-names></string-name></person-group>. <article-title>A comprehensive survey of loss functions and metrics in deep learning</article-title>. <source>Artif Intell Rev</source>. <year>2025</year>;<volume>58</volume>(<issue>7</issue>):<fpage>195</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-025-11198-7</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ghosh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Misra</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ghosh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Podder</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Utilizing social media for identifying drug addiction and recovery intervention</article-title>. In: <conf-name>2020 IEEE International Conference on Big Data (Big Data)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2020</year>. p. <fpage>3413</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Demszky</surname> <given-names>D</given-names></string-name>, <string-name><surname>Movshovitz-Attias</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ko</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cowen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nemade</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ravi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>GoEmotions: a dataset of fine-grained emotions</article-title>. In: <conf-name>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL)</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2020</year>. p. <fpage>4040</fpage>&#x2013;<lpage>54</lpage>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rani</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>K</given-names></string-name>, <string-name><surname>Subramani</surname> <given-names>S</given-names></string-name></person-group>. <article-title>From posts to knowledge: annotating a pandemic-era Reddit dataset to navigate mental health narratives</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>4</issue>):<fpage>1547</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app14041547</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Turcan</surname> <given-names>E</given-names></string-name>, <string-name><surname>McKeown</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Dreaddit: a Reddit dataset for stress analysis in social media</article-title>. <comment>arXiv:1911.00133</comment>. <year>2019</year>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cambria</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Suicidal ideation and mental disorder detection with attentive relation networks</article-title>. <source>Neural Comput Appl</source>. <year>2022</year>;<volume>34</volume>:<fpage>10309</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-021-06208-y</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ringeval</surname> <given-names>F</given-names></string-name>, <string-name><surname>Schuller</surname> <given-names>BW</given-names></string-name>, <string-name><surname>Valstar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cowie</surname> <given-names>R</given-names></string-name>, <string-name><surname>Kaya</surname> <given-names>H</given-names></string-name>, <string-name><surname>Amiriparian</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>AVEC 2019 workshop and challenge: state-of-mind, detecting depression with AI, and cross-cultural affect recognition</article-title>. In: <conf-name>Proceedings of the 9th International on Audio/Visual Emotion Challenge and Workshop</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2019</year>. p. <fpage>3</fpage>&#x2013;<lpage>12</lpage>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gratch</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lucas</surname> <given-names>GM</given-names></string-name>, <string-name><surname>King</surname> <given-names>A</given-names></string-name>, <string-name><surname>Morency</surname> <given-names>L-P</given-names></string-name></person-group>. <article-title>The Distress Analysis Interview Corpus of human and computer interviews</article-title>. In: <conf-name>Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC 2014)</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2014</year>. p. <fpage>3123</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kirkpatrick</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pascanu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Rabinowitz</surname> <given-names>N</given-names></string-name>, <string-name><surname>Veness</surname> <given-names>J</given-names></string-name>, <string-name><surname>Desjardins</surname> <given-names>G</given-names></string-name>, <string-name><surname>Rusu</surname> <given-names>AA</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Overcoming catastrophic forgetting in neural networks</article-title>. <source>Proc Natl Acad Sci U S A</source>. <year>2017</year>;<volume>114</volume>(<issue>13</issue>):<fpage>3521</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1073/pnas.1611835114</pub-id>; <pub-id pub-id-type="pmid">28292907</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lopez-Paz</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ranzato</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Gradient episodic memory for continual learning</article-title>. In: <conf-name>NIPS&#x2019;17: Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc.</publisher-name>; <year>2017</year>. p. <fpage>6470</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hoiem</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Learning without forgetting</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2017</year>;<volume>40</volume>(<issue>12</issue>):<fpage>2935</fpage>&#x2013;<lpage>47</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2017.2773081</pub-id>; <pub-id pub-id-type="pmid">29990101</pub-id></mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rolnick</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ahuja</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schwarz</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lillicrap</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wayne</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Experience replay for continual learning</article-title>. In: <conf-name>Proceedings of the 33rd International Conference on Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc.</publisher-name>; <year>2019</year>. p. <fpage>350</fpage>&#x2013;<lpage>60</lpage>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>De Lange</surname> <given-names>M</given-names></string-name>, <string-name><surname>Aljundi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Masana</surname> <given-names>M</given-names></string-name>, <string-name><surname>Parisot</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>X</given-names></string-name>, <string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A continual learning survey: defying forgetting in classification tasks</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2021</year>;<volume>44</volume>(<issue>7</issue>):<fpage>3366</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2021.3057446</pub-id>; <pub-id pub-id-type="pmid">33544669</pub-id></mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zenke</surname> <given-names>F</given-names></string-name>, <string-name><surname>Poole</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ganguli</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Continual learning through synaptic intelligence</article-title>. In: <conf-name>ICML&#x2019;17: Proceedings of the 34th International Conference on Machine Learning; 2017 Aug 6&#x2013;11</conf-name>; <publisher-loc>Sydney, NSW, Australia</publisher-loc>. p. <fpage>3987</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chicco</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jurman</surname> <given-names>G</given-names></string-name></person-group>. <article-title>The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation</article-title>. <source>BMC Genomics</source>. <year>2020</year>;<volume>21</volume>(<issue>1</issue>):<fpage>6</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s12864-019-6413-7</pub-id>; <pub-id pub-id-type="pmid">31898477</pub-id></mixed-citation></ref>
</ref-list>
</back></article>















